Full transcript
The $200 Power User Confession
0:00I was a Claude Max subscriber, I paid $200
0:02a month, I loved it.
0:03I was a beast power user, using hundreds
0:06of millions of tokens per day sometimes.
0:08I can easily say that
0:09just over the last 12 months,
0:11Claude Code went from pretty neat to "OMG
0:14can't live without it".
0:15Something as a builder and solopreneur I
0:17genuinely couldn't do without.
0:18And that is exactly why I
0:20canceled my subscription.
0:21Because I realized there was something
0:23seriously wrong with a
0:24dependency on Enthropic.
0:25Let me explain.
0:26You see over the last 30
0:27days of my subscription,
0:29on two of them I surpassed 700 million
0:31tokens per day, and other
0:32days were well on their way.
0:34If we do all the math here
0:35with the general API cost,
0:37those two days alone
0:38would have cost over $50,000,
0:40around $2,000 if you
0:42factor in all the token caching.
0:43But we'll get to that later.
0:44But even with the discounted cash tokens,
0:46we're still looking at
0:48approximately $1,800 per month,
0:49just for two days of usage, and my
0:51subscription only costs $200.
0:53Something is seriously
0:54wrong with the economics here,
0:55and I don't want to be hooked on Cloud
0:56when the other shoe drops.
0:58How was I getting all
0:59of this for $200 a month?
1:00Someone asked to foot
1:01the energy bill here.
1:02AI is not free.
1:04Yet the more I use Claude Code, the more
1:06my work and business depended on it.
1:08Until I finally started thinking about
1:09what would happen if that $200 plan
1:12disappeared next week.
1:13What if they changed the
1:14limits or the product itself?
1:15At that point, how easy
1:17would it be for me to leave?
1:18And over time that
1:19question started bothering me.
1:20A lot.
1:21Because a large part of my life was
1:22revolving around Claude.
1:23Fortunately, there's a way out of this.
1:25Today we have open-weight models that are
1:27not only surprisingly good,
1:28but they offer results
1:29comparable to Opus and Fable.
1:31Meanwhile, there are more and more
1:33open-source agent frameworks and
1:34harnesses that are model-agnostic, which
1:36allow you to get
1:36creative past the handful
1:38of models offered by Claude, or OpenAI, or
1:40any American frontier family of models.
1:43So I decided to try something.
1:44I want to see whether I can replace my
1:46$200 Claude subscription
1:47with open models and a
1:48smarter workflow, without giving up the
1:50things that made Claude so damn useful to
1:53me in the first place.
1:54Let's start with the token inference.
The Token Math ($50,000 For Two Days)
1:56As I said, those two days alone would
1:57have cost approximately $1,800.
2:00But that's only two days of my 30-day
2:02subscription that I pay, in total, $200.
2:05So I started digging into
2:06that because it seemed strange.
2:08The API prices are public.
2:10If you took the amount of input and
2:11output tokens from those days
2:13and priced them as normal API usage,
2:15well, you see these high
2:16numbers in the tens of thousands.
2:17So either their API margins are obscene,
2:20or they're burning VC money
2:21to subsidize user inference,
2:23get us hooked, and hiding the
2:25actual cost of delivery here.
2:26It's probably a little bit of both.
2:28They have a good product.
2:29People want to pay for it.
2:30Claude Code has become one
2:31of the most useful products
2:32some of us have ever
2:33paid for in our lives.
2:34We use it constantly.
2:36Someone like me, I could sit down on a
2:37weekend, start building something,
2:39and not stop for 20 hours straight.
2:42The more you use it, the more you got
2:43from that easy $200.
2:45That's an incredible
2:46deal for a heavy user.
2:48But there is a flip side.
2:50Once you get used to working that way,
2:51you build your entire workflow around it
2:53and you become dependent.
2:54You stop thinking about
2:56the cost of every request.
2:57You stop worrying about
2:58which model you're using.
2:59You just open Claude
3:00Code and keep working.
3:01And I had definitely
3:02gotten to that point.
Subsidized Convenience Is The Lock-In
3:03Lazy with my AI usage because the Claude
3:06Code harness itself,
3:07with the subsidized
3:08inference, allowed me to just work
3:09and not have to really think about
3:11anything under the hood.
3:12So when I started looking at how much of
3:13my actual work was being done with Claude,
3:15I started wondering what
3:16happens if the deal changes.
3:18Because we've seen AI products or
3:20anything in tech
3:21change their limits before.
3:22A plan that looks generous
3:24today can look very different
3:25once the economic reality of
3:28running a business kicks in,
3:30which is often what happens when a
3:31company is about to IPO.
3:33But that's another story.
3:34So I'd already seen how
3:35quickly my usage could explode
3:37with over 700 million
3:38tokens in a single day.
3:39That's just what my
3:40workflow started to look like.
3:41So I had to ask myself
3:43the uncomfortable question.
3:44If Anthropic changed the price tomorrow,
3:46or tighten the limits, or
3:47change the way Claude Code worked,
3:49how much of my workflow would I have to
3:51rebuild or start from scratch?
3:52And that question bothered
3:54me a lot more than the $200.
3:57So I used Claude Code to do a lot.
3:58Build landing pages, sales pages,
4:00automated meta ad campaigns,
4:02generate visual assets for videos or
4:04presentations with
4:05emotion and hyperframes,
4:07and handle parts of my content workflow,
4:09including repurposing, metadata
4:10generation, et cetera.
4:11I did a lot.
4:12And then there were my side projects.
4:14I had a project I was
4:14working on with my son,
4:15where he created a world, built a story,
4:18and we turned it into a
4:18book, and a YouTube channel,
4:20and an educational alerts or
4:21read game on its own website.
4:23We made our own vibe-coded Patreon.
4:25And Claude was sitting
4:26right in the middle of that.
4:27If I didn't have the subsidized free
4:29inference, essentially,
4:30to build that, I wouldn't.
4:32Because I wasn't
4:32guaranteed to make any money.
4:33It was just a fun project.
4:35And it was with this project with my son,
The Audit: How Much Of My Life Ran On Claude
4:36I realized I would open Claude Code before
4:39I'd even think about
4:40whether there was another
4:41way to get the job done.
4:42I already knew how to use it.
4:43The tools were there.
4:44Context and memory was there.
4:46I built up all my workflows around it.
4:48So starting something new was easy.
4:50And that convenience
4:51adds up as an ignorance debt
4:53that eventually we must pay back.
4:55Because you've stopped
4:56questioning the tool,
4:57because it just works.
4:59And after enough time, you
5:00inevitably feel locked in,
5:01whether you realize it or not.
5:03The cost of switching
5:04just becomes too high.
5:05And that's what Anthropic wants.
5:07That's what any business wants.
5:08It's good business.
5:09Because it turned into even the thought
5:11of me canceling my Cloud subscription.
5:13It was more than just
5:14cutting another subscription
5:15from my life.
5:15I'm not just going to
5:16stop building with AI.
5:18I have to figure out an alternative.
5:19There's a cost to switching.
5:21Some of my workflows
5:22would be easy to move,
5:22but some of it would not.
5:23And truthfully, I didn't
5:24like how little I'd actually
5:26thought about that before.
5:27I'd spent all this time
5:28making Cloud more useful to me,
5:30while quietly making
5:31my own workflow harder
5:32to separate from it.
5:33And that's a weird position to end up in.
5:35You start with a tool
Why "Replace Claude" Is The Wrong Question
5:36because it's convenient,
5:37then you build around it.
5:38And eventually, that
5:39convenience becomes the reason
5:41you don't want to leave.
5:42It's like the relationship
5:43you know you need to end,
5:44but it's just too difficult.
5:46So eventually, two things
5:47were really bothering me--
5:48the future economics of the subscription
5:50and the company, the vendor lock-in,
5:52the fact that I built
5:53so much around a system
5:54that I didn't really control.
5:56So at this point, I asked myself,
5:58I don't want to just be stuck with Cloud,
6:00and I don't want to just resign to that,
6:02especially with all
6:03the geopolitical drama
6:04Anthropic has been cooking up.
6:05So I started looking for a way out.
6:07Now, the question arises, what the hell
6:09do we use instead of Cloud or OpenAI?
6:11Because I'm talking about Cloud here,
6:13but the truth is this
6:14applies to any AI provider,
6:16especially one that
6:17doesn't respect your privacy
6:18and is giving you something for free.
6:20Remember what they say,
6:21there is no such thing
6:22as a free lunch.
6:23And if lunches were tokens, well, I've
6:26gotten a lot of free
6:28lunches with Cloud code.
6:29But the answer to this
6:30question is not to replace Cloud
6:32or Codex with one thing.
6:34And that's how I was
6:35thinking about it the wrong way
6:36at the beginning.
6:37I was looking for a single tool that
6:38could replace Cloud.
6:39Oh, is Codex better?
6:40Is Gemini CLI better?
6:42But that doesn't really make sense,
6:44because why should one
6:45model or one family models
6:47have to provide everything?
6:48If I'm writing code,
6:50maybe I want just one model.
6:51If I'm planning a
6:52project, I might want another.
6:54And this is where things got interesting,
6:56because the options have
6:57changed a lot since I first
6:59started using Cloud code.
The New Stack: Open-Weight Models + OpenCode
7:01Open-weight models now
7:02can do some serious work,
7:03from Kimi K3 to GLM 5.2, GLM 5.3,
7:08DeepSeek V4, Q1 3.8,
7:10and you can connect all these to open
7:12source agent harnesses.
7:14And a tool like Open
7:15code, which has quickly
7:17become my Cloud code replacement,
7:19is what they call model agnostic.
7:20I can connect any
7:21provider, any model into there.
7:22And it's completely
7:23compatible with my Cloud code agentic
7:25workflows.
7:26So once I started
7:27thinking about it that way,
7:28the whole setup changed.
7:29I was able to start using
7:30Kimi K3 as more of a planner,
7:33an orchestrator.
7:34It's more expensive, but
7:35it's super intelligent.
7:37And it's actually a
7:37lot friendlier and easier
7:39to work with than Fable.
7:40Meanwhile, GLM 5.2 is cheaper, could
7:42handle a lot of the coding.
7:43And it's a very good
7:44creative writer as well.
7:46And for even simpler jobs, DeepSeek V4
7:48Flash may as well be free.
7:51It's fast, intelligent, and while I
7:53wouldn't want it planning my projects,
7:55it can handle the manual tasks perfectly.
7:58And getting comfortable with this
8:00meant that I could
8:01move away from the belief
8:02that I need to Cloud code
8:04to achieve the productivity
8:06that I desired.
8:07But there's another
8:07advantage I want to double click on.
Privacy And Profiling (And My Home Network)
8:09And that is actually
8:10choosing where the models come from.
8:11Now, as I was talking about, when
8:13I'm creating something
8:13for my business content,
8:14like what does the privacy really matter?
8:16Well, in a lot of ways,
8:17the output itself, whatever.
8:19Privacy, no privacy, whatever.
8:20But there's the actual creative process
8:23that comes from me, my
8:24unique way of thinking
8:25as a human being that is
8:27being profiled by tools
8:28like Cloud Code, Codex, Gemini.
8:30They're learning my psychology
8:32and making a profile about me.
8:34And once you start going deep into that,
8:35it's kind of freaky.
8:36So there's another reason to shift
8:38from these proprietary frontier services.
8:40And that begins to align with the
8:42inevitable next step,
8:43if you're a builder with AI, is what
8:45you can start using these harnesses
8:47for inside your own house.
8:48I started messing around
8:49with home security systems,
8:51with Raspberry Pis in the
8:52living room hosting video game
8:54emulators.
8:55And once I hit the stage, I realized
8:56I don't want an
8:56anthropic in my home network.
8:58No, thank you.
8:59And the models I can host
9:00locally just aren't there yet.
9:01So that's why I use
9:02Venice as a provider of all
9:04of my favorite open weight models.
9:06Because they have zero
9:07data retention agreements
9:08with the providers.
9:09Meaning nothing is
9:10being stored by the provider
9:12of my inference.
9:14No profile is being built.
9:16And even if there's a
9:16nefarious entity inside the GPUs
9:19that my provider doesn't even know about,
9:21they're not going to
9:21know it's coming from me.
9:22So this started
9:23creating a comprehensive view
9:25of how I want to
9:26approach my AI setups now.
9:28Instead of opening
9:29Cloud Code and asking it
9:30to do every little thing,
9:31I can now build workflows
9:33that use different
9:33models, different tools,
9:34different providers for
9:35different parts of the job.
9:37And that brings us into a
Smart Model Routing In The Workspace
9:38topic that most Cloud Code users
9:39have probably not thought about much.
9:41And that is smart model routing.
9:42I had no choice but to
9:43start digging into this.
9:44I'd gotten into the
9:45habit of treating Cloud
9:46like the default answer to every problem.
9:48So once I stopped doing that, the
9:50economics actually started
9:51to look very different.
9:52I could have one model think
9:53through what actually needs
9:54to happen and then hand
9:55off the plan to another model
9:57for implementation, a cheaper model.
10:00And if it's not a coding thing,
10:01I could use even simpler models.
10:02And then when I need a
10:03serious final review or audit,
10:05then I bring the big boys back in.
10:07So the most interesting part
10:09of something like smart model
10:10routing is that it can be
10:11handled in the workspace itself.
10:13So what might not be
10:14obvious as a Cloud Code user
10:15is that your AI agent can
10:17make calls to other models
10:19with instructions, with a
10:20system prompt, whatever.
10:22And what's really interesting about this,
10:23if you've only used Cloud Code or Codex
10:25and you've never
10:25built your own workspaces,
10:27is that the workspace
10:28itself can handle a lot
10:30of this smart model routing.
10:32The better your
10:32instructions are in AgentsMD,
10:35the smarter your agent can
10:36be in sending certain tasks
10:38to certain models.
10:39This will not only give you better
10:41performance and results,
10:43but it will save
10:43money, which is important
10:45when we're no longer
10:46using the subsidized inference
10:48by Anthropic.
10:49A well-structured
10:50workspace means my agent is going
10:52to be smart about the
10:53tasks it does versus the tasks
10:56it sends subagents to do.
10:58So as an example, I could
10:59use Cloud Code for a task.
11:01I could use Open Code for a task,
11:02and they might use a
11:03completely different amount of tokens,
11:05even if I'm using the same model.
11:07So let me explain.
11:08If I throw everything
11:09at Cloud Code and say,
11:10hey, figure this thing out, just do it.
11:12It'll probably use a lot of
11:13tokens in a fresh workspace
11:15with no instructions.
11:15It'll make its own instructions.
11:17It's so smart, it'll figure it out.
11:18It'll probably do a good job.
11:19But it's gonna burn
11:20through way more tokens
11:22than if you use the
11:23same model, Opus, Fable,
11:25in a tight workspace with
11:27instructions, with skills,
11:28with subagent profiles, et cetera,
11:30because it has context to work with.
11:32Not only that, but
11:32within the model workspace,
11:33you can integrate model routing.
11:35So if something needs to
11:36be done by a different agent
11:38in your instructions, you simply say,
11:39when we do this thing,
11:40send it to that model
11:41from that provider,
11:42here's the API key, done.
11:44You could do this with Cloud Code,
11:45but why would you even think to do that
11:47when Cloud Code is just so good
11:48and you're not paying
11:49for the actual inference?
11:51And a whole generation of AI users are
The Generation Getting Handicapped
11:53gonna be handicapped
11:54when reality comes knocking.
11:55And I don't wish that for you,
11:57so I'm glad you're watching this.
11:58And now is probably a good time to
12:00subscribe to my channel
12:02if you do wanna learn
12:03how to make workspaces
12:04that are not dependent on Cloud Code
12:05because that's what I explore in my
12:07videos on the channel.
12:07So thanks for hitting subscribe.
12:09So once I started
12:09understanding the different economics
12:11of running different
12:13models in a workspace,
12:15I wanted to know how far I could take it.
The Experiment: Venice Max And 300+ Models
12:17Could I actually
12:17rebuild enough of my workflows
12:19that I wouldn't miss Cloud?
12:21The answer was way more obvious than I
12:22could have expected.
12:23So I decided to actually test it out.
12:25I've canceled my Cloud subscription.
12:27It's time to move on.
12:28I'm running a Venice Max account,
12:30which is also $200 a month,
12:32and it gives me $225 worth
12:34of credits with their service.
12:35But there's over 300 different models
12:37that I can choose from
12:38from the top of the
12:39line open weight models
12:40like human K3 to GLM 5.2 to
12:43DeepSeek V4 Flash, et cetera.
12:45There's a lot to play with here
12:46that will allow me to save
12:48money and retain performance.
12:51It just means
12:51sacrificing the convenience.
12:53My goal is not to
12:54recreate Cloud perfectly,
12:55but it is to get similar results
12:58without having to spend a lot more money.
13:00The whole point of my
13:01experiment is to find out
13:03where the gaps in my own
13:04workspaces and workflows are.
13:06I have to be present
13:07with every single word
13:08and every single
13:08prompt because all of those
13:10are costing me tokens, which cost money.
13:12If a task doesn't
13:13need the strongest model,
13:14I shouldn't send the task there.
13:15If a bunch of context is
13:17gonna be used, can it be cached?
13:18If I'm recycling
13:19instructions, I need to make skills
13:21and put them in my
13:22behavioral context files.
13:23This is the stuff I'm most interested in
13:26because this is the true
13:27architecture of owning your own AI.
Own The Workspaces, Not The GPUs
13:29Sure, I might not be
13:30running the actual intelligence
13:32on a machine I own, but theoretically,
13:34in the not so distant future, I could.
13:36And more importantly, as
13:37long as there are vendors
13:38out there willing to provide
13:40for me the models at a cost
13:41that I wanna use, then the
13:43most important aspect for me
13:44is owning the workspaces themselves
13:46because they can be mine
13:48and no one can take those away
13:49from me, no one can shut those off.
13:51I'll just find another model provider.
13:52We live in a world of markets.
13:54Find the service
13:55provider that works for you.
13:56This way, if a particular
13:58tool or model or provider
13:59stops working for me,
14:00the rest of my workflow,
14:01the rest of my life with AI
14:03doesn't have to come down with it.
The Honest Exit: Will I Resubscribe?
14:05I don't know yet how far this will go,
14:07but that's what the
14:07next few weeks of my life
14:08are gonna all be
14:09about, not using cloud code
14:11and moving into agentic
14:12workspaces that I control
14:13with inference that I'm paying for,
14:15every single prompt accounted for.
14:18A responsibility that I
14:19have chosen to stop outsourcing
14:24can open weight
14:25models, open source agents
14:26and smart model routing
14:27replace what I was getting
14:28with cloud code, or am I going to end up
14:31resubscribing to
14:31Claude tail between my legs
14:33in a few weeks from now?
14:34That's the experiment.
14:36Stay tuned because I
14:37will be posting updates,
14:39letting you know how it goes.
AI Captains Academy (Come Say Hi)
14:40Finally, if you're interested in learning
14:42how to build these
14:43agentic workspaces yourself,
14:45I invite you to check out
14:46the AI Captain's Academy.
14:47There is a whole
14:48curriculum inside for understanding
14:49from the ground up how AI agents work,
14:51how to select between
14:52models and providers
14:54and agent frameworks, et cetera,
14:55how to approach building your agents,
14:57including knowledge, infrastructure
14:59and context engineering.
15:00We do two calls a week.
15:01One is a workshop, one is a
15:02questions and answers session.
15:04Come say hi, being a
15:05community of builders
15:06that use AI that love AI like you and me,
15:09and we share our
15:09solutions to these problems
15:11and challenges we face,
15:12such as replacing cloud code.
15:14So if that sounds interesting to you,
15:15check out the link below
15:16and I hope to see you there.
15:17Thanks for watching and I
15:18hope this has inspired you
15:19to reconsider your own
15:21Claude or OpenAI subscription.
25:00[Music]
27:00[Music]