Anthropic Subscriptions Offer 5x+ More Value Than OpenAI
newsletter.semianalysis.com
newsletter.semianalysis.com
This is clear manipulation to prevent me from complaining about lower limits until it's out of the news cycle, but I don't care. As soon as those are gone, if my $200 account doesn't give me the value it used to, I'm out.
The random resets are also clear manipulation in this regard as well. Rather than just give me a fixed weekly limit which I could then get a sense of and know if it's getting reduced, they throw out a ~0-3 resets a week at random, unexpected intervals so that I'll never know.
Do you have beefy hardware? i dream of using a local LLM, but neither my 24GB RAM M4 Air nor my PC with a RTX3080 and 16GB of RAM seem usable (yet).
i could label them usable if "prompt it and then wait for 6 minutes for it to change a single line" is deemed usable, which i don't.
spent the last weekend in a rabbit-hole of which models to use in what kind of setup, but i think with my hardware i'm just out of luck for now.
for private matters I can use Cloud AIs, but not allowed to use it for work-related matters, which is where i could use it the most. they do provide us an isolated AI environment where we can use Claude etc. for work stuff, but heavily rated (about 20 prompts per week per Claude model).
You could run some decent models on your PC l, but nothing that would totally replace serious work. Qwen 3.7 35b is pretty okay. Maybe you could try the new bonsai quant for 3.8 27b but I doubt it would be awesome.
maybe in 1-2 years my hardware will be enough for fast local-AI. maybe by then i can afford new hardware. (think with current trends option #2 gets less realistic each quarter).
not sure if i lost a bit of curiosity and spark. there's so many models, tools, harnesses, tweaks for harnesses, edited models and so many things. ultimately i don't care enough, it seems vapid to stay bleeding-edge informed about LLMs, if one's career does not directly depend on it. i mean, what the fuck even is a bonsai quant, this stuff makes me feel like an uninformed excel boomer even tho i absolutely don't am one :D
not a software engineer per se so i don't need 16 agents running 24/7 with openclaw. i want a local buddy who helps me write ansible/python faster, better, helps me analyze bugs, refactors small things. basically what Claude Web does, with the benefit of it reading the files itself and running locally.
maybe i just ignore the whole AI/LLM part and just buy some new hardware so i can use higher graphics settings while gaming, WAIT NO, that market was nuked by AI as well, guess i'll continue using my i5 from 2019.
Local models and harnesses are the new ~2010s js front end frameworks.
Either accept there's always going to be a new thing of the week and don't over-invest in expecting reliability, or work in another space.
I know it's anathema to "normal" reliable software dev, but that's what new industries look like...
Last week I've have about 4 resets in 48 hours:
- reset done by OpenAI around Friday
- banked reset expiring on Saturday
- banked reset expiring on Sunday
- regular weekly reset expiring on Sunday
So basically I was unable to use them, given weekend and all at once, unless you have the software factory ready to spin up...
I'm not complaining, they are free after all, but it's clear they are not randomly distributed.
As soon as I received them and realized they expire this year, I switched to a $20 membership (billing cycle starts tomorrow), assuming you don't.
After consuming them, I will decide to whom subscribe next.
At this point we conclusively know that.
What frustrates me is the constant unshakeable belief here that subscriptions must be running at a loss, based on zero evidence.
Anthropic have said they have 85% margin, which presumably means their normal API pricing. 85% margin is insane. No business deserves that. We can easily assume the subscription is running at a more sane margin like 5-15%.
I don't know about everyone else, but there is no way in hell I would pay even $200 for API-priced tokens for personal use. So at least for my sample size of one, Anthropic's revenue would not be higher if they dropped their subscription plan (they would get $0 from me instead of $200 per month).
I respect Anthropic (and OpenAI to a lesser extent) but I’m not going to play these games.
And a local AI solution is less capable and also quite expensive, but for some worth trying.
For 2-4k you can have opus 4.8 at home running faster than anthropic. In 8-20 months you have broke more than even.
I think for almost anyone it's worth trying. Especially if you already have hardware.
So trade-offs are there, but do those matter to all people the same way ( are they not equal in the same way to everyone )?
Because there are a large number of use cases (and large sub-portions of others) where genius-level AI isn't required.
Which means once that threshold is surpassed in people's relevant fields, available margin on that is going to collapse to commodity levels.
It's difficult to see how either pure-play AI company maintains its valuation once that happens. Their TAM is based on capturing a big chunk of all work, not just that which requires the highest intelligence.
All this indicates to me that loads of GPUs were only purchased on paper or are sitting in warehouses unused. This almost certainly means a drop in orders followed by a drop in RAM prices (though that would indicate a demand drop to investors, so maybe it’s better for stock prices to keep paying to bills bigger warehouses and stuff them full of unused GPUs).
The second RAM prices normalize, local LLM becomes much cheaper. A machine that doesn’t make much sense at $10-15k suddenly becomes a lot more feasible at $3-5k. If those stored GPUs flood the market, we might see even bigger price drops.
First, the form factor of GPUs in data centers aren't the same as desktop GPUs, so you couldn't use them even if you wanted to in a normal rig.
Second, from a business perspective that doesn't make a lot of sense. It's much more likely that GPUs are going to the highest bidder/large contracts who are scooping them up to populate/upgrade data centers that are in operation because they are going to get more money per gpu on newer hardware. The "old" hardware might go to a warehouse to be auctioned to the highest bidder or go to a data center coming online, but I highly doubt they are sitting on market wrecking amounts of GPUs just waiting to flood the market. When they go bankrupt and they have to sell a datacenter or two, those datacenters being parted out as part of a bankruptcy deal I could see. Warehouses full of unused GPUs doesnt make any sense to me though.
With eGPU PCIe 3.0 or 4.0 x4 links (or even Thunderbolt), that essentially doesn't matter for inference.
You pay the bandwidth hit on model load / unload (a few seconds), but it's irrelevant for post-loaded inference.
Anyone hosting a high wattage GPU is going to be fine hosting in an external enclosure, most of which have generous extra-spec room.
And if they don't, if a flood of cheaper DC GPUs hit the used market, you can bet Chinese manufacturers will have enclosures that fit them available the day after.
Yeah but it's unlikely to happen in coming 2-3 years (at least) and then the SOTA models are going to be a completely different animals.
The flat monthly subscriptions seem too tempting for the providers to screw with. I prefer to paygo and to be responsible with my consumption. I also want the ability to scale substantially beyond what a typical consumer plan may offer on occasion.
My monthly usage ranges from $10-$1000. I don't have to worry about quotas or anything. If I need to use several thousand dollars worth of tokens, I can just pull out the Amex and everything works. It's constant performance all day every day. I have long since maxed out my org level with OAI, so it would be very difficult to exceed any realistic limits.
Must be nice to have money to burn. But if I had to guess, this doesn't seem like representative consumer behavior that Anthropic or OpenAI should rely on holding up at scale.
I lost my job recently, so gave myself an AI budget of $100 to help with search and applications.
If I put that in OpenAI or Anthropic, I’d hit limits quickly and lose whatever I didn’t use.
Or… I could put the same money in OpenRouter, use open models at 1/25th the price, and only need to pay more when I’ve spent what I put in.
Back in January when Claude Cowork was new and Claude Code was one of the only performant harnesses, it would have been a tougher decision.
But Hermes, dsh, agy, codex, opencode, pi… they make it so easy to achieve so much with such a low budget.
I know it’s a cliche nowadays to say “just use cheaper models” but the value they offer is SO much greater. And i can switch to GPT6, or Opus 5.5 in two clicks for tasks anyway.
My point is this: OpenAI and Anthropic are pulling stunts like this because they don’t care about you and your subscription. So stop caring about them.
Exactly where I landed, I had a personal subscription last year that shifted between OpenAI and Anthropic depending on who had the better model for that month.
This year? I don't need that, I can do the same as you: budget and pre-pay for some tokens in OpenRouter, and use very cheap models for absolutely anything I need on my personal projects. I can use open source harnesses that give me similar results, my projects do not need the absolute most-expensive frontier model at all, that's just a waste.
And if absolutely needed to use some frontier capability I can pay the tokens for that instead of committing to US$ 200-500 for a subscription that they can just pull the rug from me at any point.
I still have my job where they give me access to all the shiny expensive models with their enterprise agreements about data retention, the legal stuff that a company cares about and my personal projects don't, if I keep my usage under the newly implement monthly budget no one will bother me about it and so I just use what I'm told to.
But you give up privacy, because now your data is in the hands of 2 service providers, not 1.
Hmm,yeah but then come the data brokers who conveniently tie everything back together ...
And there’s also a major error in the calculation: it doesn’t account for the “generous” resets that are mentioned at the beginning of the article!
Taking the 1 to 2 weekly resets into account and comparing tokens, the ChatGPT Pro 200 plan works out to be between 1.2× and 1.7× more cost-effective per token than Claude’s for Astra/Fable, while for Sol/Opus it’s more like 0.7× to 1.04× as valuable.
However, this also doesn’t take into account the fact that, as another commenter pointed out, GPT-6.1 Sol currently uses fewer tokens to perform the same task (according to that comment, 5× fewer tokens, but I haven’t verified those numbers).
My figures are very rough, but the conclusion in the title is clearly false and clickbait. I get the impression that OpenAI offers better value for now, even if that means we’re relying on those “generous” resets continuing to happen.
We should take advantage of it while it lasts.
The only way we get to predictability is by paying what subscriptions are actually worth, and this entire game hinges on the fast that AI companies are not convinced that most people are willing to pay the 3-5x (or however much it is) multiple on what these accounts actually cost to run.
Evidence?
It benefits them to have sensational headlines that people might pay to have earlier/better/more of.
> Token efficiency is also an extremely relevant factor, but the industry unfortunately lacks reliable data here. Many people like to cite this chart from Artificial Analysis, but we do not believe the benchmark tasks in the AA Intelligence Index are at all representative of real work people do with LLMs.
Right, so they just ignore the fact that different models use very different number of tokens to achieve the same thing. Ignoring that means they can't say anything about "value". This is like comparing numbers with different units. They're Atokens and Otokens.
day2day performance difference is negligible, there's less refusals and you always have glm 5.3 for whenever opus 5.5 arbitrarily decides to terminate your conversation.
This is enough to run 2 concurrent sessions 16 hours a day.
What happened with your subscription when they banned you? Do you lose what's left of the month without a refund?
As a Claude subscriber, I am quite enjoying the value I get from Anthropic. But I am not forgetting that this is their loss leader, and API access is still their main driver. So strike while things are good, but expect some pullback in the future.
There lack of accounting for token efficiency and tokens/task. Makes this whole thing are less useful.
Token doesn't have the same meaning at OpenAI vs Anthropics, they use different tokenizer. How can this be used for comparison?
Can't find the link now it was a the pi harness author saying you could load the entire harness into context with those budgets IIRC.
Anthropic has way more inflated api token priced than OpenAI what is plainly visible on cost charts.