Banning third-party tools has nothing to do with rate limits. They’re trying to position themselves as the Apple of AI companies -a walled garden. They may soon discover that screwing developers is not a good strategy.
They are not 10× better than Codex; on the contrary, in my opinion Codex produces much better code. Even Kimi K2.5 is a very capable model I find on par with Sonnet at least, very close to Opus. Forcing people to use ONLY a broken Claude Code UX with a subscription only ensures they loose advantage they had.
Google AI Pro is like $15/month for practically unlimited Pro requests, each of which take million tokens of context (and then also perform thinking, free Google search for grounding, inline image generation if needed). This includes Gemini CLI, Gemini Code Assist (VS Code), the main chatbot, and a bunch of other vibe-coding projects which have their own rate limits or no rate limits at all.
It's crazy to think this is sustainable. It'll be like Xbox Game Pass - start at £5/month to hook people in and before you know it it's £20/month and has nowhere near as many games.
Google has made custom AI chips for 11 years — since 2015 — and inference costs them 2-5x less than it does for every other competitor.
The landmark paper that invented the techniques behind ChatGPT, Claude and modern AI was also published by Google scientists 9 years ago.
That’s probably how they can afford it.
Google already has a huge competitive advantage because they have more data than anyone else, bundle Gemini in each android to siphon even more data, and the android platform. The TPUs truly make me believe there actually could be a sort of monopoly on LLMs in the end, even though there are so many good models with open weights, so little (technical) reasons to create software that only integrates with Gemini, etc.
Google will have a lion‘s share of inferring I believe. OpenAI and Claude will have a very hard time fighting this.
You've described every R&D company ever.
"Synthesizing drugs is cheap - just a few dollars per million pills. They're trying to bundle pharmaceutical research costs... etc."
There's plenty of legit criticisms of this business model and Anthropic, but pointing out that R&D companies sink money into research and then charge more than the marginal cost for the final product, isn't one of them.
My point was simpler: they’re almost certainly not losing money on subscriptions because of inference. Inference is relatively cheap. And of course the big cost is training and ongoing R&D.
The real issue is the market they’re in. They’re competing with companies like Kimi and DeepSeek that also spend heavily on R&D but release strong models openly. That means anyone can run inference and customers can use it without paying for bundled research costs.
Training frontier models takes months, costs billions, and the model is outdated in six months. I just don’t see how a closed, subscription-only model reliably covers that in the long run, especially if you’re tightening ecosystem access at the same time.
They can totally lose money on subscriptions despite the costs of inference, because research costs have to be counted too.
Of course they are losing money when you factor in R&D. Everybody knows that. That is not what people mean when they say that they "lose money" on subscriptions.
I don't really think that view is as widespread as you believe.
But this is how every subscription works. Most people lose money on their gym subscription, but the convenience takes us.
5h allowance is somewhere between 50M-100M tokens from what I can tell.
On 200$ claude code plan you should be burning hundreds of millions of token per day to make anthropic hurt.
IMHO subscription plans are totally banking on many users underusing them. Also LLM providers dont like to say exact numbers (how much you get , etc)
For small personal projects it’s great value for money. Cheapest subscription was like 3$ during new years, token quota is acceptable to me (my guess it’s about 50-100M tokens per 5h)
Dunno how it would be with big projects, but with “personal project” things it feels to me that GLM-4.7 is 80-90% of Claude Opus 4.5. Just a tiny bit of more hand holding for GLM.
You can buy a GPU that's been used to mine bitcoin for 5 years with zero downtime, and as long as it's been properly taken care of (or better, undervolted), that GPU functions the exact same as a 5 year old GPU in your PC. Probably even better.
GPUs are rated to do 100%, all the time. That's the point. Otherwise it'd be 115%.
You don't run your gaming PC 24/7.
The only reason they're "perishable" is because of the GPU arms race, where renewing them every 5 years is likely to be worth the investment for the gains you make in power efficiency.
Do you think Google has a pile of millions of older TPUs they threw out because they all failed, when chips are basically impossible to recycle ? No, they keep using them, they're serving your nanobanana prompts.
Why do people keep saying inference is cheap if they're losing so much money from it?
And cost of inference tripled from $3B in 2024 to $10B in 2025, so cost of revenue linearly grows with number of users, i.e. it does not get cheaper.
The interesting question is: In what scenario do you see any of the players as being able to stop spending ungodly amounts for R&D and hardware without losing out to the competitors?
It was sticker price of $33,000 adjusted for inflation:
https://en.wikipedia.org/wiki/Ford_Taurus_%28second_generati...
I don't think it would even feel safe to drive at all compared to what we have got use to with modern cars. It broke down 3 times while I had it and stranded me on the road. No cell phone of course to call anyone.
These were the mythic "good ol days".
I recently encountered this randomly -- knives are apparently one of the few products that nearly every household has needed since antiquity, and they have changed fairly little since the bronze age, so they are used by economists as a benchmark that can span centuries.
Source: it was an aside in a random economics conversation with charGPT (grain of salt?).
There is no practical upshot here, but I thought it was cool.
It’s also false that the technology has changed very little.
The jumps from bronze to iron to steel to modern steel and sometimes to stainless steel all result in vastly different products. Not to mention the advances in composite materials for handles.
Then you need to look at substitute goods and the what people actually used knives for.
A huge amount of the demand for knives evaporated thanks to societal changes and substitute goods like forks. A few hundred years ago the average person had a knife that was their primary eating utensil, a survival tool, and a self defense weapon. Knives like that exist today but they’re not something every household has or needs.
This is a good example of why learning from ChatGPT is dangerous. This is a story that sounds very plausible at first glance, but doesn’t make sense once you dig in.
With that said, if it is a hallucination (and it sounds like it was), it's one of the more interesting ones I have encountered. It almost has the shape of a good idea.
Blade and handle material has certainly changed over the years, but I think good arguments about how relevant that is could be made both ways. They remain handled cutting tools, used in the same general way, for the same general purposes (though as you posted out, some use cases have gone away). Basically anyone from any of these periods would recognize a knife from any other, and be able to pick it up and make immediate use of it for all their normal knife related purposes.
To be clear though, I am now siding with the clankers and arguing for a hallucination. It's an interesting thing to think about, but it sounds like it's not an established concept in any way shape or form.
Enterprise products with sufficient market share and "stickiness", will not.
For historical precedent, see the commercial practices of Oracle, Microsoft, Vmware, Salesforce, at the height of their power.
The software is free (citation: Cuda, nvcc, llvm, olama/llama cpp, linux, etc)
The hardware is *not* getting cheaper (unless we're talking a 5+ year time) as most manufacturers are signaling the current shortages will continue ~24 months.
Yes, that's the time I'm talking about.
You also had a blip with increasing hard disk prices when Thailand flooded a few years ago.
If you factor in the cost of integration and ongoing maintenance - by humans or llms - it is not free. But it certainly has never been cheaper.
Despite the high price, the Bentley factory is running 24/7 and still behind schedule due to orders placed by the rental-car company, who has nearly-infinite money.
We see vendors reducing memory in new smart phones in 2026 vs 2025 for example.
At least for the moment falling consumer tech hardware prices are over.
I also think we're, as ICs, being given Bentleys meanwhile they're trying to invent Waymos to put us all out of work.
Humans are the cost center in their world model.
If AI was truly this productive they wouldn't be struggling so hard to sell their wares.
Finance 101 tldr explanation: The contribution margin (= price per token -variable cost per token ) this is positive
Profit (= contribution margin x cuantity- fix cost)
i will not be as bullish to say they will no colapse (0 idear how much real debt and commitments they have, if after the bubble pop spending fall shraply, or a new deepseek moment) but this sound like good trajectory (all things considered) i heavily doubt the 380 billions in valuation
"this is how much is spendeed in developers between $659 billion and $737 billion. The United States is the largest driver of this spending, accounting for more than half of the global total ($368.5 billion in 2024)" so is like saying that a 2% of all salaries of developers in the world will be absorbed as profit whit the current 33.3 ratio, quite high giving the amount of risk of the company.
Is my goto reference for debt numbers etc.
If the answer is not yes, then they are making money on inference. If the answer is no, the market is going to have a bad time.
The sounds like a confession that claude code is somewhat wasteful at token use.
I find that competitive edge unlikely to last meaningfully in the long term, but this is still a contrarian view.
More recently, people have started to wise up to the view that the value is in the application layer
https://www.iconiqcapital.com/growth/reports/2026-state-of-a...