Case in point, claude code seems hell bent on increasing usage at all cost. Which makes sense in the growing phase (get people hooked) but it does not make sense given the hardware shortage. So, which is it?
Case in point, claude code seems hell bent on increasing usage at all cost. Which makes sense in the growing phase (get people hooked) but it does not make sense given the hardware shortage. So, which is it?
The most widely accepted estimates (though I disagree with them) are that Anthropic’s margin on API inference is ~70% from which people extrapolate what their token usage would cost via the API and compare that to what their plan costs.
https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...
(edit: better link https://newsletter.semianalysis.com/p/anthropic-growth-and-b...)
re: increasing usage with resets, it’s because they’ve overblown usage and need to show that usage is growing ahead of the IPO. They’re increasing usage on fixed price plans without increasing the cost, the only plausible explanation is they have unused capacity. If they were capacity constrained then the last thing they would do is give away more usage for free.
A lot of people are using Claude and ChatGPT for all kinds of minor things at work, and they probably wouldn't be if they were paying the true cost of the product. And this is all the while their work product is suffering because AI is not a great fit for a lot of use cases.
No they did not. Dario has repeatedly stated if they stopped training new models, they would be very profitable.
“The majority of the cost is inference, not necessarily the training of the model.”
It's not an assumption when the companies hosting the models publicly announce that they're subsidizing costs.
Ignoring current state of shortages, the interesting part (for me at least) is how much a SOTA token costs on current hardware if you exclude the costs for paying for current TSMC shortage and without factoring in the cost of new datacenters being rushed and renting hardware from your competitors (who are also supply limited).
The "actual" cost is of course also relevant and interesting but it is another thing entirely.
The former does not pass the smell test when there are random providers selling tokens for competitive open weight models.
But it does not tell you what and how much either is subsidized...