Input: $30 / 1M tokens
Output: $60 / 1M tokens
GPT-5.5:
Input: $5 / 1M tokens
Output: $30 / 1M tokens
Costs have been reducing by over 5x year over year. Inference cost concern is mostly performative.
https://simianwords.bearblog.dev/conclusive-proofs-that-llm-...
Edit: can't reply but companies aren't selling inference at loss. In the blog post I point to third party hosting of open models like Deepseek which are also going down. They are not VC backed.
I also point to Gemma 31B which you can run on your laptop today that beats most models from 2024.
We will only know the actually situation once Anthropic goes public and we can look at their books.
That blog post is not very compelling either. Without knowing details of the architecture, comparing the various frontier models to open models doesn’t make sense.
Why do you need to know the architecture? Just compare Deepseek V4's performance with GPT 4 and treat internals as a blackbox. Deepseek is much cheaper and way more performant. If you can agree to reasonable assumptions
1. that closed source models are more efficient than open source
2. Deepseek is served at a profit and not a loss
Then it is pretty clear that the prices have gone down. If the prices have gone down more than 20x-30x then surely it is not _still_ subsidised is it?
I think this amount of skepticism is not warranted here. Every reasonable explanation or proxy is met with "but you don't know what they really do" is naive.
It is borderline conspiratorial to believe it this way.
Not a reasonable assumption for a variety of reasons.
> 2. Deepseek is served at a profit and not a loss
Not a reasonable assumption either.
> Why do you need to know the architecture? Just compare Deepseek V4's performance with GPT 4 and treat internals as a blackbox.
Because the internals are what actually matter and what drives inference cost.
It would be entirely reasonable to expect that GPT-5.5 has some sort of optimizations or changes to the architecture to make it easier to train, or to make runtime ablation easier, or to better handle large batches, or whatever.
Those changes, particularly if they are non-public, can easily result in worse inference performance than a comparably sized model without those changes.
> It is borderline conspiratorial to believe it this way.
It's not any sort of conspiracy. It's how land-grab tech companies have always worked. To presume otherwise is silly.
Pricing has no correlation with profit. It can be artificially lowered to kill competition, and artificially inflated to maximize profit.
GPT-4.1 Input: $2.00 / 1M Tokens Output: $8.00 / 1M Tokens
It would be more surprising if the surrounding architecture hasn't significantly diverged. If it _hasn't_ significantly diverged, then given the performance difference it would imply that the frontier models have significantly greater param counts, which would result in a higher cost.
We also have to assume that these operators are correctly pricing GPU depreciation, and the market is so new there is no reason to believe they are.
> Open source models are 3-6 months behind.
On the benchmarks included in their training set yes, not in real life
Assertion assertion assertion wishful thinking assertion.
Show, don't tell. Show us that we're wrong and this isn't a VC black hole. The CEO of Enron as late as September 2001 could've called every critic a sad dark loser with nobody challenging him publicly. Jim Cramer famously yelled anyone pulling their money from Bear Sterns in 2008 was "silly, do not be silly" exactly 8 days before their collapse and a -92% stock drop. In COVID, calling everyone paranoid and sensationalist about some mythical new flu was popular in December 2019 and gone by March 2020. How about Uber, the seeming go-to for how VCs can turn a money hole into a profitable business? The average price increase is now 18% per year and still going up, with an over 60% increase in 5 years. Does anyone still talk about the "sad dark HN loser path" of those who doubted VR in 2018? How's your VR startup doing?
Meanwhile there are layoffs everywhere, childcare costs keep rising, products shrinkflate.
> it costs OpenAI less money to serve GPT-5.5 than GPT-4
> Ppl don't understand how much efficiency gains are being made
I guess "ppl" also don't understand then, with all the supposed "efficiency gains" and "tokens getting cheaper" how come MS GH Copilot is switching everyone to token-based billing? Must be because those tokens are so damn cheap, innit?
Previously they used "premium requests" which would allow you to make a request to one of the more expensive models. People abused the shit out of this because a request was disconnected from tokens.
You could make one request which used tens of dollars worth of tokens, obviously not the intended usage pattern and obviously unsustainable.
Tokens for a given intelligence level are becoming much cheaper very quickly, but everyone wants to use the smartest frontier models so tokens are not dirt cheap. Even frontier models are a bit cheaper in absolute terms than they previously were, and much cheaper in terms of intelligence.
Having used it for > 4 years and having paid for it for > 2.5 years, I think I know full well how it's previous billing worked.
> You could make one request which used tens of dollars worth of tokens, obviously not the intended usage pattern and obviously unsustainable.
Gee, thanks Mr. Obvious! It never occurred to me this was the reason Microsoft recently removed Opus 4.6 and added a 15x multiplier in front of the inferior, but less token-intensive Opus 4.7!
Microsoft's previous model was not linked to tokens at all. Complete anomaly among coding agent providers. It's not representative of token economics at large. Claude Code recently announced increased limits. Codex does regular limit refreshes.
Tokens are pretty damn abundant even though they're not bargain basement cheap yet.