> For the data center build outs, demand for tokens is still exceeding supply.
Can you provide any numbers for this?
> For the data center build outs, demand for tokens is still exceeding supply.
Can you provide any numbers for this?
Now we don't know the true size of any of the proprietary models, but my educated guess is that Sonnet is in about the same parameter range, just with better training and much better fine tuning and RLHF. Yet API pricing for Sonnet is $3/MTok input + $15/MTok output, exactly six times as expensive. Even Haiku is twice as expensive as Kimi K2.5.
I find it difficult to believe in a world where those API prices aren't profitable. For subscription pricing it's harder to tell. We hear about those that get insane value out of their subscription, but there has to be a large mass who never reaches their limits. With company-wide rollouts there might even be a lot of subscription users who consume virtually no tokens at all.
This is false. We may assume it's the most efficient way of generating revenue given their GPUs, but their overall profitability will just be a guess. They would still have incentives to run hardware at maximum, even when it's uncertain to eventually recoup costs.
> a world where those API prices aren't profitable
A lab with employees and models in training has other costs than the operating expenses of a GPU farm.
If they're losing money and have no VC backing, they'd just turn off the lights.
But that's moving the goalposts? The original claim was on inference itself, not the whole company.
> The cost to serve tokens is absolutely profitable today and that’s been true for at least a year.
Everything is profitable if you ignore the costs.
Are you sure? Surely there is a lot of interesting data in those LLM interactions.
It's fair to say that if all these operators are competing for tokens, that the OpenRouter token operator (not sure the exact phrase but the people running the models) are accounting for some level of margin.
However, how many of these are running their own data centers and GPUs?
If they are running their own infrastructure, then it's not a simple equation of if each specific token set is profitable, since it needs to account for the cost of running the data center. It could be that they believe that it is profitable in the long term by utilizing the long tail of asset depreciation, but that isn't guaranteed.
IF they aren't running their own infrastructure, then it's much easier to claim that it's profitable and has a margin (outside of running their servers to manage the rented infrastructure).
HOWEVER, a lot of data centers have some pretty crazy low prices for GPUs that may be vying for user base and revenue over profitability. In these cases, if data center growth starts slowing due to slower buildout then it's very likely GPU prices go up and inference stops becoming profitable for the open router owners.
So long term it's not clear how profitable even these open models are.
OpenAI and Anthropic definitely fall into the latter category too. Their infrastructure requirements are much higher than the open models, and they are being given huge discounts so Microsoft/Amazon/Google can all claim revenue (since they have profitability coming from other parts). It's not clear if OpenAI and Anthropic models would be profitable at inference if they were paying rates that cloud hosts would make a profit from.
There's just way too many dimensions to this scenario to flat out state that open router proves inference is profitable at scale.
That gives you a very good estimate of "how much can you serve the tokens of a model of the size N for while making a profit".
Now, keep in mind: Kimi K2.5 is 1T MoE. Today's frontier LLMs are in the 1T to 5T range, also MoE. Make an estimate. Compare that estimate with the actual frontier lab prices.
In the current volatile environment, the API prices are more of a baseline where we can assume it can't be much cheaper to operate these models.
This is why switching to local open weight models saves a lot of money. (Even though it’s not apples to apples.)
This is the opposite of an AI bubble burst.
300k tokens for that hour.
OpenAI charges $6.
Those are pessimistic assumptions.
If not, you aren't breaking even.
(1) Rent your GPUs.
(2) Pay list price, no volume breaks.
(3) Get only 85 tokens/sec. Realistically, frontier models would attain 200+ tokens/second amortized.
Inference is extremely profitable at scale.
You're generating about 36 million tokens/hour. Cost of Mixtral 8x7b on Open router is $0.54/M input tokens. $0.54/M output tokens.
You're looking at potentially $38.88/hour return on that H100 GPU. This is probably the best case scenario.
In reality, inference providers will use multiple GPUs together to run bigger, smarter models for a higher price.
“Technically correct. The best kind of correct”. So inference may technically be _capable_ of being profitable, but I have question’s about them being profitable in _practice_.
For supply look at outages and growth rates at companies like openrouter. The demand is growing every week.