Why current LLM costs are not sustainable
aditya.patadia.org
aditya.patadia.org
1. We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs.
2. There hasn't been a real incentive to work on cost optimization for data centers and the hardware they contain. When/if price hikes happen and send people scrambling to use other models or drastically reduce AI usage, this will suddenly need to happen.
3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that.
4. Open models are a wildcard that could completely change the calculus.
This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow.
My claude code environment shows me cost per token used in that session, according to API costs. It regularly exceeds $200. I pay $200 a month for my claude subscription. That's fairly obviously subsidised, unless you genuinely believe their unit costs are 100x less than what they're charging.
For example, Microsoft 365 Copilot. A company might get a significant discount on the price per user for that SKU. What does that translate to? OpenAI is the actual brain behind it. So who ends up getting the money at the end?
https://blogs.microsoft.com/blog/2026/04/27/the-next-phase-o...
OpenAI needs Microsoft's infra for training and market penetration. MS needs something AI to slap onto their products until they develop their own in-house.
Now, whether it's actually necessary or not for an enterprise is a question. There's a lot of FOMO spending. Maybe 5-10% of your workforce are programmers that need great AI. The rest of the admin, finance, other people? TBD.
Feels like arguing that it's not clear if Bugatti's losses came from selling the Veyron instead of designing and developing the Veyron.
I missed the word don't, originally. Apologies.
> What is happening here is that leading AI labs are charging not only for inference but also for research in model architecture, training data collection and curation, model training cost (which can be tens or even hundreds of millions of dollars), paying their employees and recovering the marketing costs.
That's what's being subsidized.
As said before: Interference costs are not the only operational costs. Same as electricity costs for the craftsman. Running a power drill is not the the whole expense to consider. The craftsman has to eat, AI company's employees have to eat. The craftsman has to learn about new building standards, the AI company has to train their models because no one wants to use a product stuck in time (that's not "research", just maintenance). If not even interference was recovered in revenue, nobody would even start to argue about sustainability.
I can't debate this further, because HN is rate limiting my account for dissenting opinions in the past.
How does that figure look if you count in the current unprecedented LLM/AI-driven price inflation on both hardware, services and software? I don't believe we're exactly in the "$5 airport uber" era if you count that into your total.
And in this analogy you need to spend a lot more when buying a car, no matter if it's a new or 2nd hand one, following the price inflation caused by cheap Ubers. So in essence, my question is how much have those cheap Uber rides then cost you in reality, when factoring in the directly related price increases for the things you need and buy? Is it a net positive or negative at the end of the day for anyone other than the very few at the very top of the system?
No.
Right now it's silly to default to frontier models, but it won't bankrupt your company. I believe in the short-medium term future, we'll need to be more deliberate about model choices.
In the long-term, of course, tech costs tend to plummet. Is there a future where in 15 years, my Apple Watch locally runs an Opus 4.8-class model? Maybe. And that would obviate this whole discussion.
We've already seen price hikes / token limits earlier this year, with suddenly some people running out of budget on the first day of the month. This will likely keep going for a while.
On the other hand, costs will drop too - open models and specialized hardware, as the article notes. The long question will be whether the companies will get a return on their invested billions. I don't think they will, not with the amount of competition they're facing, and I don't think any one company or model (series) has a monopoly yet. Popularity sure, but I'm confident a competitor may appear tomorrow and people will switch.
I think it's like that, but not quite. The people who have a subscription but barely use it were probably never doing any serious work with AI in the first place. I.e., why would they get a subscription when their one or two chat questions (or, "make a picture of me as a superhero" prompts) per day can be had for free?
Especially with Claude, I think people who subscribe skew very heavily towards people that can very easily make more than $20 worth of queries in a month. And then there's the not-insignificant number of people who are tokenmaxxing.
It's like the gym membership model except ten percent of members are able to spend 72 hours per day at the gym while the rest spend 8 IMO.
Likely even the E4B, which is really both fun and impressive.
That is clearly a big component of Apple's bet, anyway.
It's still more like programming than telling a chatbot to go make you GTAVI in JavaScript and make sure the graphics are as good as the original.
Maybe a safer prediction would be that most people will be fine just using hybrid agentic programs that run the models locally(probably with extra spyware). I think this is Apple's bet.
It's actually worse, the AI explosion just hikes hardware prices faster than capacity can catch up - and it will likely not catch up in a while because investments are both expensive, long, and might not seem all that good idea while bubble is still bubbling.
The massive push frankly also made it unsustainable. If RAM didn't cost 3x and compute manufacturers would have to compete instead of selling every unit instantly at whatever price they want the frontier model tokens might've costed closer to sustainable amount
Inference is not exactly cheap. Based on what do you think this is "heavily subsidized" still? What would to token cost have to be, with current models, for it to not be that? What do you know that has you make such a claim?
Do they? It's free right now at chat.com. After that it's $20/month which isn't much in the US. Three Starbucks or two meals at McDonald's will run you more than that these days.
This is nonsense that AI providers want to peddle. Inference is wildly gross margin profitable - likely 90%+ gross margins. It's very easy to work out the cost structures bottoms up. All providers can drop costs to a third and still keep positive gross margins.
The problems are 1. It possibly still doesn't pay out on training investment in a reasonable time frame without a massive expansion of the 90% gross margin.
2. There is no moat. As we see Mac Mini & High End GPUs stock outs and the pricing offered by DeepSeek and Qwen, the performance of Open Weight models are good enough that people can and are already shifting many inference workloads out of these 90% margin players
If you can use a subscription with any of the SOTA models, do that.
Instead of around 4k EUR in token costs, my Opus usage costs me 108 EUR (with taxes) per month with their Max 5x plan. It's the same with OpenAI, those are heavily subsidized.
It doesn't make sense to pay per-token, unless you must.
> What is happening here is that leading AI labs are charging not only for inference but also for research in model architecture, training data collection and curation, model training cost (which can be tens or even hundreds of millions of dollars), paying their employees and recovering the marketing costs.
Chances are, they're never getting that money back. Best case scenario, the hype around AI slowly declines, worst case - it crashes and takes a part of the economy with it.
Also anyone doing distillation with hundreds or thousands of those subsidized attacks is probably winning big. Especially as the model architectures (e.g. DeepSeek V4) are more oriented towards efficiency.
> Last but not least and in fact the most important factor, is the ability of users to run local models. So far, almost everyone is using cloud-hosted models and local models are either too big to deploy or too slow to work with. With advancements in chips, this will change in 4-5 years’ time.
Currently beefy hardware to run them fast enough to be competitive with the cloud (at least 60 tps) is expensive and even then the small local models quite suck compared to SOTA or even DeepSeek V4 Pro and GLM 5.2, though they're way better than they used to be (compare Qwen 3.6 with 2.5 for example).
Those subscriptions plans are for private use only! If you are running a business you are not allowed to use them actually. Anyway..
This is at work where I don't work on greenfield or parallelize feature development.
I cannot see the agent burning through $50 for one moderately sized TypeScript cleanup in my setup. This sounds like something that can be improved on OP's side.
There have been rumors about a potential Sonnet 5 model release in the near future, which hopefully tilts the cost/benefit ratio further in our favor.
I have absolutely seen stuff like this happen. Think about it, when you point Claude at a bunch of files, it has to suck them all up (tens of thousands of tokens), spend some proportional number of tokens doing stuff, and spit them back out (tens of thousands of tokens) for each pass in the "cleanup" loop. I had a similar situation occur a few months ago. Very small "add Javadoc to these dozens of classes" scenario. Sonnet rapidly rate limited my $20 plan so I switched to extra usage. A very small (IMO) number of changes later I had spent like $7 in tokens.
The main problem is you really have no idea ahead of time just how many tokens a given task is going to take. I suggest you try spending a day running your Opus 4.8 High effort on API pricing to see just how much your $200 subscription is being subsidized before you confidently state that $50 for some TS cleanup task isn't possible.
$50 would be 10M input tokens, not tens of thousands.
Well I saw $200/month and thought you were talking about a max plan, sorry. But I will say unless you're using that top end model extremely judiciously $200 for 2-4 weeks of work is similarly hard to believe (see the other poster breaking down their usage). What are you typically doing? Must be pretty hardcore stuff if you need to use the baddest available model. How many interactions per day? Care to share your token usage stats?
> $50 would be 10M input tokens, not tens of thousands.
Two things. One, input tokens are but one component, and the cheapest. Output tokens include the tens of thousands being spit out for file changes AND the thinking/crunching that you don't see. And that's the most expensive part. And remember, that's per iteration, not everything is one-shot (especially with tasks like "fix this large part of my codebase).
It is not my experience that you need to do 'hardcore stuff' to require the use of a large model. The difference in productivity between babysitting Sonnet and trying to get the result into a good shape compared to using Opus 4.8 seems large to me.
At home, unfortunately I only have the stats from the official apps rather than granular ones, and it looks like the Claude Desktop app is buggy: it was showing 17M tokens total in the last 30 days, but even just clicking on a conversation in my side bar increased the counter to 19M. It's clearly not working.
Codex shows up to 900M tokens total/week.
Here's my usage, from the ccusage tool (slightly shortened for readability):
┌──────────┬───────────────┬────────────┬─────────────┬─────────────┬───────────────┬────────────────┬────────────────┬─────────────┐
│ Month │ Agent │ Models │ Input │ Output │ Cache Create │ Cache Read │ Total Tokens │ Cost (USD) │
├──────────┼───────────────┼────────────┼─────────────┼─────────────┼───────────────┼────────────────┼────────────────┼─────────────┤
│ 2026-06 │ - Claude │ - opus-4-8 │ 13,635,792 │ 32,562,574 │ 177,985,265 │ 5,265,814,971 │ 5,489,998,602 │ $4665.09 │
└──────────┴───────────────┴────────────┴─────────────┴─────────────┴───────────────┴────────────────┴────────────────┴─────────────┘
Now obviously that is all with the Max 5x subscription, other agents and models excluded.So per day that'd be around 155 USD (including weekends), which doesn't seem that far off, as long as the example cleanup takes up around 1/3 of one's daily work (or needs a lot of review/test iterations, or needs to review a lot of the existing code etc.).
I do not know whether that is typical, or indicative of conversations with too many turns.
Not that I would worry about this on a subscription plan, but at work where we are billed at API rates, I try to move to new conversations as often as possible.
For example, if you make Claude Code explore a codebase, write a plan based on it and your requirements, do a few iterations of further specifying and altering it, and afterwards let it work for let's say 2-4 hours.
Sub-agents and dynamic workflows do alter the numbers a bit, but not to a crazy degree in the long run.
So, given the SOTA providers with even larger models also need to continously be using considerable resources for training their next models, to fund future data centers, and make profit, the token costs are more likely reflecting the real costs, rather than the subscription costs.
So what's the price difference, 3000x?
One thing we do know from OpenAI's leaked financial document is that they are already profitable on inference, though that data is not broken down by cost and revenue of API vs. subscription. One important factor is that subscription inference can be optimized in ways to reduce cost (e.g., usage limits, batch optimization around API-prioritized inference, etc...). I think simply we do not know the actual cost of subscription interference for SOTA models.
Sources https://openai.com/business/pricing/#api says for GPT-5.5:
Input:$5.00 / 1M tokens Cached input:$0.50 / 1M tokens Output:$30.00 / 1M tokens
and for https://docs.fireworks.ai/serverless/pricing DeepSeek V4 Pro: Input: $1.74 / 1M tokens Cached input: $0.145 / 1M tokens Output: $3.48 /
Ratios are: 2.8, 3.4, 8.6So as these numbers seem reasonably comparable to SOTA, and the SOTA vendors have additional overhead, then I think it is fair to deem that the alternative explanation offered here is not the explanation:
> Why do you think that subscriptions are subsidized and not that enterprise tokens are sold at 3000% margin?
As it does seem like the GPT-5.5 API tokens do not have significant margin based on the overhead-free companies selling inference for smaller models at prices of the same scale, I think we can believe that the subscriptions must be heavily subsidized.
It should be noted though that DeepSeek itself sells this even cheaper, but they may also be in it for the getting market share.
The fact that they'll milk corpos that actually have money is obvious, compared to me because I'm broke, as are many other subscription users. The large AI labs don't seem to be profitable so I bet for the regular users there's plenty of subsidizing going on to at least get people to use the tech (and maybe that'd lead to some conversions at work or API usage eventually):
> OpenAI's net loss ballooned from $5 billion in 2024 to a staggering $39 billion last year, as it continued to spend heavily on AI model development and securing compute capacity, the Financial Times reported on Tuesday, citing audited financial figures confirmed by its sources.
https://finance.yahoo.com/markets/stocks/articles/openai-fin...
Inference itself might be profitable, but is just funding the training and other stuff.
So how these companies and people manage to use these absurd amount of tokens is a mystery to me. It feels like this are just running huge amount of non-vetted data to the LLM's and or running loops against the LLM's which only produce fractional results if not wasted results for insane cost.
So really it is the equivalent of just burning money, or heating your house in the winter while having all your windows open.
But try running Claude Opus at API prices through a 'clever' RAG based intermediate system 'managing' a 2024-era context size window completely unaligned with 2026 frontier model tool use expectations, that results in 100% cache miss and content coherency destruction on every single interaction. There's your typical 'Enterprise Agreement' GenAI setup.
I only really discovered this when trying to find out how my Enterprise friends' AI experiences were so completely opposite from my own successes as I could not believe how poor their results were even though on the surface it looked like we were using the same model, and I know they aren't 'bad' software engineers and developers.
Absolutely!
I know some colleagues who are routinely spending thousands of dollars worth of tokens, I can't see to even max out the subscription limits even if Im working all the time. Curiously enough their output is lower too.
Personally I find it faster to figure out the hard parts by myself and then give a few smaller tasks to Claude.
Fire and forget. They run multiple agents in parallel 24/7. AI isn't just a rubber ducky for them, its their main (only) tool at that point.
Yes that also means I get to inspect and confirm every step of the way, to ensure the design is followed, we are not making unneccesary changes, we have thought about edge cases, testing etc. And I also keep an understanding of what is produced, because I will manually copy it in, I will manually read through it. I will do secondary review of it myself in PR's whatever.
But I guess a lot of people just don't and just blow claude code through the roof on ad libitum infinity loop?
On subscription, I just checked, I have the Pro Plan, which for Claude I believe is the equal of the normal one?
But if you use Claude Code or Codex, you will blow through your pro plan quickly. If you don't use them, you are not really using AI. I know how that sounds, but that is how it is.
These models are smart now. Really smart. Yes, they hallucinate, but usually not without reason. I am having long discussions with these models before generating code, and generate markdown from them. These are then the basis for the generated code. I am trying to give the model as much background as possible. I read the generated markdown: if there is something that feels off, like I don't really know what it means, then you need to fix that first, by discussing it with the model. Often, these are real problems in how I was understanding something, the model wasn't really getting it, and just made something up that it hoped would kinda work.
And I prefer Codex over Claude Code (prior to Fable, Fable is something else!), it behaves more like a helpful PhD-level colleague and just feels sharper. Claude Code sounds a bit like a mix between an HR person and a therapist that is on vacation too often.
I am still looking at code, but only if something came up during high-level discussions with the model that I want to pin down exactly. Otherwise I just talk about the high-level intention of the code, usually not looking at it.
What REALLY helps is coming up with the right theoretical frameworks for your work, with practical implementations that the model can use, and that allow some kind of verification. Let's say you want to parse something. For a one-off the model is great at generating "hand-rolled" parsing code, but for something disciplined, giving the model a way to generate context-free grammars and giving it a way to check them for determinism gives great results.
It is not that hard. Just launch 10 different windows and make sure to loop back in after every turn and you will be burning billions of tokens per month in no time.
Are in you sending it to work in nested for loops? If yes, what sort of work would that be?
There's additional advantages that everything you query, all of your context cache and everything it outputs stays private and can't be arbitrarily turned off by external interference.
Personally I think it would be a fairly good bet that something with the 1TB of RAM needed to properly self-host GLM5.2 will still be a very usable piece of hardware in 4 to 5 years from now. There will be even larger, newer models available, sure. But there will also be better models that continue to fit in the same size.
So you could see small LLM co-operatives working out, yeah.
But my thinking is that this four-to-five-year scenario just won't come to fruition, because the whole concept of needing to run these massive, massive models will slightly more likely be rendered moot by smaller models with better reasoning capacity, and possibly even in that timescale by hardware innovations.
One of the biggest problems I have with the whole "we won't be profitable until 2030" model is that 2030 is almost exactly as far into the future as the launch of ChatGPT is in the past, and in that time, models far more capable than that first ChatGPT have been made available to freely download and run on desktop hardware that existed before it launched, and the entire non-model surrounding functionality of that original ChatGPT plus many more functions is now not much more than a routine weekend coding project.
I don't know why the market would entertain the idea that no upset like that is possible in the same period of time again.
On top of this, people are constantly coming up with better ways of running models on less special hardware and "good enough" models are now existing for most tasks.
So where does that leave the frontier labs? Drug discovery? Maybe some hard math problems? I mean it's not that big actually...
We're in a brief window where this is profitable, like batch computing was in the 70s. However, once your own device can do it, you're going to start migrating.
Only on a pay-per-token basis, I think. Unless it's a very tight-knit circle of folks. Fixed monthly subscription costs I doubt would work in that model. Because you'll get the inevitable: someone pegging the service 24/7 because it's "unlimited" while everyone else suffers.
You can presumably hard-limit LLMs the same way — total, burst quotas etc.
(Suddenly getting a very fun flashback to the environment in which someone first explained Markov chains to me — MediaMOO. A text-based chat environment with configurable limits on the number of CPU "ticks" you were allowed in order to do things)
we are there, Gemma4/Qwen3.6 are GPT-4 level models runnable on a fancy laptop.
but expectations shifted, nobody wants a GPT-4 level model anymore
If all of global spend on Anthropic/OpenAI/Gemini APIs just switches over to DeepSeek then easily we can decrease total AI spend by 10x
But billions? A bit exaggerated.
Its the ridiculous cached input token price of $0.00028/1 M
https://www.reddit.com/r/DeepSeek/comments/1twesxe/comment/o...
Count me in, I'll be testing that. Thank you!
That is possible inside the US. How do you do it all over the world? You have to convince every country in the world not use frontier models? Even worse how do you convince all the countries to not build their own models?
China or your local one?
The difference is DeepSeek and other Chinese models are open weights.
As an European, today I certainly classify USA as a "hostile superpower" because the actions from the last few years of both the US government and of certain big US companies have stolen a lot of money directly from my own pocket, by artificially limiting competition in several important markets, like smartphones, SSDs and memory modules, thus greatly raising the prices in comparison with what they would have been in normal market conditions (i.e. if the US government had behaved after the same rules that they had forced upon the other countries for decades, by various methods of propaganda, bribing and blackmailing).
China has (so far), never done that to me.
This is not some hypothetical hostility, but billions of humans from all over the world have been losing more and more money in recent years and in recent months, from their own pockets, much of which eventually reaches US companies, like Qualcomm and Micron (or South-Korean companies, who have also benefited from the US policies).
Of course, China is not trustworthy, but until now, unless you are a neighbor like Taiwan, their hostility is only hypothetical and in the future, not real and in the present, like for USA.
China's hostility is not hypothetical. The US's hostility is not universal. I think my point about ignorance or propaganda is proven by your statement.
Literal race on twitter posting to increase token throughput and drive down costs on these Chinese open source models
> To give an example, just doing Typescript type fixes with this model across 50 files cost me $54 this afternoon.
That's all because it ran through the most expensive frontier model for a mechanical task that a cheaper model could easily handle.
What hardly gets mentioned is that most people don't actually measure what each task costs them on a granular level. A lot of the waste comes from running everything through one expensive model without considering breaking the big task into smaller tasks farmed out to cheaper models.
Whether the labs' economics hold is above my paygrade. It costs what it costs. What I can control is my own usage.
Like everyone else leaning in heavily on AI usage, my tokens started running out mid week...sometimes within a couple of days. I had to do something about it or double my spend. So I started tracking my cost per task type a few months ago and it completely changed my workflow. The lowest hanging fruit was the mechanical stuff. Moving that to the cheapest models was a game changer, and much faster to boot. Mid tier models take the workhorse tasks. The frontier heavy hitters are now only used for judgment calls like reviewing and planning. Spend dropped dramatically.
Freeing up all those tokens made me even more ambitious to explore parallel ways of working, to get even more out of what I was already paying for.
I really believe that in the near-term future we will run our LLMs in hardware, not in software. Hardwire a capable model into a device the size of a graphics card, embed it into a laptop, and you have something that uses less power, does faster inference, doesn't require additional CPU or memory, doesn't cost a monthly fee, and will probably eventually be available for under a (few) hundred bucks.
Fable seemed very clearly a step up in my one afternoon of usage. I gave it several bugs that other models had failed at repeatedly (in a mess of a vibe coded side project) and it fixed them each in one prompt.
No, but progress not stopping doesn't mean it's not plateauing. I believe 'plateauing' is understood as the process of approaching a plateau, not being stuck on a plateau already. So, the question is about the rate of progress, not its existence.
HN commenters have been saying that LLMs plateaued ever since the first ChatGPT release.
6 months ago:
> LLMs are amazing, but they have reached a plateau.
https://news.ycombinator.com/item?id=46109534
1 year ago:
> generative AI has languished in the same place, even in my kindest estimations, for several months, though it's really been years.
https://news.ycombinator.com/item?id=43085885
2 years ago:
> 2024 has seen nothing substantially good and the only notesworthy thing is this article finally hitting into the public consciousness that we are past of the AI peak and beyond the plateau and freefalling has already begun.
https://news.ycombinator.com/item?id=42125888
---
Many more that I haven't time to look up.
I think the present just always feels slow.
> HN commenters have been saying that LLMs plateaued ever since the first ChatGPT release.
Your earliest example was 2 years after release, when LLMs were already widely used and there is literally a source to support the claim. Now, you need to show research efforts, time and resource investments, ... are producing proportional results, disproving diminishing returns. If there are diminishing returns, LLMs are plateauing.
Also, manual capability extensions, or use/edge case adaptations, which may improve subjective usability are not exactly advances in AI as technology. LLMs still hallucinate. LLms still fundamentally struggle with certain classes of problems (e.g. counting), but you increasingly need to come up with different problem dress-ups because of targeted interventions to manufacture hype and a limited supply of test cases. Can you make the case AI got actually more intelligent, fundamentally? That is, not an increase in case specific usability, but a decrease in fundamental limitations. And is this proportional to improvement efforts?
I think you're setting up expectations here that nothing could falsify.
Any advance could be dismissed as "case-specific", couldn't it?
What's something concrete that you would judge as a fundamental step forward?
My own claim would be that the increase in long-horizon tasks without heavy harnesses is a sign of a decrease in "fundamental limitations".
METR's time-horizon tests are going up, and that seems to match actual problems LLMs are solving in daily usage. Eg. Agent workloads. Or games: see the newest models being able to beat longer-term tasks like Pokemon from just a minimal vision-only harness. Models last year had cheating harnesses and yet still failed.
> I merely pointed out your argument wasn't a sound rebuttal. Criticizing subjective experiences to assess the situation is, but that cuts both ways.
My original "doesn't mean progress stopped" comment was replying to the "no improvement" guy, not the plateau guy.
Yes, I agree there's vibe-based judgement happening on both sides. I was quoting old HN not to make a "boy who cried wolf" argument, but to say that feeling a vibe of slow progress is not enough, since people repeatedly have expressed that vibe.
> Your earliest example was 2 years after release, when LLMs were already widely used
Fair. I was going backwards and thought a few would suffice. There was a lot of doubt earlier too if you search. From March 2023, four months after chatgpt "We are heading towards a limit and the AIs aren't getting much better." https://news.ycombinator.com/item?id=35332537
> Now, _you_ need to show research efforts, time and resource investments, ... are producing _proportional_ results, disproving diminishing returns. If there are diminishing returns, LLMs are plateauing.
Yes, it's a good point, though we don't have good numbers on investment so it's hard to answer in either direction. But the METR speedup in 2024-25 appears to be mainly from RL/algorithm gains not training compute, judging as an outsider.
"Is progress efficiency per dollar decreasing" is a different claim to the earlier "is progress slowing?", though.
> LLMs still hallucinate. LLms still fundamentally struggle with certain classes of problems (e.g. counting),
Measured hallucination rates have gone down, and don't seem to be an issue for achieving things, as per METR and other results. I don't see hallucination issues in practice as all real work I know about includes the ability to validate.
If by "counting" you mean the "counting r's" issue, that's a tokenisation artifact (the models are blind to individual characters): ask it to count in Python and it's fine (ie. reasoning is not the problem).
(the viral online counting videos of voice-mode are using an ancient gpt4o-based model that is trained to keep answers short. I do think this is giving much of the public the vibe that current text models are worse than they actually are)
and
> If by "counting" you mean the "counting r's" issue, that's a tokenisation artifact (the models are blind to individual characters): ask it to count in Python and it's fine (ie. reasoning is not the problem).
The problem is a lack of transparency. That's on the AI companies, not me. Everybody can attach a mathematics engine to an LLM and say "Look, it does math now!". The intelligence part hasn't changed though, the model didn't figure out mathematics, didn't get any closer to AGI.
I am also not saying, testing is impossible (making claims unfalsifiable), but you have to acknowledge it's inherently becoming harder to argue either way. Now, if I asked to count to 100 the LLMs deliver. I can even ask to count to 100 in base16. But it falls apart in less common bases like base33. The problem clearly isn't solved fundamentally. For counting in arbitrary bases you need a knowledge transfer and creative adjustment since you may exceed the symbol space. I think, this is also a good indicator because for the same reason there is no training material. Again, this can be addressed algorithmically, so this test is not futureproof.
Regarding character counting, the excuse about tokenisation is a bit questionable in this context, I think, because inherent limitations to the current approach is exactly the matter debated. Do we know, the extent of this particular problem space in the real world?
>What's something concrete that you would judge as a fundamental step forward?
I want reproducibility and transparency how a feature happened. If you threw a bazillion dollars onto Wolfram Alpha 15 years ago, you would also have a very useful tool for most questions. Spending a lot of money to implement edge case solution is not technological progress. That's always been possible. I mean, in the beginning there were thousands of humans curating chatgpt answers. That's not intelligence, just reckless spending, exploitation and borderline fraud. And it shows exactly how much "customer experience" can be delivered just by throwing money around.
> "Is progress efficiency per dollar decreasing" is a different claim to the earlier "is progress slowing?", though.
Well, I guess it is a matter of opinion. In my opinion, a "thermodynamic" perspective is useful. If you hire 100000 more workers to harvest a small field, you get it done quicker, but is it a breakthrough in farming? That is, trivially, if we spent more energy, resources, money we can expect to get more anything. However, those investments are not spent elsewhere, the progress of technological/scientific advancement is slowing down elsewhere. So, going by relative contributions to overall progress things are slowing down when progress is disproportionally low in an resource intense field. After all resources on Earth are limited and there is only so much sunlight refueling our energy supply. If the AI hype isn't panning out, the damage by hardware prices alone will be enormous. You need compute for more than AI and a lot of that is unaffordable right now. AI has to be 10X for every field, or it's gonna be a net global set back. Look at the volatility and social disruption in the US. It's a gamble, not covered by growth, but debt. Putting progression into relation to effort is required IMO.
> Measured hallucination rates have gone down, and don't seem to be an issue for achieving things, as per METR and other results. I don't see hallucination issues in practice as all real work I know about includes the ability to validate.
Are you in tech? Because I think there is a bias related to the above, since agentic AI can compile/run code to validate answers and most problems are documented ad nauseam. In my experience, if you search for very specific information in other sciences, you still frequently get confident non-sense. Outside of IT, you can't really have a "fact checking" co-routine, you need actual cognition. Of course, if you "chain AIs" things get a little better, but I claim again that's multiplying effort, not fundamentally making anything smarter. If we compare to biological intelligence, we absolutely do not see the same relationship. You may think this argument is moving the goal post, but if I ask you this: If we double the AI workload (eg. agentic AI) to get 10% better results, doesn't this sound like approaching a plateau to you? We are literally putting the fabric of civilization and humanity's future at risk for this (by social disruption, increasing energy/resource expenditure and not spending the money elsewhere in most critical times). Expectations must be high.
Anyway, thank you for expanding the argument, nice discussion.
In terms of running the model locally vs a service provider, that will be down to convenience more than anything else for the same reason why not everyone is hosting their own website at home on their own box.
If we continue this year with a2a, agentic layer and co, there is probably a huge bulk coming up with a lot more agents running a lot longer and talking to each other to solve issues which will increase token usage significanlty.
The price for tokens will become a proxy of consumed energy in my mind, i.e. tps will be something like kWh almost directly correlated in terms of cost.
This is obviously untrue, both with GPT-5.4, and Claude Fable as examples in the last 6 months.
Like i still used plan mode 6 month ago now I don't.
I would argue that with every model release we have a new learning phase.
The AI haters have been saying this for 2 years now.
>amusing side note: >Was in a meeting reviewing a potential new product, it was going well until they showed us that they had added AI to it (of course they have). It was pretty obviously just shoehorned in, and one part of that obviousness was that they had a column that showed how many tokens it took to make each query.
>I asked who is paying for the tokens, they said its included in the license. I said, so is there a budget or is it all you can eat. they said good question they didnt know and would get back to me. I said the reason i asked was just one query there had a 250k token burn on it. and it was a fairly simple query about one device.
>then, one of the execs on their side was heard saying out loud "Why are we even showing this to the customers?"
>it have us quite a chuckle. But lesson learned... the cost of adding AI to anything isnt really being accounted for let alone the true cost of actually running the AI.
>all things AI are going to get more expensive. even if you dont want the AI aspect.
Of course they do. How else do you expect them to pay for that? If you buy a Foo from Acme, Inc, you aren’t only paying construction costs, either.
> On the other hand, once an open weight model is released, any inference provider can easily host it and just do some markup on inference cost. This proves way cheaper than running a frontier AI lab.
The only logical conclusion for commercial AI labs is to never release their models as open data, and try to stay ahead of open models. One way to do that is by having better models, another by having more users (because that decreases the per-user costs of creating the models, decreasing the price difference with companies running open models). The frontier labs are aiming for a combination of both.
1. Chat, being 3 yr old, is a fairly mature and solved problem today. Top companies aren't even talking about it anymore! Gemma 31B does it amazingly well (for $0.4/1M token output). Practically every near-SoTA and SoTA model does simple "chat-like" QA amazingly well -- summarization, basic question answering, single- or few-step search.
2. Tasks -- or knowledge work on a computer -- are the new frontier. Computers have become competent only recently, and only for some of the tasks so far. I'd guess another 2-3 yr development cycle, after which "el cheapo" models will be virtually indistinguishable from SoTA.
As tasks are the new game in town, AI labs can still charge a premium for it. That premium has disappeared already for chat; most users cannot tell 99% correct answer from 95% correct answer; nor do they always wish for maximum accuracy.
3. What comes after Tasks? I think today's AI startups should figure that one out and solve it before everyone else.
That has been the story of 200+ years of industrialization: new technology eliminates some jobs, but it also creates new industries, new demand, and new kinds of work.
We heard the same panic about radiologists. In 2016, Geoffrey Hinton famously suggested we should stop training radiologists because AI would outperform them. Yet in 2026, we need more radiologists, not fewer. The job is changing, not disappearing.
You even see a similar dynamic with immigration. Immigrants don’t just “take jobs”; they also create demand, start businesses, pay taxes, and expand the market. Remove them, and the economy often shrinks — meaning fewer jobs overall, not more.
TL;DR: AI is not simply “coming for your job.” Yes, the nature of work will change. We no longer employ “human calculators,” but society didn’t run out of work. We created better, more productive jobs than doing arithmetic by hand all day.
Who in hell would actually do this? That's a level of problem that any of the flash-class models can solve.
Hand that sort of thing to GPT-mini, Haiku, or DeepSeek Flash, and save the big guns for big architectural problems.
I agree that text capabilities are maybe hitting the limit of available training data.
But, the big AI platforms are now being used by so many people for so many things that this becomes the new data source for new and different types of abilities.
Theres also the fact that the AI story means they've been able to fund huge data center builds and hardware innovation. This will lift the existing AI/ML applications (e.g. robotics, sensing) in themselves, as well as the fact that they can be integrated with the text models in probably really useful ways.
So I think, maybe text abilities are nearing the end, but intelligence and other interfaces with the real world still have a lot of space to grow.
Not trying to be harsh, but that sounds like a skill issue. You have the language server to lean on; easy feedback loop; sub agent per type.
If the subscription is gutted by factor 2/5/10/20/65 to make it more profitable for Anthropic it will be harder for users to justify the subscription.
On the other hand 13,000 USD in API credits can go a very long way if used ergonomically. For instance using a max context length of 200k is multiplying your reach in comparison to 1m context.
The Chinese open weight models were always winning the AI race to zero where as the likes of Anthropic and OpenAI have no choice but to increase token costs.
Even Microsoft wants to use some of the Chinese models only realizing how expensive both the frontier models are. It turns out that Jevon's paradox does not exist in the US (it exists in China).
This "Tokenmaxxing" marketing stunt was a scam for the frontier models to raise even more money at unsustainable valuations.
The OpenAIs and Anthropics are going to get eaten by open source, I don't see prices going up, prices are going to crater. The models are going to be more and more commoditized.
The author mentions $54 in costs but the reality is that developers are paid around this much per hour.
What is likely to happen: LLM performance goes even higher and can do tasks that take humans days to accomplish. You then have to compare LLM cost with human cost - something the Author has forgotten in their analsys.
Sure, but imagine a situation where you've spent an hour going back and forth with the LLM trying to fix a problem and at the end of it you've only made minimal progress. Now you've spent an hour of your time AND $54 with little to show for it. It's a metric I don't think many people track: the cost of going in circles with an LLM for an extended period of time while burning tokens and still not resolving the problem.
I know the number of times I tried to do something where the answer was simple but I took a few days to get there.
And the reality is that other industries aren’t finding the use for LLMs as much as programmers are. Sure there are some benefits but you can’t fire your marketing department and replace it with AI
I feel the only ones losing are the AI startups and Google. This is why they're trying to morph into a social-media like experience of simulated human interaction that can monetize a certain demographic of vulnerable people.
1. How much it costs in terms of programmers' salaries?
2. Can DeepSeek do this (I bet it can) and how much it costs?
The fact the author ever had the idea of using a SOTA to solve do this means LLMs are actually quite cheap.
claude 5 mythos and fabled are 50$/MOutTok. Previous models were priced at 75$, so presumably they found the "too expensive" price point.
I don't believe in the model where OpenAI or Anthropic will own the whole value chain. They'll try and probably fail. They'll own lots of infrastructure and there's going to be a shortage of that for some time. But most of the value creation will happen upstream from them.
I also believe that before any real companies are running these models locally, they will already have some kind of agentic layer.
With the current frontier model lab progress, i do not see any real company which makes real money, running local models.
Running local models is easy for me, for sure not that easy for any company. Your DC needs to be able to host GPUs, it needs the cooling power, you need to have a DC. Without a DC, you need to have someone maintaining critical infrastrucutre, taking care of model evaluation etc.
For external parties, there might become a new business model: You might not hire an external anymore, but a token budget and the 'operator of the token budget'.
The current chip fabs are full, developing a high end / cheapisch local LLM Chip will still take a few years as long as the DC GPU demand is still as high as it is.
But for sure there will be use cases of very critical data, but at the end the question will still be how big they are in comparision to the rest of the market.
These cricial workloads also have the cost issue, right? so will they reduce workforce to compensate for the budget?
That's the only way I can see frontier labs charging high enough to sustain the cash flow needed to operate as racing to the bottom is not possible for them.
It is interesting to think whether this is another "Cambrian" era like the smartphone OSes when you had Symbian, Android, iOs, Windows Mobile and so many others competing.
So the hyperscalers already won for now probably.
At the end of the day, you send a lot of personal data to these endpoints. If you already host everything through microsoft already, LLM hosting is then a no brainer.
This isn't how coding models get better though. Why would this have anything to do with plateauing?
The other alternatives with LLMs becoming more expensive in an Uber-like move may not work due to a lot of competition. I also don't think usage will increase 10x. I don't always have coding tasks for an LLM despite it being good.
My reasons to believe so are outside of what interests HN community and I am neither endorsing this behavior, nor I think it is that simple. But US also has a huge debt that it must service. Wouldn't it be convenient if it was suddenly halved in actual value?
OpenAI and Anthropic will just go back to entirely healthy valuations of ~$5-10B each and the industry carries on.
anyone got a source? sounds juicy
8xB200[1] costs around 250k DIY and 450k from an enterprise builder so that will be our cost factor, these consume around 7.8kw at 100% load with median load of around 7kw (optimistic) which means that a 240kwh solar installation would be enough to supply it (72kwh buffer for bad weeks / winter) and that will set you back around $240k: this includes battery storage, installation and inverters, diy cost would be lower at around $160k.
This puts the cost of the entire system anywhere from $410k to $690k. This does not take in any property tax or land ownership into account since honestly it varies too much. The solar is simply used to provide a fixed cost basis for powering hardware instead of monthly recurring payments. Financing a 5 year loan would cost anywhere from $8,313 to $15,700.
Now let's do the math for glm-5.2[3], a fully optimized theoretical build can do around 1200tok/s which means that's around 13-14 streams of ~90tok/s on average, pushing batching further and limiting context size to around ~300k with ~150k median) you can achieve up to 37 streams at around ~40tok/s pushing performance envelope to 1400tok/s. This means you are able to generate 2.5B to 2.9B tokens in 4 weeks.
Which means putting the numbers together you can serve 1m tokens at $2.86 to $3.32 per million output tokens all else being equal. Considering that glm-5.2 is approaching opus level intelligence it's pretty safe to say that same applies for frontier labs. Input/cache write/cache reads are very difficult to price, so this assumes you're providing input / cache for free[4]. As a very heavy user I generate around 2M to 5M output tokens a day which would put me at $5.72 to $6.6 of cost per day totalling $200 a month[2].
What I also don't mention is that frontier labs have BY FAR the lowest cost per token out of any provider out there due to the amount of money they also invest into efficiency gains. This was proven by the fact that anthropic saw a huge exodus of openai users put strain on their systems and with efficiency optimizations alone they managed to mitigate a bulk of capacity issues, altho they did run into limits and had to begin spreading out the duck curve, but I have zero doubts they're getting percentage points of improvements month to month.
[1]: H300's are unobtanium unless you're building rack-rooms, H200's are not that cost effective and saturate too fast while having poorer efficiency, only capable of running flash tier models.
[2]: Okay, I didn't expect to arrive at the $200, this is kind of entertaining.
[3]: fp8, z.ai serves fp8 according to openrouter.
[4]: Assuming you want to charge for input / cache, cache reads make up roughly 30% of the cost, output 20% 50% input so to price it out it would be roughly $.3 for 1m input, $.015 for cache reads and $1.5 for output. Judging by https://openrouter.ai/z-ai/glm-5.2#pricing, appears that my math checks out.
a surprisingly large fraction of production workloads can be handled by smaller models with the right scaffolding. it's often easier to switch to a larger model than to engineer those pieces, so many teams never bother.
my intuition is that a lot of the current "ai cost crisis" is really an orchestration problem rather than a model pricing problem. before asking whether frontier pricing is sustainable, i'd first ask how much of that spend is simple tasks being sent to the smartest available model by default.
my bet for the next few years is that the model itself stops being where the value is. frontier models will become more like commodities, and the real difference will be the layer around them as routing each task to the cheapest model that can do it well, verifying the output, and only escalating when needed.
eventually, asking "which model do you use?" will sound a bit like asking "which cpu do you use?" the engine still matters, but the system built around it matters a lot more.