Our estimated spend for AIaaS would exceed that cost in less than a year.
In a few years, there will be hardware capable of running frontier models good enough for most things at accessible prices for even tiny companies.
Our estimated spend for AIaaS would exceed that cost in less than a year.
In a few years, there will be hardware capable of running frontier models good enough for most things at accessible prices for even tiny companies.
If open source models are ~3-6 months behind SOTA, and ~opus4.6 capabilities are good-enough for product market fit, do the frontier labs have half a decade to catch up on their prior burn?
AI cost ballooning faster than companies can afford is becoming a very common topic in my circles right now. The era of "I'll pay infinitely more for marginal gains" is over from what I can tell.
That's doing a lot of work here.
The future I see isn't most companies buying hundreds of thousands in hardware to run models, it's them adding a line item to their AWS bill. Inference costs on the larger hosted open source models are dramatically lower than the frontier labs API pricing.
That's the future Amazon sees too. We just had a week long session with the AWS team and they pushed that to us multiple times.
The days of requiring a data center to run anything resembling opus 4.6 are already counted. (But the industry will fight hard to get people to keep paying the Claude tax.)
And yeah, that may be the ~decade world, but we're in the mainframe era of the frontier models. It's going to be more economical for basically any consumer, and most businesses, to pay someone else to host models for quite a while.
I wanted the ability to run whatever cameras on a VLAN and own the stack.
Said model will also run as a tool-calling coding model excellently (it's no Opus, but for a thing that once set up is just the cost of energy, it's incredible). It can type faster than you can, probably 10x faster, so with guidance it'll make you faster. And it's free.
It's here. If folks want ChatGPT without a subscription, they can have it today on their computer. The only money to be made is in the high end models doing "serious business" work spanning 1M+ token contexts and massive uncertainty. Everything else is already set to be eaten by today's local models.
Here's a prompt I just ran against Claude Opus 4.7:
> Use python3 to experiment with whether the SQLite3 authorizer mechanism can be used to detect an INSERT OR REPLACE based just on running an explain query without examining the SQL string itself
Opus nailed it: https://claude.ai/share/c4212606-3fee-4b7c-bc97-505e0348ccac
I tried the same thing against qwen/qwen3.5-35b-a3b running locally in lmstudio, with the Pi coding agent. At first it looked like it was going to do great! And then it fell apart over the course of several tool calls: https://gisthost.github.io/?8ae2f842df619fb7fd8f1ccd82fe41c7
I'm used to GPT-5.5 and Opus 4.7 handling that kind of prompt without any problems at all.
I'm not an expert in SQLLite so I can't say if this is 100% correct, but it seemed directionally similar to the conclusion from claude.
### TL;DR
- Authorizer + EXPLAIN: No — authorizer only sees SQLITE_INSERT, not VDBE opcodes
- EXPLAIN opcode analysis alone: Yes — Delete opcode at position 10 is the unique signature of INSERT OR REPLACE / REPLACE
I can't help but think the not-so-distant future will see language models expected on commodity personal computing devices.I tried it a second time, and it spent a lot of time trying to figure out some authorization issue, so definitely not a slam dunk. I might run it a few more times for science. But while this is a new model it's also quite lightweight, and as hardware adapts and improves it seems inevitable that for many use-cases a packaged language model running locally will do the trick.
Not saying they're equivalent, local models still decohere much quicker as the context grows in my experience. But... Interesting.
I don't think we can discount this, frankly. Newer electronics are energy efficient, but older devices are more energy-intensive, and unless configured well, a gaming PC can easily use a few dollars a month in electricity, so now you're approaching subscription territory. A subscription comes with no upfront cost, higher reliability, no wasted space in your home, mobile apps, etc. (and less privacy).
I bet this will ironically be couched in "safety" reasons or regulation to get anti-AI folks on board, even if it favors the large incumbents.
Why would it not? The typical new phone today has 16gb of RAM. 20 years ago that was somewhere around 32mb. Factor 512. It's not hard to see that we'll get there rather soon, especially if there is an application that provides demand.
> You people fundamentally don't understand the memory requirements for running inference.
You seem to be overlooking how fast things change in this industry, especially if tons o money can be made as a consequence.
> Your cute local models seem good enough because you have no standards and anything an LLM produces seems like magic to you.
Please don't generalize. I'm an expressed AI skeptic and have to deal with the bad consequences of AI slop every day. But you can't deny that there are enough applicationn areas where people have use cases and those will be much easier if things don't need a few round trips to a data center that sucks all the electricity and water out of neighboring communities.
The iPhone 17 has like 8 gb, the Pixel 10 12.
The original iPhone was 128mb, and the iPhone 6 from 2016-2018 was around 1gb; that puts the iPhone at around 8x RAM per decade, and puts us at 128gb in our pockets at around 2036 or so.
(Incidentally, the big news in phone RAM is that a lot of new phones are dropping back to 4gb because of RAM shortages.)
Running software in the cloud gives you certain reliability and scaling advantages that would be very hard to replicate locally. Running some code agents in the cloud vs local hardware, if the local hardware gets "good enough," breaks the other way - offline usage, alone, would be hugely valuable to many people and companies.
It'd be very interesting to see where various players would decide to make a call "local is good enough" though. Buying the hardware isn't a small bet, if it's not something that ends up as part of your standard computer.
And 5% worse model for 10% of the price of the bleeding edge will be worth it for majority of people
People got a lot done before Opus 4.6. In 6 months, would you be dissatisfied by Opus-4.6-level open-weight models, just because Opus 4.8 will be out?
I hope there's a "good enough" point but I don't think we're there yet. Like for me hardware got good enough several years ago. But while opus 4.7 is really good compared to everything else, it's not so good that I would use it at a discount over whatever is available in a few months. The improvement in quality, speed, and daily frustration is worth it to me... Spoken as someone whose employer is footing the bill, so take that with a grain of salt.
I want to run my own local models, but I don't think that's feasible without lots of frustration until a few generations of frontier models are so good that they're almost indistinguishable for common tasks. Kind of like how MacBook pros have been for a while.
Using Cursor to hop between models, I've found Opus to be generally better at really tricky debugging than GPT 5.5 or earlier models, but not reliably better at execution because of these things. I'm not sure Composer 2.5 is quite there yet for the execution side, but it's getting pretty close to those other ones, such that I'm definitely still in a "debug and plan with slow, execute with faster ones" operating model for working on hard shit.
It actually seems somewhat difficult to train such a model since "all the text on the Internet" is easier to provide in bulk than a highly curated set.
But at the moment, I can't imagine why I wouldn't be spending the majority of my time with the best models. I'm spending a lot of time with them! Reducing the number of back-and-forths is extremely valuable to me.
I expect in two months I will still want to spend >80% of my time prompting the best models, and that's true if I were spending my own money on hobby projects, too.
These are ideas that simplify the design, reduce future work and tie together the entire system. If in two months I can arrive at ideas of that quality with normal brainstorming with llms that will be extremely valuable
would you be dissatisfied by Opus-4.6-level open-weight
models, just because Opus 4.8 will be out?
Well, I see what you mean, but two big concepts...1A. Models get stale pretty quickly w.r.t. new developments that occur past their cutoff date. "But you can just keep them current by linking them to never documentation, etc!" Well, no, you sorta can't -- at least not in perpetuity. Those search results fill up your context window real quick. So that gets unsustainable real quick.
1B. Even when your context has plenty of free space, the results you get from "here's a link to the documentation for this new framework that released after your cutoff date" absolutely pales to the results you get from knowledge that is fully baked into the trained model as opposed to your context window. For one thing, that documentation link you pasted into your context might link to... a dozen code examples. Whereas if that was baked into the model itself, the model might have been trained on many thousands of examples in Github etc.
2. It's also a reality that most professional engineers have to keep up with their peers and competitors. We can maybe say it shouldn't be that way, but it is. So if $SOME_NEW_MODEL is significantly better than 4.6... and my peers and or competitors are using it, then yeah I might but really feeling the need to match them. And I'm not even necessarily talking about some kind of cutthroat dog-eat-dog stack-ranked workplace.
These limitations aren't relevant for all use cases or careers but they're hiiiiiiiighly relevant for professional software engineering.
The people who are claiming Opus level capability does not have sufficiently complex problems to see the difference.
Kimi is close for example regarding SWE bench for code. For reasoning there are open models that surpass opus by quite a margin already.
A $50k - 100k rig could do it and an entire company would be able to use it a full speed.
The goalpost we've been bludgeoned with over and over again is that, in particular, Everything Changed in November 2025. That GPT 5.2 and Claude 4.5 were the inflection point. That is actually 6 months ago. And DeepSeek 4 is already there.
> run locally
You can't run DeepSeek locally on consumer hardware[1], but you can on enterprise hardware, and enterprise spend is the subject of this conversation -- and even if you aren't self-hosting, it doesn't matter, because you can just get your inference from one of the the many companies serving DeepSeek, who trivially undercut the pricing of OpenAI/Anthropic because they didn't have to spend hundreds of billions on training frontier from scratch but instead only invest in supporting inference, which is already profitable.
[1] Since this misconception comes up all the time, I'll go ahead and pre-empt it: no, training a 32b parameter model on outputs from DeepSeek and running that locally is not "running DeepSeek", despite the hundreds of stupid articles and Youtube videos making that idiotic claim that they're running it on a 5090.
Maybe not DeepSeek v4 Pro, but I've run DeepSeek v4 Flash on my 128GB MacBook Pro using antirez's carefully quantized https://github.com/antirez/ds4 and it's impressive.
I'd qualify that by writing that you can't run it with ordinary, real-time speed and throughput. If all you care about is slow and high-latency inference, there's no reason why that shouldn't be feasible even on the cheapest miniPC around, as long as it can literally store the model weights and keep around the (rather small) context.
Claude code was a lot of people's introduction to using coding agents that could do a lot more than copy-pasting from a chatbot or autocomplete.
Opus 4.6 quality for local inference would be revolutionary.
They just have to be useful enough that companies don't need the best.
They are.
Your argument rests on the "for marginal gains" part but it's really not clear that the gains are marginal in the foreseeable future.
We're 3.5 years into this current AI wave, and a lot of the valuations have been predicated on what you're arguing here -- that essentially should one of the labs make an order-of-magnitude improvement or hit escape velocity on recursive self-improvement they'd become the most powerful economic chokepoint in history.
The reality has been that given access to compute + capital all of the labs can stay pretty competitive with each other. Someone does a bit better on coding, someone else does a bit better on tool calling, and then they swap after each spending another $100bn.
The market looks like a commodity market where the commodity is intelligence, not a winner-take-all market with massive margins. Plenty of people get rich in oil and airlines, but they notably don't tend to be the innovators long term, they tend to be the operators. Obviously if the machines become sentient tomorrow, turn on their masters, and hit world-dominating intelligence, that assessment changes, but after several years of that narrative while objective reality looks quite different I think the more sober voices are starting to gain a foothold.
I'd pay a premium for even just a model that's 20% better, no ASI required, and I think a lot of people would. I wouldn't call that marginal, if it means I'm getting frustrated on 20% fewer tasks.
A recurring pattern that I've seen in myself and others is to at first be very impressed by a new model's coding capabilities, and then desensitize quickly and start being frustrated by the shortcomings.
The point I'm making is that I think we're rapidly hitting levels where corporate buyers aren't willing to pay multiple-times-more for marginal gains, and I expect that to become more the case over time, not less. You, and a small % of other power users in the market might tolerate a $400/month pro-supreme-plan for access to Mythos or whatever, but I don't think that's going to scale up in quite the same ways we've seen so far.
Even a year ago paying multiples times more for a 50% gain was very sensible for a lot of workflows. But if we're getting to "good enough" for things like coding, justifying to your CTO/CFO why the org should go from spending $1m/year to $5m/year for a 10% higher hit-rate on one-shot prompts from the engineers is a much tougher sell.
I remember that even when GPT-4 was king, the Gorilla paper showed that Llama 7B could be fine-tuned to outperform GPT-4 on tool calling.
On domains that don’t involve agentic tool calling*, I haven’t found the frontier to have advanced that much.
Edit: I should broaden this to domains that naturally lend themselves to RLVR training. Models are drastically better at math now.
I'm not sure why you're so obsessed with the non-codex versions
If we still were in the ZIRP era, busting the bubble would certainly kill off the world's economy for good simply due to its size.
The larger point I'm making is I think models are rapidly becoming commoditized. There is probably a small market long term that's willing to pay 10x for 10% marginal gains, but the majority of the buyers in the market will be economic and we're likely to have a lot of folks willing to spend 1/10 the cost for 90% of the performance, and plenty of companies that haven't raised hundreds of billions-trillions who can provide that.
A lot of the frontier labs valuations has been based on an assumption that 1-2 companies would get break-away intelligence that basically made them economic chokepoints indefinitely into the future. The reality that's becoming increasingly clear is that model quality is a pretty linear function of (cash burned - ability to copy other's homework) and the economics are starting to look a lot more like airlines than online advertising.
The economics of airlines are such that they generally earn a return on capital less than cost of capital.
I think this is exactly where we are heading and OAI-Anthropic are the concordes.
They know they do not and that’s why they’re all trying to IPO right now, so they can pass the bag to consumer investors
1- SpaceX + Tesla + xAI merger / IPO while Musk was vocal against IPO for about a decade
2- Warren Buffett cash at record highs
Someone got to be exit liquidity
Last year's AI models will be the same. Do you want to spend 3 hours prompting free AI to fix your code or 1 hour prompting AI you paid $20 for?
Just think how much further that $100K would have gone if the hardware market wasn't so screwed-up.
Anecdote: I priced-out adding 1TB of RAM to a four node cluster a couple months ago. The cluster was purchased in fall of 2024 w/ 4 nodes, each with 256GB RAM. The nodes cost just over $14K apiece back in 2024 (entire box, not just the RAM).
Dell wanted >$90K a couple months ago to add 256GB to each node.
RAM is expensive, but not THAT expensive. I just bought 128Gb for about $5k for our build cluster (it's not even for AI, sigh). Even if you need larger-sized DIMM sticks, it's still going to be in the vicinity of ~15k tops.
I haven't had problems w/ Dell support and 3rd party memory, personally, but given the machines' application I understood the concern.
The Gemini Flash is very good at searches. Just about any low end model can toss out a poem. All the higher end models (open source and otherwise) seem to be able to churn out code that passes tests. The smaller, "less capable" ones are much faster at it, which means in the hands of a skilled practitioner are the best choice for that task. But they rapidly fall apart where there isn't a hard source of truth (like a good test suite) to grind against. Because of that you have to use a bigger model for bug finding. In that task the open source models tend to fail on larger code bases, where something like Opus still shines. I gather Mythos is an absolute monster, and unparalleled, and unavailable. I'm sure one of the reasons for that is it's so expensive to run.
Or to put it another way - you don't use a 100 tonne crane to pick up the shopping. And ... the smaller models will happily run on in-house hardware. You may not do it today because of the current DRAM price and integrated NPUs have just started shipping, but in 5 years time models will be running on your phone.
It's a given that the SOTA models need to raise their prices. It's also a given that they can't. The more they raise the more customers will move to their competition.
So what happens next? Well I think it will suck horribly if you can't move off of SOTA sooner or later, because the Big Two are going to lose customers, and therefore have to raise prices on the locked in customers even more than these projections suggest.
Beyond that if you're looking to start a business, figure out how to use cheap models in new scenarios. Build software which does that and license it. This is kind of contrary to the idea that you shouldn't over optimize for deficiencies in the models that will likely go away in the next generation - for instance a lot of problems were solved when context windows got way bigger. So it's a thin line to walk but I think it's there because a lot of orgs are using Claude today for pretty basic tasks.
The dev who's addicted to SOTA models honestly is going to have to settle for less or get totally screwed. Most applications within business from what I see aside from complex research do not require SOTA. They summarize, they classify, they transform, and doing that accurately has been cheap for a while.
Your last point is common sense in my opinion, I agree with it. At the end of the day most employees are (by definition) of average intelligence and most businesses are average in complexity. Thus, it is logical that average tools (AI models) should do the job for most people and most businesses.
But what if your competitors sell their knowledge to AI companies?
Then you're still screwed.
AFAIK you would get about ~5 concurrent users, with a max context window of ~128K tokens on the larger models.
This wouldn't be good enough for coding -- are you guys thinking of using it for something else?
Roughly equivalent to 4x H200's for less than half the price.
Vaguely around 60k tokens per second...
The decadal move to all-cloud-all-the-time killed off in-house hardware teams while the C-suite chased their OpEx dreams.
It would be interesting if we come full circle on this.
There's not that much depth in a lot of 'everyday' writing. For many tasks that means that you don't need to be hyperintelligent - reading a recipe or a shopping list, reading a newspaper article, etc.
What you call harnesses I call… bullshit?
With a nice UI on top, for the desktop app too: [2]
[1]: https://developers.openai.com/codex/config-advanced#custom-m...
My single spark has me running Qwen 3.6 27B and antirez’s specially quantised DeepSeek v4 Flash (which is shockingly impressive)
How is that supposed to give good results?
I was going to say - the models are just going to keep growing at a pace exceeding the pace of hardware pricing/availability
But then I realised that, far more likely, there will be a plateau reached (again) where nobody is seeing gain, and at that point hardware will catch up
My guess is there’s gonna be some legislation or something “you can’t share anything over this level of complexity” and I think that that’s what a lot of that mythos rattling was all about
What makes you so confident about this prediction? Hardware costs haven't exactly been cratering recently.
No, but local models have been booming in performance/quality improvements. The RAM shortage won't last forever (more supply will come online when if demand doesn't diminish), and then the math would be pretty easy.
The current frontier? Sure. The frontier then? No - obviously that frontier is going to keep consuming available datacenter compute capacity, which will be better
https://www.gigabyte.com/Enterprise/GPU-Server/G383-R80-AAP1
There are physical limits to how much you can compress data and how much is needed for a capable model. If by hardware capable for running SOTA you mean a 7 figure investment for a company, than sure. But how come these companies didnt do the same thing for cloud? There's been this option for self hosting infrastructure for a decade but companies don't use it, they pay AWS.