So yeah, I think models on local hardware will be quite common soon among the tech savvy (such as people creating software).
A100 -> H100 was >3x tokens per joule, H100 -> B200 >10x. There are significant low-hanging fruit still available in architectural efficiency, and the vendors are chasing them.
This is the big risk for AI companies that I feel is not being sufficiently priced in. Almost none of the investments they are making are durable, the depreciation schedules for everything but the real estate should be less than 24 months. Until the hardware is stable enough that you only get double-digit % improvements per generation, it should almost be counted as opex.
As it stands there's way more demand than supply. The new GPUs are going to run frontier models while the older ones serve smaller ones.
That said some of these are running in tents hooked up to mobile turbines. I can see some of those going away but generally I think you'll see them used until they start to fail in 5-10 years.
E.g. grok isn't truly multi-modal, it has a callable tool that is a separate VLM it invokes on image URLs or files (for a long time it was grok-1.5v, but I think they have upgraded now, it was pretty bad).
And then you have the small summarizer models for the CoT/thought traces, the guidable summarizer models for the standard browse tools, etc.
There's a ton of stuff that can use an aging GPU.
I do hope you're right that it will get cheaper over time (it should), but right now 32GB of VRAM is not affordable to a lot of people. You're talking ~$4500 just for the GPU, or $800 ish used if you can find one.
It's a tad less efficient and a bit more of a hassle, but still a good experience for only a fraction of the price.
Gotta remember inflation here.
$1K in 1995 was roughly equivalent to $2K now and wouldn't have been a particularly "good" machine then.
In 1982 the Commodore 64 started at about $600 bucks, also roughly around $2K today.
If you outgrew that, beefier machines back then were A LOT. It was easy to find $2k+ towers and (especially) laptops even into the 2000s, and a lot of those would be $5K+ equivalent today.
I imagine having multiple providers competing will drive down hosted versions of open weight models drastically.
And we've barely started to scratch the surface on helping open-weight models "be the best they can be", with cloud burst parallel sampling and prompt mutation. Looking for best probabilistic results for a prompt, and looking for best prompt variants for a task. Adaptively scaling computes at generation, not just training.
And speculatively, if agentic coding is naturally a multiplicity, what UX might enable human devs to dance with that quantum superposition? Rather than quickly collapsing to one monkey and its keyboard.
I was thinking more of the providers of inference on open weight models that openrouter proxies to.
It's definitely worth investing in self-hosting the agent infrastructure around the model though: all the documents, knowledge base, all the connectors, the agent itself to run on your hardware
https://huggingface.co/Jackrong/Qwen3.5-27B-Claude-4.6-Opus-...
My use of the term model here is confusing
Especially because the world is likely to persist, at least for a while, in state where computing hardware demand drastically exceeds supply resulting in high prices for hardware. So why wouldn't you want to max out utilisation and amortize costs, at least for typical (non sensitive) use cases.
Started with computers around 2009 and later bought an oldish computer (a pentium 4 PC) for the equivalent of 50 usd. Codeblocks and Python Idle were free at the time (C and Python were the first languages I learned). The barrier to programming has always been low as the only thing you needed was books (the internet made things easier) and access to a PC (I had friends with laptop and my school lab).
Certainly the transistors/chip or transistors/$ or flops/$ have not been progressing at the same exponential rate as during 1970-2010. There is still progress, but it's rather slower.
As you point out it's really cost per transistor or cost per flop that we mostly care about. I'm finding it hard to find a succinct and clear plot, but I believe one is provided by Our World In Data on "GPU computational performance per dollar" [1] which, to my eyes, clearly shows exponential growth in computational power per dollar.
The picture for storage is a little more muddied but if you squint just right you might still be able to recover an uninterrupted exponential growth [2].
In my view, it's pretty clear that advances in AI have progressed so quickly because GPUs have been keeping up with the exponential growth of computational power (per unit cost).
Exponential growth in this area is usually characterized by "S-curves", where one technology gets saturated but the exponential increase in power or decrease in cost is picked up by another, adjacent, technology, that allows the growth to continue. For compute it's CPUs to GPUs. For storage it's platter drives that are now being overtaken by SSDs.
The more general phenomena is called Wright's law, or experience curve effects [3].
[0] https://en.wikipedia.org/wiki/Moore%27s_law
[1] https://ourworldindata.org/grapher/gpu-price-performance?ySc...
[2] https://ourworldindata.org/grapher/historical-cost-of-comput...
Possibly it's the same price range, allowing for inflation.
> It was only in 2025, as memory prices began an unprecedented surge, that the memory makers started to build new fabs targeted at HBM, all slated to start producing chips in 2027 or 2028.
If you want to argue that this is different from all previous RAM shortages, you can, but the burden of proof is on you to show the difference.
this time demand doesn't stop. there is an exponential demand for tokens.
[citation needed]
There is certainly economic pressure to create an exponential demand for tokens, but we've already seen a pullback from the costly "token maxing" companies were pushing last year.
It's also pretty unclear to what degree the RAM shortage is driven by inference (versus by training). We're rapidly approaching the point where frontier models are "good enough" for everyday use, are at some point we're going to hit diminishing returns on training new trillion-parameter models...