Based on what? The RAM requirements alone are extraordinary.
No, running large models on shared, dedicated hosted hardware at full utilization is going to be vastly more cost-efficient for the foreseeable future.
I've got a 128GB strix halo staying warm at home, it has nothing on top models with big budget. It's good supplement to low end plans for offloading grunt work / initial triage
Thanks for suggestion tho, tool by antirez is always going to pique interest, I'll check it out when I'm finally home again
Tho says Metal / CUDA, so doesn't seem friendly to Linux AMD system
I wish this was true but it is not. And I am working on open source models so if anything, I would have a bias towards agreeing with you.
Frontier closed models (GPT/Claude) are gaining distance to everybody else. Even Google, once the king.
Your claim is a meme coming from benchmark results and sadly a lot of models are benchmaxxed. Llama 4, and most notably the Grok 3 drama with a lot of layoffs. And Chinese big tech... well they have some cultural issues.
"Qwen's base models live in a very exam-heavy basin - distinct from other base models like llama/gemma. Shown below are the embeddings from randomly sampled rollouts from ambiguous initial words like "The" and "A":"
https://xcancel.com/N8Programs/status/2044408755790508113
---
But thank god at least we have DeepSeek. They keep releasing good models in spite of being so seriously resource constrained. Punching well above their weight. But they are not just 6 months behind, either.
[0] US AI firms team up in bid to counter Chinese 'distillation' (Apr 7) https://finance.yahoo.com/sectors/technology/articles/us-ai-...
Case in point: North Korea, with far, far fewer resources.
Gemma 4 was a major improvement is self-hostable local models and Qwen-3.6-A34B is a beast, and runs great on an MBP (and insanely well on a 4090).
The biggest lift is combining these models with a good agent harness (personally prefer Hermes agent). But I’ve found in practice they’re really not benchmaxxing. I’ve had these agents successfully hand a few non-trivial research projects that I wouldn’t have been able to accomplish as successfully even last year.
When you add in the open-but-not local models, Kimi, GLM, Minimax, you have a lot of very nice options. For personal use anything I don’t use local models for I give to my Kimi 2.6 powered agent.
Over-promising is a very stupid thing. Nobody will value the intermediate steps. Nobody will value all the effort because they will always compare us with frontier models made with billions and we will become a running joke. So please stop.
At what tps? You can run the new gemini flash or 5.3 codex spark at 1000+tps and run circles "open" models. You can't run anything useable locally without at the very least a blackwell 6000 if not two
Sure you can run qwen 3.6 at 20tps on a mac 128gb but let's not pretend this will get you anywhere
That is only true right now because hundreds of billions of dollars are being burned by these AI companies to try to win market share. If you paid what it actually cost, your comment would likely be very different.
Digital sovereignty laws may mandate/remove access to LLMs of other countries on economic and national security grounds.
They'll be controlling lights and temperature, they'll be adding calendar reminders that show up on your phone and your fridge. Your phone and devices might sync pictures and videos there instead of the large cloud providers. They'll also be a media server, able to stream and multiplex whatever content you want through the home. They'll also be a VPN endpoint, likely your home router, maybe also a wifi access point.
I think this makes quite a bit of sense. I don't think they'll be ubiquitous, but they could be.
This distributes the power demand where local solar generation can supplement , gives the home user a lot of control, and claims overship of the user data from big tech.
Maybe I'm imagining things but this is what I think is coming.
It's the lmm/data heart of the home. A useful digital tool.
We'll have this massive machine to do "home automation", something that by all rights should be possible with less computing than is deployed in smartwatches today. Yuck...
That's just one piece of the puzzle. If you're running the LLM there's no reason your family's mobile devices couldn't use said home LLM box to save battery life on their devices while maintaining control of their data, searches, photos, files, etc.
I use a local LLM with it, but you can use a hosted LLM if you like.
The core home automation stuff can run on a potato. The LLM just writes new automations when I ask it, or acts as a natural language interface.
I use a pretty small 4B parameter local LLM, on a fairly modest mini PC. It doesn't take a frontier model to do that kind of work.
I agree that the market for local AI is basically limited to nerds at this point, but that's because nobody's really explained why local AI is a good thing and also because the vast majority of people need the $20 paid plan at most. How much time and money would it take to get something half as good as OpenAIs products running locally?
There are a lot of good things that need to be explained to people, but nobody ever managed to. I don't think this will be any different.
> because the vast majority of people need the $20 paid plan at most
Exactly, people are not gonna invest time and money when there's already something else that satisfies their need.
Local AI will need to be both better and more convenient in order to be adoped by the masses.
Moreover local AI is not free. You'll need proper hardware to run it which you have to pay extra for it.
However, if you can deliver 90% of the value of AI for 90% less cost, that is a really big incentive. Companies will spring up to fill that kind of gap.
Nobody can undercut the big AI players right now because they are all over-funded by VC money. Once the frontier companies try to match cost to expense, suddenly they become very, very vulnerable.
They’re still pricey, the world is still scaling up memory production, and a lot of code isn’t yet built for AMD, but we went from the Wright’s brothers first airplane to jet engines in 27 years.
I’m not sure “it’s only a few years away” but we are sure moving there fast.
Nitpick: more like 36 years, from Wright Flyer in 1903 to Heinkel 178 in 1939. Still quite impressive.
The print shop can’t replicate the practicality of local printing and I can’t replicate their scale of investment. Both coexist perfectly.
Cynically: it’s become an executive-level gpu measuring contest. If you’re not making huge commitments on data centers, you can’t be a serious player.
Realistically: It’s a mix of the two. The recent Claude caps for agentic usage suggest that demand exceeded their immediate compute supply. That they can alleviate it with additional capacity from the existing and small-ish xAI facility suggests that either demand may not be rising quite as fast as anticipated, that they’re okay in the short term until more capacity comes online, or a mix of both.
Open questions:
1. At what price point does demand fall, and are the frontier providers overall profitable before that price point?
2. At what price/performance point do on-prem local models make more sense than cloud models?
I must say that the largest dedicated hosted hardware providers now, like Amazon or Google, to a large extent do not produce the software they are offering as a hosted solution (like Linux, Postgres, Redis, Python, Node, etc). Similarly I'm not sure if the producers of the frontier models are going to keep their lead as the service providers for the most widely used models. They would need to have quite a bit of an edge above open-weights models.
Also, models are given very sensitive data to process. For large organizations, the shared dedicated hardware may look like a few (dozens of) racks in a datacenter, rented by a particular company and not shared with any other tenants.
I take it you haven’t actually run any of the current gen local models?
They all fit on fairly accessibility hardware, and their performance is at least on par with what I was paying for last year.
I have one of my agents running entirely from a local model running on a MBP and it has repeatedly shown it’s capable of non-trivial tasks.
Playing around with another, uncensored, local model on my 4090 desktop has me finally thinking about canceling my personal Anthropic subscription. Fully private, uncensored chat is a game changer.
For work it’s still all private models but largely because, at this stage, it’s worth paying a premium just to be sure you’re using the best and it saves the time of managing out own physical servers. But if we got news tomorrow that Anthropic and OpenAI were shutting down, a reasonable setup could be figured out pretty quickly.
Edit: ah I see the models mentioned in another comment of yours
I also have an agent using Kimi 2.6 as a backend (which is open, but not local) and for some coding tasks as well.
Maybe your statement is true for smaller codebases and shorter conversations, but I’d be surprised if you actually achieve good results on millions of lines of code with a million token context.
Granted if your setup works well for your workload then that’s all you need.
At the same time, $100 a month is A LOT of RAM.
No one can deny that right now these new compact models are not as good as frontier models but for the first time we actually have competent local-first models. If I give you a local model that runs on your current hardware and performs at 75% of the ability of a frontier private paid model, would you still pay for frontier? More importantly, would you hand control of your processes and code to them knowing enshitifcation and price-hikes are always lurking nearby?
For businesses, I get it you want to compete. But personally, it's over. Even if I considered for a second paying OpenAI/Claude, not gonna happen now.
They have to keep getting better to stay ahead of each other and open weight.
Which means it's the opposite of a timebomb, the article has it completely backwards, tokens at current level of reasoning will continue to get cheaper.
I'm not sure 'local' will be the end state, as hardware needs are high. But certainly competitive forces tend to push profit margins toward zero.
Extended discussion on this topic:
https://corecursive.com/the-pre-training-wall-and-the-treadm...
Boss is happy, very happy. We're rolling it out more widely now.
But this is the future.
I seriously doubt it. Scaling is already strained (don't buy into the "exponential" hype). And, in any case, the competition will be against the frontier models that will exist in two years.
We are only 2-4 years away from consumer grade immutable-weight ASICs.
You might be interested in the tiny tape out project, which guides you through the process of getting your own design etched on silicon. If you only need larger features and not the next gen single digit nanometer stuff, you may not be so supply constrained.
The issue is the very huge amount of DRAM and high bandwidth these model require.
Also, how many companies will just buy an M6/M7 MacBook Pro with 32GB+ of RAM in a couple of years and get “free” AI along with the workstation they were going to buy anyway?
The big question I'd be asking if I was investing in one of the big players is if those changes are "it can do 99% instead of 97% of the tasks a user will throw at it" (at which point going local and taking back cost control/ownership makes a lot of sense, especially for companies) OR "it will fully replace a human with better output"?
I already don't need Opus for a lot of my tasks and choose instead faster/cheaper ones.
The former is a company that's gonna be trying to sell mainframes against the PC. The latter is a company that is in potentially huge demand, assuming the replaced humans end up with other ways of getting money to still be able to buy stuff in the first place. ;)
But even if scaling plateaus for the frontier models, maybe distillation will improve to the point where smaller more manageable models can reach the same plateau. That would be great for local.
Unless there isn't some important breakthrough in hw production or in models architecture, it's quite the opposite: bigger, more expensive and more energy-intensive hw is needed today compared to 1 or 2 years ago.
And how many tokens would that buy?
I hardly doubt there will be consumer grade HW to run it in 2 years either. And deep seek v4 pro is not even close to OAI or anthropic frontier models.
Already today is not possible to run deep seek v4 pro locally, and I cannot imagine that in 2 years we will be.
Local models never reach the % utilization that cloud providers have (80%+), and they’re always going to be much better than local models for this reason.
It’s not unreasonable to suppose that in 2 years time an opus 5 quality model will be etched into silicon for high performance local inference. Then you just upgrade your model every 2-3 years by upgrading your hardware.
Is an example startup in this area claiming 16k tok/s on an asic for llama 8b. Qwen has a 27b model at opus 4.5 quality.
It's a measure of a very thin sort of "value/$" that excludes a lot of other things that could be of value to a business, like control, predictability, and availability.
Thin clients have been going away for a long time. The trend has been to continue to push higher levels of compute into ever-smaller and ever-more-portable devices.
But in many cases self hosted or dedicated boxes are cheaper than cloud.
Eventually, we'll see. Frontier models still need some pretty serious hardware which will slowly come down in cost. Smaller models are becoming more capable, which will presumably continue to improve.
I think there's still a pretty big gap, though. Claude estimates Opus 4.6 and GLM-5 need about 1.5Ti VRAM. It puts gpt-5.5 around 3-6Ti of VRAM.
That's 8x Nvidia H200 @ ~$30k USD each. Still need some big efficiency improvements and big hardware cost reduction.
same with models.
It would cost me $300 in normal deepseek v4 pricing (non discounted) PER DAY, but I get it all for $500 worth of subscriptions.
To run deepseek v4 class model, you would need to spend $120k just in gpus.
Not even when that site calls itself "market" to create plausible deniality.
And yet, less than 0.01% of the population (made up number, but I am more likely to be overestimating than underestimating) do so.
Running local models to do real work is likely to be another niche hobby.
AI is the future operating system of every computer everywhere