And as far as I understand it, an Air with an M3 is perfectly capable of running larger models (albeit slower) if it had the memory.
That said they have some elasticity when it comes to the DRAM shortage.
An HP Zbook with an AMD 395+ and 128Gb of memory apparently lists for $4049 [0]
An ASUS ROG Flow z13 with the same spec sells for $2799 [1] - so cheaper than Apple, but still a high price for a laptop.
[0] https://hothardware.com/reviews/hp-zbook-ultra-g1a-128gb-rev...
[1] https://www.hidevolution.com/asus-rog-flow-z13-gz302ea-xs99-...
You don't necessarily need to go the maxed up SKU.
Thanks for helping me see it!
From what I understand, getting a non-Apple solution to the problem of running LLMs in 64GB of VRAM or more has a price tag that is at least double of what you mentioned, and likely has another digit in front if you want to get to 128GB?
> However, for the average laptop that’s over a year old, the number of useful AI models you can run locally on your PC is close to zero.
This straight up isn’t true.
Also, macOS only has around 10% desktop market share globally.
https://www.mactech.com/2025/03/18/the-mac-now-has-14-8-of-t...
Though maybe it depends on what you're doing? (Although if you're doing something simple like embeddings, then you don't need the Apple hardware in the first place.)
Do you work offline often?
Essential.
https://pmc.ncbi.nlm.nih.gov/articles/PMC12067846/
Who cares if result is right / wrong etc as it will all be different in a year … just interesting to see a test of desktop class hardware go ok.
I found that for this method the smaller the model, the better it works, because smaller models can generally handle it, and you benefit more from iteration speed than anything else.
I don't have hardware to run even tiny LLMs at anything approaching interactive speeds, so I use APIs. The one I ended up with was Grok 4 Fast, because it's weirdly fast.
ArtificialAnalysis has a section "end to end" time, and it was the best there for a long time, tho many other models are catching up now.
I found only one great application of local LLMs: spam filtering. I wrote a "despammer" tool that accesses my mail server using IMAP, reads new messages, and uses an LLM to determine if they are spam or not. 95.6% correct classification rate on my (very difficult) test corpus, in practical usage it's nearly perfect. gpt-oss-20b is currently the best model for this.
For all other purposes models with <80B parameters are just too stupid to do anything useful for me. I write in Clojure and there is no boilerplate: the code reflects real business problems, so I need an LLM that is capable of understanding things. Claude Code, especially with Opus, does pretty well on simpler problems, all local models are just plain dumb and a waste of time compared to that, so I don't see the appeal yet.
That said, my next laptop will be a MacBook pro with M5 Max and 128GB of RAM, because the small LLMs are slowly getting better.
Most laptops can run at best a 7-14b model, even if you buy one with a high spec graphics chip. These are not useful models unless you're writing spam.
Most desktops have a decent amount of system memory but that can't be used for running LLMs at a useful speed, especially since the stuff you could run in 32-64GB RAM would need lots of interaction and hand holding.
And that's for the easy part, inference. Training is much more expensive.
I have 16GB ram. I use unsloth quantized models like qwen3 and gpt-oss. I have some MCP servers like Context7 and Fetch that make sure the models have up to date information. I use continue.dev in VSCode or OpenCode Agent with LM Studio and write C++ code against Vulkan.
It’s more than capable. Is it fast? Not necessarily. Does it get stuck? Sometimes. Does it keep getting better? With every model release on huggingface.
Total monthly cost: $0
Hello, from outside of California!
but it’s more than I have!
However, I agree with the article that people will run big LLMs on their laptop N years down the line. Especially if hardware outgrows best-in-class LLM model requirements. If a phone could run a 512GB LLM model fast, you would want it.
Uber is economical, too; but folks prefer to own cars, sometimes multiple.
And how there's market for all kinds of vanity cars, fast sportscars, expensive supercars... I imagine PCs & Laptops will have such a market, too: In probably less than a decade, may be a £20k laptop running a 671b+ LLM locally will be the norm among pros.
When LLM use approaches this number, running one locally would be, yes. What you and other commentator seem to miss is, "Uber" is a stand-in for Cloud-based LLMs: Someone else builds and owns those servers, runs the LLMs, pays the electricity bills... while its users find it "economical" to rent it.
(btw, taxis are considered economical in parts of the world where owning cars is a luxury)
An Uber car can provide several.
One time I took an Uber to work because my car broke down and was in the shop and the Uber driver (somewhat pointedly) made a comment that I must be really rich to commute to work via Uber because Ubers are so expensive
The amount of compute in the world is doubling over 2 years because of the ongoing investment in AI (!!)
In some scenario where new investment stops flowing and some AI companies go bankrupt all that compute will be looking for a market.
Inference providers are already profitable so with cheaper hardware it will mean even cheaper AI systems.
That surprises me, do you remember where you learned that?
Here's a few good ones:
https://github.com/deepseek-ai/open-infra-index/blob/main/20... (suggests Deepseek is making 80% raw margin on inference)
https://www.snellman.net/blog/archive/2025-06-02-llms-are-ch...
https://martinalderson.com/posts/are-openai-and-anthropic-re... (there's a HN discussion of this where it was pointed out this overestimates the costs)
https://www.tensoreconomics.com/p/llm-inference-economics-fr... (long, but the TL;DR is that serving Lllama 3.3 70B costs around $0.28/million tokens input, $0.95 output at high utilization. These are close to what we see in the market: https://artificialanalysis.ai/models/llama-3-3-instruct-70b/... )
> The amount of compute in the world is doubling over 2 years because of the ongoing investment in AI (!!)
All going into the hands of a small group of people that will soon need to pay the piper.
That said, VC backed tech companies almost universally pull the rug once the money stops coming in. And historically those didn't have the trillions of dollars in future obligations that the current compute hardware oligopoly has. I can't see any universe where they don't start charging more, especially now that they've begun to make computers unaffordable for normal people.
And even past the bottom dollar cost, AI provides so many fun, new, unique ways for them to rug pull users. Maybe they start forcing users to smaller/quantized models. Maybe they start giving even the paying users ads. Maybe they start inserting propaganda/ads directly into the training data to make it more subtle. Maybe they just switch out models randomly or based on instantaneous hardware demand, giving users something even more unstable than LLMs already are. Maybe they'll charge based on semantic context (I see you're asking for help with your 2015 Ford Focus. Please subscribe to our 'Mechanic+' plan for $5/month or $25 for 24 hours). Maybe they charge more for API access. Maybe they'll charge to not train on your interactions.
I'll pass, thanks.
> All going into the hands of a small group of people that will soon need to pay the piper.
It's not very small! On the inference side there are many competitive providers as well as the option of hiring GPU servers yourself.
> And historically those didn't have the trillions of dollars in future obligations that the current compute hardware oligopoly has. I can't see any universe where they don't start charging more, especially now that they've begun to make computers unaffordable for normal people.
I can't say how strongly I disagree with this - it's just not how competition works, or how the current market is structured.
Take gpt-oss-120B as an example. It's not frontier level quality but it's not far off and certainly gives a strong redline that open source models will never get less intelligent than.
There is a competitive market in hosting providers, and you can see the pricing here: https://artificialanalysis.ai/models/gpt-oss-120b/providers?...
In what world is there a way in which all the providers (who are want revenue!) raise prices above the premium price Cerebas is charging for their very high speed inference?
There's already Google, profitable serving at the low-end at around half the price of Cerebas (but then you have to deal with Google billing!)
The fact that Azure/Amazon are all pricing exactly the same as 8(!) other providers as well as the same price https://www.voltagepark.com/blog/how-to-deploy-gpt-oss-on-a-... gives for running your own server shows how the economics work on NVidia hardware. There's no subsidy going on there.
This is on hardware that is already deployed. That isn't suddenly going to get more expensive unless demand increases... in which case the new hardware coming online over the next 24 months is a good investment, not a bad one!
No one is leaving an H100 cluster not running because the power costs too much - this is why remnants markets like Vast.ai exist.
But the operational costs are much lower than some people in this thread seem to think.
You can find a safe margin for the price by looking at aggregators.
https://gpus.io/gpus/h100 is showing $1.83/hour lowest price, around $2.85 average.
That easily pays running costs - a H100 server with cooling etc is around $0.10/hour to keep running
And a massive overbuild pushes prices down not up!
which is funded by the dumping
when the bubble pops: these DCs are turned off and left to rot, and your capacity drops by a factor of 8192
What dumping do you mean?
Are you implying NVidia is selling H200s below cost?
If not then you might be interested to see that Deepseek has released there inference costs here: https://github.com/deepseek-ai/open-infra-index/blob/main/20...
If they are losing money it's because they have a free app they are subsidizing, not because the API is underpriced.
The situation might be a lot different than people selling ex-crypto mining GPUs to gamers. There might be a lot of effective scrap that is no longer usable when it is no longer part of a some companies technological fever dream.
I don't see why consumer hardware won't evolve to run more LLMs locally. It is a nice goal to strive for, which consumer hardware makers have been missing for a decade now. It is definitely achievable, especially if you just care about inference.
so stop it
You need 96gb or 128gb to do non trivial things. That is not yet 749 usd
While browsing the Apple website, it looks like the cheapest Macbook with 64 GB of RAM is the Macbook Pro M4 Max with 40-core GPU, which starts at $3,899, a.k.a. more than five times more expensive than the price quoted above.
You can easily run models like Mistral and Stable Diffusion in Ollama and Draw Things, and you can run newer models like Devstral (the MLX version) and Z Image Turbo with a little effort using LM Studio and Comfyui. It isn't as fast as using a good nVidia GPU or a cloud GPU but it's certainly good enough to play around with and learn more about it. I've written a bunch of apps that give me a browser UI talking to an API that's provided by an app running a model locally and it works perfectly well. I did that on an 8GB M1 for 18 months and then upgraded to a 24GB M4 Pro recently. I still have the M1 on my network for doing AI things in the background.
Yes, the models it can run do not perform like chatgpt or claude 4.5, but they're still very useful.
If you try to get them to compose text, you'll end up seeing a lot less variety than you would with a chatgpt for instance. That said, ask them to analyze a csv file that you don't want to give to chatgpt, or ask them to write code and they're generally competent at it. the high end codex-gpt-5.2 type models are smarter, may find better solutions, may track down bugs more quickly -- but the local models are getting better all the time.
Anyway, I'm on a mission to have no subscriptions in the New Year. Plus it feels wrong to be contributing towards my own irrelevance (GAI).
The article kinda sucks at explaining how NPUs aren’t really even needed, they just have potential to make things more efficient in the future rather than depending on the power consumption involved with running your GPU.
A Lenovo T15g with a 16gb 3080 mobile doesn’t do too badly and will run more than just Windows.