> However, for the average laptop that’s over a year old, the number of useful AI models you can run locally on your PC is close to zero.
This straight up isn’t true.
> However, for the average laptop that’s over a year old, the number of useful AI models you can run locally on your PC is close to zero.
This straight up isn’t true.
Most laptops can run at best a 7-14b model, even if you buy one with a high spec graphics chip. These are not useful models unless you're writing spam.
Most desktops have a decent amount of system memory but that can't be used for running LLMs at a useful speed, especially since the stuff you could run in 32-64GB RAM would need lots of interaction and hand holding.
And that's for the easy part, inference. Training is much more expensive.
I have 16GB ram. I use unsloth quantized models like qwen3 and gpt-oss. I have some MCP servers like Context7 and Fetch that make sure the models have up to date information. I use continue.dev in VSCode or OpenCode Agent with LM Studio and write C++ code against Vulkan.
It’s more than capable. Is it fast? Not necessarily. Does it get stuck? Sometimes. Does it keep getting better? With every model release on huggingface.
Total monthly cost: $0
Though maybe it depends on what you're doing? (Although if you're doing something simple like embeddings, then you don't need the Apple hardware in the first place.)
Do you work offline often?
Essential.
https://pmc.ncbi.nlm.nih.gov/articles/PMC12067846/
Who cares if result is right / wrong etc as it will all be different in a year … just interesting to see a test of desktop class hardware go ok.
I found that for this method the smaller the model, the better it works, because smaller models can generally handle it, and you benefit more from iteration speed than anything else.
I don't have hardware to run even tiny LLMs at anything approaching interactive speeds, so I use APIs. The one I ended up with was Grok 4 Fast, because it's weirdly fast.
ArtificialAnalysis has a section "end to end" time, and it was the best there for a long time, tho many other models are catching up now.
I found only one great application of local LLMs: spam filtering. I wrote a "despammer" tool that accesses my mail server using IMAP, reads new messages, and uses an LLM to determine if they are spam or not. 95.6% correct classification rate on my (very difficult) test corpus, in practical usage it's nearly perfect. gpt-oss-20b is currently the best model for this.
For all other purposes models with <80B parameters are just too stupid to do anything useful for me. I write in Clojure and there is no boilerplate: the code reflects real business problems, so I need an LLM that is capable of understanding things. Claude Code, especially with Opus, does pretty well on simpler problems, all local models are just plain dumb and a waste of time compared to that, so I don't see the appeal yet.
That said, my next laptop will be a MacBook pro with M5 Max and 128GB of RAM, because the small LLMs are slowly getting better.
Also, macOS only has around 10% desktop market share globally.
https://www.mactech.com/2025/03/18/the-mac-now-has-14-8-of-t...
Hello, from outside of California!
but it’s more than I have!