A Practical Guide to Running Local LLMs
spin.atomicobject.com
spin.atomicobject.com
On macOS it's worth investigating the MLX ecosystem. The easiest way to do that right now is using LM Studio (free but proprietary), or you can run the MLX libraries directly in Python. I have a plugin for my LLM CLI tool that uses MLX here: https://simonwillison.net/2025/Feb/15/llm-mlx/
Whisperfile is also amazing. I used it to transcribe a podcast I hosted 20 years ago and it was fast and the errors it made were acceptable.
Upon getting a model up and running though, I quickly realized that I really have no idea what to use it for.
That's okay, it's not like anyone else has any clue what to do with LLMs, either.
There are a bazzillion and one hardware combinations where even RAM timings can make a difference. Offloading a small portion to a GPU can make a HUGE difference. Some engines have been optimized to run on Pascal with CUDA compute below 7.0, and some have tricks for newer gen cards with modern CUDA. Some engines only run on Linux while others are completely x-platform. It is truly the wild-west of combinatorics as they relate to hardware and software. It is bewildering to say the least.
In other words, there is no clear "best" outside of a DGX and Linux software stack. The only way to know anything right now is to test and optimize for what you want to accomplish by running a local llm.
Anyone experimented with local LLMs and compared the output from ChatGPT or Claude? The article mentions that they use local LLMs when they're not overly concerned with the quality or response time, but what are some other limitations or differences to running these models locally?
Compare that with the commercial models where a lot of that is done on a large scale for you.
I'd say it's suprisingly good at exacting reminders and todos from natural language into json 'action objects', or turning what is essentially a run-on sentence of a transcript into a markdown formatted text.
What I've found most fun to play with is to get it to extract metadata like an 'anxiety score' and tags.
Overall, it's clearly 'dumber' than the hosted big models, and in my case I have to deal with a small context window. In general my 'vibe' is that I have to be clearer and more explicit, and it's usually better to just do multiple passes over the same text with very targeted questions.
Oh, and in my case I definitely notice my laptop screeching to a halt when it's processing a big transcript, but in my case I can specifically delay those jobs to a time where I'm not at my computer.
End user experience should start by selecting the models of interest to run and output hardware builds with price tracking for components.
I'm wondering about multi-modal models, or generative models (like image diffusion models). For example, I was wondering about noise removal from audio files and how hard it would be to find open models that could be fine tuned for that purpose and how easy they would be to run locally.