I’m open to suggestions, but the alternatives outlined in the blog post ain’t it.
I’m open to suggestions, but the alternatives outlined in the blog post ain’t it.
> LM Studio gives you a GUI if that’s what you want. It uses llama.cpp under the hood, exposes all the knobs, and supports any GGUF model without lock-in.
> Jan(https://www.jan.ai/) is another open-source desktop app with a clean chat interface and local-first design.
> Msty(https://msty.ai/) offers a polished GUI with multi-model support and built-in RAG. koboldcpp is another option with a web UI and extensive configuration options.
API wise: LM Studio has REST, OpenAI and more API Compatibilities.
So no, they are not alternatives to ollama
As other posters report, now llama-server implements an OpenAI compatible API and you can also connect to it with any Web browser.
I have not tried yet the OpenAI API, but it should have eliminated the last Ollama advantage.
I do not believe that the Ollama "curated" models are significantly easier to use for a newbie than downloading the models directly from Huggingface.
On Huggingface you have much more details about models, which can allow you to navigate through the jungle of countless model variants, to find what should be more suitable for yourself.
The fact criticized in TFA, that the Ollama "curated" list can be misleading about the characteristics of the models, is a very serious criticism from my point of view, which is enough for me to not use such "curated" models.
I am not aware of any alternative for choosing and downloading the right model for local inference that is superior to using directly the Huggingface site.
I believe that choosing a model is the most intimidating part for a newbie who wants to run inference locally.
If a good choice is made, downloading the model, installing llama.cpp and running llama-server are trivial actions, which require minimal skills.
For a (brand new!) newbie, it's very, very likely to be information overload.
They're still at the start of their journey, so simple tends to be better for 90% of users. ;)
LMStudio is listed as an alternative. It offers a chat UI, a model server supporting OpenAI, Anthropic and LMStudio API interfaces. It supports loading the models on demand or picking what models you want loaded. And you can tweak every parameter.
And it uses llama.cpp which is the whole point of the blog post.
llama-server -hf ggml-org/gemma-4-E4B-it-GGUF --port 8000 (with MCP support and web chat interface)
and you have OpenAI API on the same 8000 port. (https://github.com/ggml-org/llama.cpp/tree/master/tools/serv... lists the endpoints)
That's what I meant by model management. I'm too tired to scroll through a bazillion models that all have very cryptic names and abbreviations just to find the one that works well on my system with my software stack.
I want a simple interface that a tool like me can scroll through easily, click on, and then have a model that works well enough. If I put in that much brain power to get my LLM working, I might as well do the work myself instead of using an LLM in the first place.
2. Choose the model they recommend
3. Run the one-liner the site gives you
Bonus: faster access to latest models and better memory usage
Do you think that this 229B parameter model will work on my consumer PC?
Stop pretending like HF is in any way beginner friendly.