llama.cpp
llama.app
llama.app
Meaning that you (and by that I mean your AI agent that has read the llama.cpp code) can write an ini file pointing to your models with parameters optimized for the specific model on your specific hardware. (Optimized by you through testing. Not that AI)
Then, any api client can just select a model and the system does the right thing.
It's great software. It just works.
__
You just need to ignore the cargo culting commandline options on social media. But you should be listening to the devs.
Have you already enabled ngram-mod (or rather just spec-default)? It is practically free.
Why not optimized by AI through testing ? Give it a test set to work on and let it loose.
Intent is the answer and AI has none.
Adding llama-swap adds a fair amount of complexity that is unpleasant to debug and its documentation isn't very good (the only real doc is the sample config file).
Personally, I don't see the value in idle timeouts or time based eviction. It doesn't improve model load performance but does mean you're more likely to incur a load penalty for any request.
I generally use the same two models for my agentic workloads and they sit comfortably next to one another in the 32GB of VRAM on my GPU. If I need to load a larger model, llama-swap can easily eject both models if necessary.
The gulf is quickly shrinking though and it seems like the need for llama-swap will disappear soon.
> but you might not be aware that llama-server can do multi-model for a while now
you will see that the sentence structure clearly implies both a change compared with a prior state and also lack of any third-party thing.
So the answer to the question has already been encoded as text available.
_
I can see the desire for explicit validation though. For that, I would propose a sentence structure like
> Oh cool! That means that llama-swap is now superseded/no longer needed?
That shows that you've read and understand the message, gives you the double-check and might on top spark a conversation about how these solutions compare. Plus that if the guy you're commenting too has spoken nonsense, they need to backpedal.
Does the llama.cpp UI provide the same? If not, it is too early to say that llama-swap is “superseded/no longer needed.”
My machine can only load one model at a time. The loading/unloading times just don't seem worth the switch. I tend to use qwen3.6 for anything and that's it. Then again, I am a simple coder.
ggerganov and the team have done a stellar job maintaining the quality while still being fast to implement new models/improvements.
Which last time I checked, both use different formats of the weights, the former GGUF while the latter .safetensors. I mostly end up using vLLM these days and I'm a bit more performance sensitive than what I used to be. Just a shame it's a hassle to share the weights between them with conversion and what not, either batched or on-startup.
I’m still waiting on 98.css to become the standard for vibe coded sites. You don’t have to read docs anyway if you’re just using LLMs! All you have to do is say “use 98.css” and you have a 10/10 site
The core problem is that some people don't even seem to notice / care.
WRT speculative decode, basically zero finetunes keep it. I'm testing with some ridiculous abliterated amalgamation so spec decode has been gone for most of its ancestry.
Fable recommended n-gram speculation so I'm working on that now.
ps. ngram didn't work for me very well, but dedicated speculative model works very well
ps. 2. in my case I'm just maintaining Makefile that does everything from update/upgrade (git pull/recompile) to starting server with different models, stuff like:
# over baseline at temp 0.6 (95 vs 45 tok/s), ~4x over naive layer-split baseline.
Qwen3.6-27B-MTP-UD-Q8_K_XL:
./llama.cpp/llama-server \
-hf unsloth/Qwen3.6-27B-MTP-GGUF:UD-Q8_K_XL \
--no-mmproj \
--parallel 1 \
--kv-unified \
--flash-attn on \
--fit off \
--split-mode tensor \
-ngl 99 \
--cache-type-k q8_0 \
--cache-type-v q8_0 \
--host 0.0.0.0 \
--tools all \
--jinja \
--ctx-size 262144 \
--spec-type draft-mtp \
--spec-draft-n-max 6 \
--temp 0.6 \
--top-p 0.95 \
--top-k 20 \
--min-p 0.0 \
--repeat-penalty 1.0 \
--presence-penalty 1.1 \
--threads 8 \
--reasoning-budget 2048 \
--reasoning on \
--chat-template-kwargs '{"preserve_thinking": true}' \
--reasoning-budget-message "reasoning budget consumed, time to answer now"
...
Qwen: Qwen3.6
Qwen3.6: Qwen3.6-27B
Qwen3.6-35B-A3B: Qwen3.6-35B-A3B-MTP
Qwen3.6-35B-A3B-MTP: Qwen3.6-35B-A3B-MTP-UD-Q8_K_XL
Qwen3.6-27B: Qwen3.6-27B-MTP
Qwen3.6-27B-MTP: Qwen3.6-27B-MTP-UD-Q8_K_XLTwo examples:
- https://github.com/ggml-org/llama.cpp/pull/25863 Someone's few lines change broke the native (ROCm) support for the AMD GPU inside Framework (and other integrated systems), and any rollback or proper fix is pending for almost a month. Fortunately there's workaround (switching to Vulkan rather than ROCm devices), but both the way the bug was introduced and the way it is not fixed just doesn't give much confidencen
- LM Studio is using llama.cpp internally for GGUF, they ship their own build with their closed source system as "runtimes". Their ROCm runtime does not enable the the AMD GPU inside the Framework, even thought the llama.cpp version would support it. So their runtime keeps telling me that there's no supported AMD GPU -- again, the solution is to use the GPU with the Vulkan devices. Not fixed since Jan at least https://github.com/lmstudio-ai/lmstudio-bug-tracker/issues/1...
I guess overall it's the worst runtime I've seen so far, except for all the other runtimes out there... I'm a fan, though in some cases I don't have enough knowledge, or I don't have access to fix things, and that feels like a bummer...
As much as I can tell, the ROCm version of llama.cpp would be a bit faster on prompt processing, but about the same on the token generation as Vulkan. Real life benchmarks don't seem to give any "ROCm or nothing" sort of vibes. And the difference between the performance of different models are way bigger than the difference between the llama.cpp versions (and versus different runtimes like the llama.cpp/GGUF and the MLX runtimes on Mac for the same models)...
I've tinkered enough with the serving, that I'd rather do something with them with, say 10% slower speed, than spending hours on seting things up again... YMMV
Also if I may ask, what does the rest of your stack look like (agent, harness etc)?
Yes ideally there would be testing every hardware + software combo but this costs engineering time and $$$ money, and you are running on master branch, no master branch of any software is stable, inherently, if you run into issues, just stick to the old hash where stuff worked, why are you insistent on both being at the bleeding edge and experience 0 breakage!
The second I didn't say it's any of llama.cpp's "fault", but it is _related_ to llama.cpp since it's being shipped in another system, aye?
Can't stick to the old hash either, because older version have different bugs. E.g. on older versions the same Qwen3.6 model reliably fails to call specific tools due to template issues, while just having the newer llama.cpp version has that fixed. So different versions - different bugs, rather than no bugs.
Why the beating you are trying to gimme, mate? :)
Sorry if I was too harsh, it’s just that my perception watching the repo has been that the llama.cpp devs are by far the most cautious and slow moving of all the inference implementations, so I found your perspective a bit surprising, I do think that the desire for stable software that never break, and software that supports the latest models and devices/device API’s are conflicting, nothing will do both, and I think that llama.cpp devs do a good job of balancing between shipping features and not breaking users.
It's a shame coz it's not even really what we want, we would obviously all be better served if we could use Vulkan or something. But I guess it's inevitable that a generic framework lags behind here.
If I was AMD I'd hire a whole ecosystem team to sit next to the ROCm people and just support big users like llama.cpp to work better on their HW, e.g. giving OSS maintainers access to their board farms. Maybe they have already done that, in which case I guess I should say I'd double the size of that team.
A team from AMD maintains Lemonade, another all-in-one setp with convenient installers for setting everything up: https://lemonade-server.ai/
These are probably better than running against llama.cpp ROCm directly as there are frequent/constant regressions on the main branch, especially for gfx1151 (Strix Halo), but RDNA in general.
There are number of AMD-focused llama.cpp forks (nathanw1014, charlie12345, ciru-ai, justinappler, etc) - as well as a few alternatives like hipfire or my hipEngine. While ROCm has gotten a lot better, one of the things I've found after writing an inference engine that has completely custom tuned/fused C++/HIP kernels, is that while it's been pretty straightforward to match/beat llama.cpp ROCm performance, that Vulkan RADV has been a lot harder since RDNA3 support for ROCm has a few issues that make it underperform ACO on some common operations on both gfx1100 and gfx1151 (see: https://github.com/ROCm/ROCm/issues/6409 )
In general, for anyone just looking to run LLM models on an AMD card, I'd just recommend going with llama.cpp Vulkan and skipping ROCm completely.
Git clone llama.cpp and build it, it's not hard.
https://github.com/ggml-org/llama.cpp/blob/master/docs/build...
literally just a few steps for the basics:
git clone https://github.com/ggml-org/llama.cpp
cmake -B build
cmake --build build --config Release
Yeah, 100% and it's becoming more and more of a thing, see rust install for example.
OTOH, if you're installing llama.cpp, you're more than likely planning to run an LLM on your Linux box with an agentic harness, so a curl into bash thing might be the least of your security concerns, :-)
The harness gets the openai-compatible endpoint fed into it to talk to llama-server across the network, but the VM has no access whatsoever to my personal files, mail, backups/deep storage, fileserver, Documents folder, etc.
Yeah I recently tried the coding harness that's recommended here, Pi, in a bubble wrap sandbox and was horrified to learn that it spams multiple warnings at you if you don't give it write access to its own config/extension folder... Everyone else is rawdogging it I guess.
How is it different than trusting any other method of installation? If URL has https and is from an author you trust i dont see the difference.
One thing I do not do as a matter of practice is install things with a ridiculous number of recursive npm dependencies.
What do you mean with this?
curl|sh is convenient for container images I guess.
I mean, sure, if there's people who can't figure that out, they're probably better off using a GUI that is a wrapper on top of somebody else's precompiled llama-server, like unsloth studio or lm studio. There's a good sized market for that and I wish them well.
This, "just run these three commands" attitude is exactly the reason why these curl|sh "installers" have become popular.
Point Claude Code at a repository and ask how to install it safely. You don’t have to know about make or cryptography of HTTPS or anything, really. It will walk you through the options and risk.
If you have questions about any part of it—i.e. you don’t recognize an acronym or deeply understand why something works—you can ask.
Or ask here! HN is filled with smart humans.
Those aren't the same assertions.
It's not a sensible assumption that a process must necessarily be difficult or complex just because you don't already know how to do it. There are an unenumerable number of tasks each of us don't know how to do and have never done before which are not difficult at all.
You really don't need to know much at all in order to compile llama.cpp. You can literally ask any half decent AI agent to do it for you or give you a step by step guide, and troubleshoot any error you get.
Just about everyone I associate with personally is polite, kind, and friendly, and everyone I work with professionally is too serious of a technologist to dismiss the most transformative technology of our lifetimes to date.
I know some people who'd oppose datacenters going up in their neighborhood or town, but nothing as bizarre or irrational as disowning someone else for merely using AI.
Cloning a repo and building it is not _that_ hard, but easy installation is often the thing that makes or breaks a product. I believe Ollama proves that point in this context.
But yes, still trusting the project with arbitrary code execution on your machine, including build formulas that pull stuff from the internet and suffer from all the above anyways
https://github.com/ggml-org/llama.cpp/releases
No need to compile unless you really need to.
brew install --cask llama-app
https://formulae.brew.sh/cask/llama-app(I still deeply distrust curlpipes in general though.)
what model was it that you were able to run with the rtx 3070?
He made such a big fuss about ollama implementing their own kernels and felt slighted about the online comments saying ollama didn't properly credit llama.cpp and it kind of left a bad taste in the mouth among the local inference community.
For me personally, it was this that made me avoid them at all costs: https://github.com/ollama/ollama/issues/11714#issuecomment-3...
llama.app is just an URL (for the "advertisement" webpages of llama.cpp outside GitHub).
> that much better than Ollama
llama.cpp is the real thing, ollama was a fork that remained inferior.
The "friends don't let friends use ollama" article linked in another comment convinced me to try llama-swap. I find it easier to directly deal with gguf files. Hard to quantify, but the outputs of the LLM seem better too. Asking the same gguf the same question with the same chat harness, I subjectively find llama.cpp does better. Might be some different defaults. I haven't dug deeply.
It’s _technically_ possible to get agents running on all kinds of setups but there seems to be an (undefined) floor for useful setups.
A lot of the stories people have about getting setups running on relatively low end hardware turn out to have huge compromises or run into issues on anything but trivial cases. I’ve found it hard to find a consensus. Or maybe I just don’t like the multiple thousand dollar price tags people are suggesting…
a solid step up from there is anything that can run Deepseek V4 Flash but the hardware ask there is a bit higher
I think the best option right now, since Apple has raised prices and Mac minis are basically impossible to get your hands on, is to build your own micro-itx machine. I actually built a mini-itx machine, but it does restrict your options a bit.
The Arc series Intel GPUs are what I think make this possible. I built a machine with an Arc b50 - it runs Gemma 26b a4b qat at around 30tok/s with their MTP head and prompt processing sits at around 500 tok/s. The really beautiful thing about this setup is the entire energy envelope of this machine sits at 120w at full load - when idle, it's at 40w and i've done some work in ubuntu to basically intelligently hibernate, which drops it to 0 watts when not in use. You can use a raspberry pi and Wake on Lan to wake the machine up for a overall draw of around 5 watts when not in use.
All in all this machine cost me 1.4k to build - but if you used micro-itx instead of mini-itx parts you could do it for under 1k - it has just 16gb of ddr5 but you don't really need more if you use models that can fit in vram.
I think it's pretty incredible that you can run an actually useful coding agent on a machine with a power envelope that is less than an incandescent light bulb. If you go up to micro-itx you can do even large cards like an intel b60 with 24gb or a b70 with 32gb and run even more powerful models. For all of these intel GPU's you'll want to compile the latest llama.cpp version with SYCL support - they are getting speedups every day, so worth staying on the edge.
Official repo, also has documentation how to configure server parameters:
https://github.com/ggml-org/Llama-macOS
Small tip, install llama.cpp with brew before llama.app, which will pick up the existing llama.cpp. That way it's easier to stay up to date with llama.cpp, since llama.app is on a slower release cadence.
Also, models installed with the hugging face CLI (hf) are picked up by llama.app automatically. The CLI will keep the model cache updated, e.g. when models get updated.
Llama.cpp became part of Huggingface recently.
The blog on the DLLM development [2].
[1] DLLM:
https://github.com/DannyArends/DLLM
[2] Teaching an AI to Know Itself: Building a Local LLM Agent in D:
https://blog.dlang.org/2026/06/07/teaching-an-ai-to-know-its...
curl -LsSf https://llama.app/install.sh | sh
and then llama serve -hf unsloth/Qwen3-4B-GGUF:Q4_0
Then I get: W load: control-looking token: 128247 '</s>' was not control-type; this is probably a bug in the model. its type will be overridden
Terminated
And the web interface says Server unavailable
Maybe it gets killed by the OS because it uses too much RAM?When I try
llama serve -hf unsloth/Llama-3.2-1B-Instruct-GGUF:Q4_K_M
It seems to work. Nice.I had a great experience with llama-cpp with Nvidia backend on NixOS.
(Sorry for being that guy.)
https://github.com/cptskippy/battlemage-llm-gateway
It's designed so that you can re-run the scripts to pull the latest updates. When Muse Glimmer was released the other day I just ran the 02 script to build the latest version of llama.cpp with support for it.
I started with "llama-server" and custom stuff around it, which is great for single model setups.. but for multi-model harness with quick switching, llama-cpp-python is peak
Must be tough not to be able to monitor your own models!
(The odds that that tagline was AI-generated seem high.)
This post (https://news.ycombinator.com/item?id=35100086) from march 2023 says in the title "Llama.cpp: Port of Facebook's LLaMA model in C/C++"
> Visit https://llama.app and follow the instructions
It's linked at the start of the README.
I can understand the desire for the llama.cpp project to want to own the end user relationship, it is true that previous to this they were a tool provider and not really owning the end user experience.