or you can just load up ollama, have it load a local model and point claude or opencode at it...
is this article old? It's not. I'm not sure why he went through all the bother of llama.cpp
is this article old? It's not. I'm not sure why he went through all the bother of llama.cpp
Original video: https://x.com/Freerunnering/status/2065275403548168398
And in the blog post there is a table showing the different speeds I got from different engines.
Slowest combo was 38.1 tk/s, and the fastest was 72.2 tk/s. All from "the same" model.
Also Ollama has other issues (like forgetting what it really is - a wrapper).