Anyone compare to ollama? I had good success with latest ollama with ROCm 7.4 on 9070 XT a few days ago
Under the hood they are both running llama.cpp, but this has specific builds for different GPUs. Not sure if the 9070 is one, I am running it on a 370 and 395 APU.
Model: qwen3.59b Prompt: "Hey, tell me a story about going to space"
Ollama completed in about 1:44 Lemonade completed in about 1:14
So it seems faster in this very limited test.
Thanks for that data point I should experiment with ROCm