HNHacker News
TopNewBestAskShowJobs

wowitsbase

3 karma · joined September 7, 2026

submissionscomments
wowitsbase··on [dead]
Mainly done through aggressive automatic cuda graph optimizations, but it also has cool custom formats, allows for huggingface pulls from cli, better quantization control, and a lot of other small things that make it easier to use than ollama.

What do you think?

wowitsbase··on I promise this Ollama replacement will give you the fastest t/s you've ever had
Somewhat, but most of the optimizations are cuda based.

I would guess it would be a little faster, like 20-40%, but not 200% like on Nvidia.

wowitsbase··on I promise this Ollama replacement will give you the fastest t/s you've ever had
A small developer team and I had worked on this for a while for personal reasons, so I have decided to port it to cpp and release it to the general public. The stats are in the github. I promise if you have a gpu this will improve your speeds by at least 50%, even if you believe your configuration is optimized.
wowitsbase··on Ollama replacement 2-4x faster for no extra compute cost
Speculative decoding was on in all of my tests, the gain came from moving the round's inputs off host memory.
wowitsbase··on Ollama replacement 2-4x faster for no extra compute cost
agreed, if not for this project I've been making I would at least be using base llama.cpp

Personally I think it comes down to simplicity, but there's no reason for it's performance drops compared to llama.cpp while it's a wrapper of it.