Mainly done through aggressive automatic cuda graph optimizations, but it also has cool custom formats, allows for huggingface pulls from cli, better quantization control, and a lot of other small things that make it easier to use than ollama.
What do you think?