On our benchmarks we approach 2x decode speeds on a variety of Mac hardware (tested most on M4 Pro and Max).
llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.