So far ollama managed to be the most performant of them all. I will get 30 to 40 tokes/sec with it when using the -mlx version of Qwen3.8.
Whatever the sauce the ollama folks baked into the mlx + MTP mix is currently working the best out of the box.
So far ollama managed to be the most performant of them all. I will get 30 to 40 tokes/sec with it when using the -mlx version of Qwen3.8.
Whatever the sauce the ollama folks baked into the mlx + MTP mix is currently working the best out of the box.
But 30-40 tokens/s would make a big difference.
I gave it a shot now:
mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit --prompt 'give me fizz buzz in rust' --enable-thinking --draft-kind mtp --draft-model mlx-community/Qwen3.8-27B-MTP-4bit --verbose
==========
Prompt: 58 tokens, 90.717 tokens-per-sec Generation: 145 tokens, 36.392 tokens-per-sec Peak memory: 17.419 GB Speculative decoding: 2.79 accepted tokens/round (1.79 accepted drafts/round, 89.4% of drafted, avg draft 2.00) over 52 rounds
Which is very close to ollama, thank you!
I'm not sure if I can get rid of the drafter model, if I understand correctly, the Qwen model already includes a built in draft headers, but just having --draft-kind mtp results in about 17 t/s.