huh.. I'm a bit of local LLM noob so I wasn't familiar with mlx_vlm.
I gave it a shot now:
mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit --prompt 'give me fizz buzz in rust' --enable-thinking --draft-kind mtp --draft-model mlx-community/Qwen3.8-27B-MTP-4bit --verbose
==========
Prompt: 58 tokens, 90.717 tokens-per-sec Generation: 145 tokens, 36.392 tokens-per-sec Peak memory: 17.419 GB Speculative decoding: 2.79 accepted tokens/round (1.79 accepted drafts/round, 89.4% of drafted, avg draft 2.00) over 52 rounds
Which is very close to ollama, thank you!
I'm not sure if I can get rid of the drafter model, if I understand correctly, the Qwen model already includes a built in draft headers, but just having --draft-kind mtp results in about 17 t/s.