"Multi-token prediction gives a free speedup of up to 2x on many models" - at the expense of halving prompt processing speed
edit: I recommend building recent llama.cpp from source, I've been updating about once a week, as there has been a fair amount of work related to MTP recently. If you're running a lot of tool calling on Qwen you might also benefit from one of the bugfixed chat templates like the Froggeric version.
spec-type = draft-mtp,ngram-mod
spec-draft-n-max = 4
I have not observed any effect on prompt processing, which is usually an order of magnitude faster than generation on my spark.