Is qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious.
It is like a ping-pong game: the advantage flips back and forth between providers.
So yeah, it's the best local model I've seen. I am going to try the Qwopus 3.6 fine tune soon with the same spec and tickets and compare the output of both.
vLLM gives me ~7000+ tok/sec with Gemma 4's MoE model. Vs ~6000 tok/sec for Qwen 3.6 MoE.
Not tried it yet but I've seen tests that suggest they've properly fixed the tool calling issues.
But there’s also the quantization of DeepSeek v4 flash called dwarfstar