Qwen3.8-27B has been the turning point for me. It's not as strong as the absolute frontier, but it's the first time I feel local coding models are actually functionally useable as daily drivers.
Man I am having a hell of a time trying to optimize 3.8 over 3.6. I don’t have a particularly powerful setup but I can usually push 20-30tok/s on 3.6 and I can barely get to 10 on 3.8. Both unsloth same VRAM/RAM distribution more or less. My 3.6 is still producing consistently better results and faster
That's interesting since both models are dense. I wonder if this is more of an optimization issue with 3.8 rather than something inherent to the architecture.
Could have sworn I read these were the same architectures the other day .... 3.6 and 3.8 at this size.
I’m also not an engineer/coder so it’s equally possible I’m just doing something wrong.
You're likely using 3.6-35B-A3B, the 3.8 is currently a 27 billion parameter dense model.
I have definitely use that to great success, and I do think it colors some of my memory here. I need to check if the 3.6 27B I was using previously was also a dense model. Good suggestion appreciate it
Maybe you're holding it wrong because it's the same architecture between the models, assuming you're using the dense 27B model in both cases. And 3.8 is a significant improvement on 3.6.