You're probably partly right (though I use the same context size for all models, so there still is a difference in the model/quantization itself), but I definitely see it "thinking itself into an extremely expensive and error-prone approach without backing out". A common "failure mode" of my benchmark with <Q6 is "okay, I will implement a simulator", and when I see that I think to myself "please, no, you don't need to do that at all, you just missed a tiny thing" (but if I would type that into the session, it would void the benchmark; the idea is that I wouldn't know the solution myself after all).
Especially because Qwen3.8 27B does not seem to be good enough to implement the simulator with all its intricacies and tiny subtleties correctly. So it's churning for hours with no real progress, where it just had to explore the initial problem space a tiny bit more.
And very interestingly, even with Q6, turning off reasoning entirely does not make it fall into that trap, or many others, and it more often solves the task in record time. Both because of what I just described, but also because the endless reasoning costs a lot of token, which significantly translates to time on a home rig...