It's not overthinking, it's the right amount of thinking necessary for such a small model to get good results. The dumber the model, the more it has to think to be smart. There's no easy way to reduce the thinking without reducing the model quality.
I do like the output from qwen when I get it, but honestly I haven't been impressed enough with it to put up with the downsides.