3.8 is absolutely quite a bit better than 3.5. My go-to quick benchmark is a bit like yours, implementing a simple but complete game in pygame. Most small models output code that's too buggy to be fixable. Qwen 3.8 wrote buggy code too, but the bugs were minor, and the game legitimately playable and fun.
That said, I've not found it very reliable, especially at the low quant required to run on my hardware. It's very capable for its size, enough to be legitimately useful, but it's never certain whether it will hit its capability peak on a given attempt, and sometimes you have to give it multiple tries.