What's a better model to run on one GPU?
Gemma3 27B gives me a rapid 1shot response, and actually works really well for the type of rubber duck brainstorming partner I often need.
Now, yes, QwQ will take a lot of tokens to get there (in one case it took it over 5 minutes running on Mac Studio M1 Ultra). Nevertheless, at least it can solve it.
Translating this to code, for example, it means that Gemma is that much more likely to pretend to solve a more complicated problem that you give it by "simplifying" it to something it already knows how to solve.