I tried that exact model, I get about 50 tokens/sec on an 64GB M1 Max MBP.
It writes pretty good code; it does seem to second-guess itself in thinking traces and I wonder if it just needs a reasoning budget and message.
Each time I test a model quickly in LM Studio, I ask it:
- to write a tiny little wordpress "last login" tracker plugin, asking me clarifying questions first. I ask it to use old PHP, break the WP coding guidelines to use inline anonymous functions in the hooks, avoid custom SQL, I see if it can write something useful in a singleton class, and what questions it thinks to ask. Qwen 3.6 does this very well, Nemotron has done OK, though it's a little less effective at reading between the lines, maybe.
- to offer an answer for a SQL puzzle about finding max score per category on old MySQL (5.0) without using subquery/derived tables -- it did a good job, picked up the nuances in the prompt that allow a particular solution, didn't go on a tangent about how it would be nice to have window functions or use subqueries, did a tool call to check like I asked. With this puzzle, if the model doesn't offer up an index for performance optimisation, I nudge it; this time it didn't volunteer one but when prompted about performance it offered an index and a bunch of other nice solutions, and only there did it round up the options for subqueries and derived tables, which is fair game.
(It did badly fail the car wash test, though, even on repeated nudging, where it gets more and more insane, doubling down and never getting the point, whereas Muse Glimmer solved it and well, with a thinking trace that didn't particularly suggest it had been post-trained)
I need to test it in Pi or opencode. I've been trying to motivate my brain to move to pi, but this model supports a longer context window so maybe opencode's overlong system prompt is less of an issue.
Of the 30B models in the last 24 hours (!) I think I prefer working with Muse Glimmer, which is slow but very good, and writes rather well with just a hint of being a bit of a cheeky monkey. Not tested that in an agentic coding setting yet.
Interesting times.