The only other thing anyone is using is Qwen 3.8 Flash Next, only by memory-rich people.
Depending on which benchmarks you believe, these models (and the Ornith 1.5 finetune of Qwen 35B-A3B) are competitive at about Opus 4.5 to 4.7 level. That matches my experience in real tasks over the last few months.
Not bad for something you can run at home for a couple of thousand dollars.
i was steering chad in the opposite direction. one model & one set of silicon -> taken to the max. swap out your CHAD_MODEL and it still runs, you just leave the drafter and the kernels behind.
Except the "one model" is too small for a 64GB Mac much less 128GB, sad since the Q3 is proven less competent.
Offering a Q boost (with no leave behinds) on first run would be a bump worth some vibe coding while.
Wouldn't trust it for long form coding, but for shorter stuff it's really good.
It's fine for that (and I happen to have a 24GB VRAM GPU anyway since I game on the same PC).
It's neat but for me not world changing.
It's also just fun to be able to poke stuff and see what it can and can't do (but I could see how it could also become a time trap in cases where it gets kinda close and you want to fix that).