I wish this was true but it is not. And I am working on open source models so if anything, I would have a bias towards agreeing with you.
Frontier closed models (GPT/Claude) are gaining distance to everybody else. Even Google, once the king.
Your claim is a meme coming from benchmark results and sadly a lot of models are benchmaxxed. Llama 4, and most notably the Grok 3 drama with a lot of layoffs. And Chinese big tech... well they have some cultural issues.
"Qwen's base models live in a very exam-heavy basin - distinct from other base models like llama/gemma. Shown below are the embeddings from randomly sampled rollouts from ambiguous initial words like "The" and "A":"
https://xcancel.com/N8Programs/status/2044408755790508113
---
But thank god at least we have DeepSeek. They keep releasing good models in spite of being so seriously resource constrained. Punching well above their weight. But they are not just 6 months behind, either.
[0] US AI firms team up in bid to counter Chinese 'distillation' (Apr 7) https://finance.yahoo.com/sectors/technology/articles/us-ai-...
Case in point: North Korea, with far, far fewer resources.
Gemma 4 was a major improvement is self-hostable local models and Qwen-3.6-A34B is a beast, and runs great on an MBP (and insanely well on a 4090).
The biggest lift is combining these models with a good agent harness (personally prefer Hermes agent). But I’ve found in practice they’re really not benchmaxxing. I’ve had these agents successfully hand a few non-trivial research projects that I wouldn’t have been able to accomplish as successfully even last year.
When you add in the open-but-not local models, Kimi, GLM, Minimax, you have a lot of very nice options. For personal use anything I don’t use local models for I give to my Kimi 2.6 powered agent.
Over-promising is a very stupid thing. Nobody will value the intermediate steps. Nobody will value all the effort because they will always compare us with frontier models made with billions and we will become a running joke. So please stop.
I've got a 128GB strix halo staying warm at home, it has nothing on top models with big budget. It's good supplement to low end plans for offloading grunt work / initial triage
Thanks for suggestion tho, tool by antirez is always going to pique interest, I'll check it out when I'm finally home again
Tho says Metal / CUDA, so doesn't seem friendly to Linux AMD system
At what tps? You can run the new gemini flash or 5.3 codex spark at 1000+tps and run circles "open" models. You can't run anything useable locally without at the very least a blackwell 6000 if not two
Sure you can run qwen 3.6 at 20tps on a mac 128gb but let's not pretend this will get you anywhere