Data at https://gertlabs.com/rankings
Data at https://gertlabs.com/rankings
MiMo v2.5 is on there, as well as the pro version.
We found a few anomalies in our evaluations, which makes sense -- if every new sub-release is better across the board in every area of the model card, that should raise alarms about benchmaxxing. But the main thing we found is that hype != performance, and I trust our benchmark methodology significantly more than the model cards the labs add to their press releases.
We didn't love the results because it draws negative scrutiny to our benchmark, but the results are real and done at scale and I think DeepSeek V4 Pro's inability to do agentic work outside of environments it was trained on is an important thing to measure, especially when so many other models can generalize to new environments just fine.
Google models also struggle with tools, but they have very strong initial answers, so there is more potential for them to bridge the gap with some better post-training.
Flash handles it fine, which I found amusing. (Since Mimo is supposed to be opus level!) But Flash seems to work even better in Claude Code...
With smaller models I always have the issue of needing to adapt myself to their preferred workflow... which sort of defeats the purpose. Price is hard to beat tho :)
When it gets stuck, I get one-shot advice from Claude or DS Pro. I’ve done massive amounts of work for cheap this way.
The issue was that my previous instructions had <command> as a placeholder. But the model started wrapping bash commands in <command></command> tags... haha. Now that it has an actual example it just works properly.