For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
Even flash reached 60.7% by step 12, and it's on step 16 now.
This is so exciting lmao.
That's misleading.
1. Having different models available is useful for A/B testing and helping improve Gemini itself.
2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.