As I’ve said before, I wouldn’t put a lot of stock in Arena’s scoring system. They have MAI Image 2.5 ranked above Gemini Nano Banana Pro, and maybe that’s true on paper but good luck using it. Microsoft’s censorship makes Google feel like the wild west by comparison.
They also have Meta’s Muse Image ranked above NB Pro, which is just patently absurd. In my own GenAI benchmark it only managed a lackluster 7 passes out of 15. Even the open‑weight Ideogram 4 scored higher than that.
For reference, here’s GPT‑Image‑2, NB Pro, and Muse compared: