I want to do a writeup on ChatGPT Image 2 but at this point I don't think people care about nuanced image generation anymore...even though ChatGPT Image 2 crushes all my existing tests.
Many of these ELO comparative tests (ArtificialAnalysis is guilty as hell on this as well) also have other problems such as a considerable number of "amateur judges" tending to prioritize aesthetics over actual instruction-following given the prompt.
Also (less a critique of Arena.AI necessarily), but the MAI models are so incredibly locked down (e.g. censored) as to be functionally useless. I have a sneaking suspicion its fallout from Tay.
This is purely about generating images with people in them, I don’t think she’s doing any logic puzzles with gotchas and specific alignments of differently colored blocks and whatnot
I was surprised the internet didn't have a meltdown when it released.
I have also have noticed that GPT Image 2 is very good and has a great cost/result ratio compared to other models, specially using the low version which can cost 0.01 and is usually good enough for many use cases. I completely replaced Nano Banana for most of my workflows because I'm using API and the cost adds up. I still haven't tried Nano Banana 2 Lite but the price may be hard to justify, although the speed bump sounds good.
NB
NB Pro
NB 2
A lot of us are still expecting Google to drop a Pro version of NB 2 but that hasn't happened yet...