I want to do a writeup on ChatGPT Image 2 but at this point I don't think people care about nuanced image generation anymore...even though ChatGPT Image 2 crushes all my existing tests.
I want to do a writeup on ChatGPT Image 2 but at this point I don't think people care about nuanced image generation anymore...even though ChatGPT Image 2 crushes all my existing tests.
This is purely about generating images with people in them, I don’t think she’s doing any logic puzzles with gotchas and specific alignments of differently colored blocks and whatnot
Many of these ELO comparative tests (ArtificialAnalysis is guilty as hell on this as well) also have other problems such as a considerable number of "amateur judges" tending to prioritize aesthetics over actual instruction-following given the prompt.
Also (less a critique of Arena.AI necessarily), but the MAI models are so incredibly locked down (e.g. censored) as to be functionally useless. I have a sneaking suspicion its fallout from Tay.
I have also have noticed that GPT Image 2 is very good and has a great cost/result ratio compared to other models, specially using the low version which can cost 0.01 and is usually good enough for many use cases. I completely replaced Nano Banana for most of my workflows because I'm using API and the cost adds up. I still haven't tried Nano Banana 2 Lite but the price may be hard to justify, although the speed bump sounds good.
I was surprised the internet didn't have a meltdown when it released.