Prompt results are graded based a weighted calculation which includes: adherence to the prompt, image fidelity, and steerability.
My comparison benchmark also tends to favor prompt adherence, which a lot of others don’t. Most of ones that I've seen tend towards rather simplistic prompts (e.g. "neon-lit city facing a robotic uprising, with high-tech battles, in anime style"), whereas the prompts I've created try to test high specificity.
I’ve been running them all the way back to SDXL.
You can compare specific models using the "View All Models" so if you want to see the progression of open-weight models, or model X vs model Y, you can do so.
Just a heads up - I haven't added Flux 3 as I'm waiting until BFL drops the open-weights version.
Generative Comparisons:
https://genai-showdown.specr.net
Editing Comparisons: