I agree. In fact the entire reason I initially built GenAI Showdown was because many of the comparative tests on places like Image Arena Leaderboard [1] are not designed to challenge models on prompt adherence. Even when they are, a considerable number of amateur judges tend to prioritize aesthetics over adherence or instruction-following.
I'll likely be redoing that particular bench with added minimum passing criteria of an anvil.
[1] - https://huggingface.co/spaces/ArtificialAnalysis/Text-to-Ima...