Aww, I don’t like the new pelican benchmark as much. I liked that the old prompt was vague and we could see how the AI interpreted it.
I think a more challenging, well, challenge, would be to offer an even more absurd scenario and see how the model handles it.
Example: generate an svg of a pelican and a mongoose eating popcorn inside a pyramid-shaped vehicle flying around Jupiter. Result: https://imgur.com/a/TBGYChc
I was inspired by Max Woolf's nano banana test prompts: https://minimaxir.com/2025/11/nano-banana-prompts/
Do you think it would be reasonable to include both in future reviews, at least for the sake of back-compatibility (and comparability)?