I think this is my favourite test. You can just tell it was programmed on smug Reddit comments talking about how Americans drive to places 50 metres away.
I'm not trying to trick it, so falling for tricks is harmless for my use cases. Does it write quality, secure code? Does it give me accurate answers about coding/physics/biology. If it gets those wrong, that's a problem. If it fails to solve riddles, well, that'll be a problem iff I decide to build a riddle solver using it.
LLMs are largely textual creatures and they fail to see things that are there or imagine things that are under certain textual patterns.
I don't think you would say a human "isn't really intelligent" because it imagines grey spots at the intersection of black squares on a white background even though they aren't there.