But that's exactly how real world works too.
But that's exactly how real world works too.
You'd get the answer to a riddle wrong or miss something and nobody would start assuming that means you lack a fundamental understanding of how the world works. There's entire fields that look into how and why we make various mistakes and riddles and puzzles designed specifically to trip people up.
If you want to test if these models can solve riddles, or where they make mistakes go right ahead that's great. It's any assumption it has a much deeper meaning that is wrong to me.
That's because with humans we assume a certain level of competency and intellectual ability. We cannot make the same assumption when testing AI systems like LLMs because their level of competency and intellectual ability is exactly the question we are trying to answer in the first place.
Note that getting an answer a little wrong, because the question looks like a question you already know the answer to, can be catastrophic in real world conditions. Tipping a frying pan over a plate on a table to serve an omelette when you've learned to do the same thing to serve a cooked shrimp works just fine and shows everyone how smart you are and how well you generalise to novel situations, right up to the point where the contents of the frying pan are on fire and you still tip them over a plate, on a table. Made of flammable wood. Oops.
Also note: a human may be confused by the Tsathoggua-Cthuga-Cxaxukluth river-crossing riddle but they'd never be confused about the danger of a frying pan on fire.
Which highlights the problem with using these riddles to assess other capabilities.
“What’s the correct way to delete an element from a Java array while iterating over it?”
“Well I suppose you’d need to avoid invalidating the iterator state but I don’t recall the…”
“BZZT!”
And btw, some arrays are fixed, but many languages call something an “array” which is dynamically-sized.