I think that definition is wide enough to include human intelligence, so their finding should be equally valid for humans.
I think that definition is wide enough to include human intelligence, so their finding should be equally valid for humans.
Which is definitely true. Human memory and the ability to correctly recall things we though we remembered is affected by a whole bunch of things and at times very unreliable.
However, human intelligence, unlike LLMs, is not limited to recalling information we once learned. We are also able to do logical reasoning, which seems to improve in LLMs, but is far from being perfect.
Another problem is how different we treat the reliability of information depending on the source, especially based on personal bias. I think that is a huge factor, because in my experience, LLMs tend to quickly fall over and change their opinion based on user input.
Baseball and bat together cost $1.10, the bat is $1 more than the ball, how much does the ball cost?
A French plane filled with Spanish passengers crashes over Italy, where are the survivors buried?
An armed man enters a store, tells the cashier to hand over the money, and when he departs the cashier calls the police. Was this a robbery?
My final example demonstrates how those cultural norms cause errors, it was from a logical thinking session at university, where none of the rest of my group could accept my (correct) claim that the answer was "not enough information to answer" even when I gave a (different but also plausible) non-robbery scenario and pointed out that we were in a logical thinking training session which would have trick questions.
My dad had a similar anecdote about not being able to convince others of the true right answer, but his training session had the setup "you crash landed on the moon, here's a list of stuff in your pod, make an ordered list of what you take with you to reach a survival station", and the correct answer was 1. oxygen tanks, 2. a rowing boat, 3. everything else, because the boat is a convenient container for everything else and you can drag it along the surface even though there's no water.
No idea what you're getting at here, though.
It's true that this is often not a big deal, but which times it is and which times it is not is not known (which itself is typically not known, once again because of the convention).
Talking about the phenomenon is also contrary to conventions, and typically extremely well enforced (as I imagine you noticed during the dispute with your incorrect classmates, or else you were smart enough to not push the issue).
This one single causal phenomenon underlies everything, yet we ~refuse[1] to examine it.
[1] Here I am kind of being hypocritical, in that I assume to some degree that humans have the base capability in the first place.
This is effectively like coming up with an algorithm and then executing it. So how good/bad are these LLMs if you asked them to generate say a LUA script to compute the answer, ala counting occurrences problem mentioned in a different comment, and then pass that off to a LUA interpreter to get the answer?
I think this is a sensible approach in some problem domains with software development being a particularly good example. But I think this approach quickly falls apart as soon as your „definitely right answer“ involves real world interaction.
And if one thinks about it, most of the value any company derives comes down to some sort of real world interaction, wether directly or by proxy.
What do cows drink?