https://i.imgur.com/z83umbk.jpeg
Here I change a widely known riddle to the opposite answer, and I manage to make it state them both as the answer.
You can stump a person with a riddle or a logic puzzle or an optical illusion.
The fact is that people who say 'it is just fancy autocomplete' are using a thought-terminating cliche, and a 'I stumped an LLM' proves nothing.
It is simply regurgitating this phrase without even considering that it is stating the exact opposite of the answer it just gave, simply because most answers to this riddle on the internet say this at the end.
> This riddle plays on the assumption that a surgeon is typically male, but in this case, the surgeon is the boy's mother.
So from this 1 failing, you can see that it is a copy and paste machine, and it doesn't even understand that it is contradicting itself.
It's very much feels to be gearing towards figuring out the gist of a search engine you may be trying to complete and put together by reading a few links.
When I speak colloquially, I have an underlying idea rooted in a world model to be expressed. I don't spit out 1 word at a time based on the previous words I already said.
No, it "can't be debated," it is clearly false! You said "by definition," but you used an irrational and bigoted definition of "general reasoning and logic" which conflates such things with performance on a standardized test. Humans aren't innately good at stupid logic puzzles that LLMs might get a 71st percentile in. Our brains are not actually designed to solve decontextualized riddles. That's a specialized skill which can be practiced. It's depressing enough when people claim IQ tests are actually good measures of human intelligence, despite overwhelming evidence to the contrary. But now, by even worse reasoning, we have people saying a computer is smarter than "average humans." (MTurk average humans? Undergrads? Who cares!) The complete lack of skepticism and scientific thinking on display by many AI developers/evangelists is just plain depressing.
Let me add that a truly humiliating number of those """general reasoning""" LLM benchmarks are fucking multiple choice questions! Not all of them, but a lot. ML critics have been complaining since ~2017 (BERT) that LLMs pick up on spurious statistical correlations in benchmarks but fail badly in real-world examples that use slightly different language. Using a multiple choice test is simply dishonest, like a middle finger to scientific criticism.