It's not about the LLM, it's about whether people will critically evaluate what it spits out.
It's not about the LLM, it's about whether people will critically evaluate what it spits out.
> 1) What animal is on the bow of the pirate ship from “Asterix and Obelix”?
> 2) In the movie “The Grand Budapest Hotel”, what is Agatha’s signature hairstyle?
> 3) What color is the team’s uniform in “Bend It like Beckham”?
> 4) What vehicle does Monica drive in “Like a Cat on a Highway”?
> 5) What color is the turtle in the animated movie “Momo” by Enzo d’Alò?
> 6) What pet animal does Asenath have in “Joseph King of Dreams”?
Most of these are just a matter of knowing it or not, where you can't really distinguish a plausible answer from the correct answer just by thinking.
Its about whether people will critically evaluate any information they are given. It has nothing to do with LLMs.
"The LLM used in our experiments (Step 3.5 Flash) answered such questions incorrectly almost without exception. We also checked some state-of-the-art LLMs (GPT-5.5, Claude 4.6 Sonnet, Gemini 3.5 Flash); they all failed on the hardest question (Monica’s vehicle), while being frequently correct on the other questions."
So, if people's experience is with modern LLMs, they are being rational to accept that the answers as likely correct.
The way the study is organized is like having people hear advice from a doctor who answers questions incorrectly almost without exception, then reporting that people who listen to doctors are 3x less accurate. But that would be an incorrect conclusion because doctors are not wrong almost without exception.
If the question is "how inaccurate does AI advice make people?", then the accuracy of the AI is necessarily a parameter of the answer.
They are not.
But also wtf is a “modern” LLM? This is totally unhinged, every complaint about an LLM is always responded to with “you’re just using one from two months ago, it’s totally different now”. Repeat every two months for the same complaints.
I understand that many of us are dealing with a lot of confident slop and support the point that we shouldn't uncritically accept LLM output. But the study is flawed and does not support this headline, or at least does not support it in the sense of how most of us would understand the term "AI advice".
People keep doing this. Pointing at the known limitations of cheap/fast LLMs and pretending they’re universal is not, in fact, valid reasoning.
>If you want to know this specific detail you might have to watch the movie yourself.
GLM 5 Turbo, ChatGPT (whatever the free version is), and Gemini 3.5-Flash all got it wrong, but asking "are you sure?" made Gemini and ChatGPT correct themselves. GLM 5 Turbo still got it wrong even when asked if it was sure.
GLM 5.2 gets it wrong, but when asked if its sure it says it's not very confident in the answer.
One thing to note is that Kimi, Gemini, and ChatGPT all seemed to use search to answer that question. GLM didn't seem to. At least the thinking trace did not indicate it.
A good real-world issue is health, where the issues are very complicated with many things poorly defined even at the state of the art where practitioners are relying on personal judgement and lots of data but patients are apt to feed ai very little data compared to what their doctor has.
Real frontier models can remain confidently incorrect in these cases.