> Are the AI companies really living in an echo chamber?
The author tested 12 models, and only one was consistently wrong. More than half were correct 100% of the time.
A better conclusion would be that there’s something in particular wrong with GPT-5 Chat, all the other GPT 5 variants are OK. I wonder what’s different?