I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
I do recall a couple of decades ago, when the Turing test was discussed as the big goal that seemed so far away. Then LLMs arguably did pass the test, and no one cared about the test anymore.
If someone sat me down today with an LLM and a human and both were trying to prove to me they were human, and I can have conversations of arbitrary length, I’d get it right every time.
We might eventually regret exposing the general population to such a new technology without almost any safeguards.
So the first thing they’d do is tell me LLMs exist and the other thing is an LLM. Obviously a true human level ai could explain that away as a fabrication to trick me. I don’t think an LLM could do even this!
Turings whole point was that through the medium of text along if the human and machine were indistinguishable then that was true intelligence. So yes conversations of arbitrary length are allowed (needed).
https://arxiv.org/abs/2503.23674
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
And people forget that sometimes humans message twice. An LLM can only respond. So it immediately fails here in a true Turing test. (You could loop the LLM but then I expect even more immediately obvious bot behaviour).
Are you serious? From the paper:
> We recruited 126 participants from the UCSD psychology undergraduate subject pool and 158 participants from Prolific (Prolific, 2025).
Each human participated in 8 rounds.
> time bound
The time bound of 5 minutes was suggested by Turing himself in his original paper.
> not reproduced
It was reproduced across two populations within the paper.
> And look at their example conversations
This is irrelevant.
And it’s all irrelevant. If one human on earth can consistently get it right then it hasn’t been passed since clearly that human can somehow determine between them (whereas no one would ever be able to determine between a true “human intelligence” by definition).
And as it stands almost everyone could tell between them when allowed to discuss whatever they want for any length of time.
The fact these researchers have to keep adding bounds shows it hasn’t been passed. If we are arguing over technicalities maybe it isn’t as obviously intelligent as claimed!
Right, the test duration was left unspecified. This means any duration is acceptable. Including, for example, the only duration actually mentioned by Turing himself in his paper. Or do you have a more authoritative source on which durations are acceptable?
> If one human on earth can consistently get it right then it hasn’t been passed
Says who? Not Turing. Probably he didn't say that because it would make the test both impractical and overly conservative.
> The fact these researchers have to keep adding bounds
What "bounds"?
The speed of the goalposts here is just amazing.
I'd figure out that it's an LLM because it's effectively superhuman. Taking that away I'm not so sure I'd be able to tell
If these things have consciousness then we are committing sadism on a massive scale.