- A recent study tested GPT-4, GPT-3.5, and 1960’s ELIZA, to see which program best mimics human conversation in a Turing test. Participants had to guess if they were interacting with a human or an AI.
- Surprisingly, the old ELIZA program outperformed GPT-3.5. GPT-4 did better than ELIZA, but didn’t reach a 50% success rate, which effectively means worse than a coin flip.
- The authors write that the Turing test still has relevance in understanding how humans interact with AI. However, we should refrain from seeing the Turing test as a barometer of AI intelligence.