It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%.
At that point I just stopped scrolling.
It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%.
At that point I just stopped scrolling.
I'm making up these figures, but the point is lower is better, or "more Human-Like". Test was specified as >50% meaning "accurately determined human vs. bot more than half the time". The site claims LLMs are now guessed correctly less than half, which is how the turing test was defined as per the site.
It makes sense, even if you disagree it's significant.
This is probably already happening within the parade of censorship systems trying to imbue the models with agency
>GPT-4 was judged to be a human 54% of the time, outperforming ELIZA (22%) but lagging behind actual humans (67%).
The incentives don't align with honesty though.