I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing. I really can't understand how this gap persists; but then, there seem to have been at least some people who couldn't sniff out ELIZA, back in the day, too.
That's mostly true for longer LLM output with all the sycophancy / LinkedIn bias thrown in.
Make it casual conversation or comments, and give it instructions on appearing casual, or even better kill the censoring and fixed-prompt (with an open model), and it's orders of magnitude more difficult, unless if you suspect it and try specifically tailored prompts to sniff it.
There's no shortage of people obliviously discussing with AI bots in comment sections.
People, many of them at least, cannot make this distinction anymore. You can, I can, but people as a whole are having problems with that.
Even this (assuming it's even true) will likely not be true in some near-term future.
>continuously claiming this or that is a bot
I see it as a contemporary form of religious thinking. Like (say) pilgrims seeing blood on a statue of the virgin, plenty of people are now seeing the hand of AI in everything they read. If you want to see something hard enough, it tends to become magically visible.
Most users aren't very critical of the output. They just want a sycophantic ear, and 4o was perfect for that task. It's not _good_ but there is high demand for it.
From the abstract: "When prompted to adopt a humanlike persona, GPT-4.5 was judged to be the human 73% of the time: significantly more often than interrogators selected the real human participant. LLaMa-3.1, with the same prompt, was judged to be the human 56% of the time"
That's because they're trained that way. If you trained a modern frontier LLM with the explicit goal of passing the Turing test, it would have no difficulty doing so.
To me it just means they can brilliantly fake human conversation - the original design goal of Large Language Models.
It's really easy to tell if you're talking to an LLM if you ask a question that requires actually knowing things, not going for the first search result of a tool call or whatever most popular answer was embedded in the weights.
For this reason even the most sophisticated models still require system prompts, skills and all that other crap.
That's exactly the criteria we used to assume for over half a century for it finally being intelligent: the Turing Test.
And what does "brilliantly fake human conversation" even mean if not some kind of intelligence? It's like saying "He is not good at math, he just brilliantly proves theorems".
You can recognize 70% of those (true positive rate) and still have a false negative rate of 30%, while thinking you got 100% of the AI ones!
The problem is that you'd be oblivious to those you don't recognize.
Great times!