Turing Syndrome (Like the Kessler syndrome) is when the amount of AI generated data on the internet surpasses human generated content to the extent that it will eventually make it impossible to distinguish between the two
Turing Syndrome (Like the Kessler syndrome) is when the amount of AI generated data on the internet surpasses human generated content to the extent that it will eventually make it impossible to distinguish between the two
Tangentially related, Mechanical Turk and its various less well known clones which come down to the exact same thing, are increasingly the status quo for social science studies. They keep making really shocking discoveries with like 99.99999% statistical certainty given the huge sample sizes these services enable. Kind of weird how often they fail to replicate.
That they can’t replicate indicates that the sampling population isn’t consistent representation of population demographics and so there’s no real signal there that can be applied to a general population like most failure to replicate studies
I saw one the other day from a professional polling firm in which exactly 7% of "the public" said they had attended a protest, no matter what the protest was about. They asked about six or seven things people might protest about, and for every single one, 7% of the panel said they'd been on such a protest. Taken at face value it led to nonsense like concluding millions of people had attended protest marches about CDBC and 15 minute cities.
Unfortunately, the panels these firms use are extremely unrepresentative of the public but are advertised as perfect proxies. People then accept this claim unthinkingly. It's a problem.
Edit: Anathem actually came out 15 years ago
I remember similar notions in his novel Fall — where the internet is so full of fake news and sponsored content that people need additional services to filter out useful information
> Early in the Reticulum—thousands of years ago—it became almost useless because it was cluttered with faulty, obsolete, or downright misleading information,’ Sammann said.
> ‘Crap, you once called it,’ I reminded him.
> ‘Yes—a technical term. So crap filtering became important. Businesses were built around it. Some of those businesses came up with a clever plan to make more money: they poisoned the well. They began to put crap on the Reticulum deliberately, forcing people to use their products to filter that crap back out. They created syndevs whose sole purpose was to spew crap into the Reticulum. But it had to be good crap.’
> ‘What is good crap?’ Arsibalt asked in a politely incredulous tone.
> ‘Well, bad crap would be an unformatted document consisting of random letters. Good crap would be a beautifully typeset, well-written document that contained a hundred correct, verifiable sentences and one that was subtly false. It’s a lot harder to generate good crap. At first they had to hire humans to churn it out. They mostly did it by taking legitimate documents and inserting errors—swapping one name for another, say. But it didn’t really take off until the military got interested.’
> ‘As a tactic for planting misinformation in the enemy’s reticules, you mean,’ Osa said. ‘This I know about. You are referring to the Artificial Inanity programs of the mid-First Millennium…’”
The reticulum (aka internet) is filled with software generated crap. Initially most of it is blatant and easy to filter out, like an over the top "but cheap v14gra" spam email.
Eventually the companies that make the crap filtration software get into an arms race with each other, they realize if they can generate their own crap that their competitors don't detect, it'll give them a competitive advantage. The reticulum becomes filled with "high quality" crap, eg text that's ALMOST correct but wrong in subtle ways. Imagine a wiki article about pi where everything is correct except the 9th digit, or an op-ed with slightly flawed logic.
Eventually crap goes beyond text, and the reticulum starts to see deepfaked images and videos that parallel news articles. One of the jobs of the ITA is to try and find the signal in all the noise.
I like the phrase, I have been wondering about this problem. Future web crawls are going to contain ever increasing amounts of gpt generated content.
A similar problem is who is going to use stack overflow when an llm can do a better job for simple problems?
It’s on the internet somewhere from me at some point previously
As the AI generated content becomes dramatically overwhelming in scale, the human content will become increasingly easy to spot (and there will be multiple cues to the human content that make it fairly obvious). There will be a crossing of the two along the way, in regards to the amount of content generated, a relatively brief time where it will be difficult to tell which is which.
The more AI content there is, and the more it advances, the easier it's going to be to play spot the human. The time in which it'll be hardest to tell them apart, will be in the middle frames rather than in the later stages.
I don't understand why spot-the-human will become easier otherwise.
As in you can’t sort it out…just like so far there’s nobody that really knows how to clean up space long term
I feel like you're not accounting for the amount of that AI content that will be deliberately and intelligently intended to masquerade as human. Every signal you can think of and may start using is a signal that intelligent humans can and will forge. And they'll get the ones you didn't think of either, because that's their job.
I don't even have to hedge or qualify this prediction, because we have decades of experience with people already forging every signal that is technically possible to forge, and pouring very substantial effort into doing so, right up to having dark businesses that provide these things as a service. "Forge all the signals that are technically possible to forge" is a product you can buy right now. For example, from 13 days ago: https://news.ycombinator.com/item?id=36151140
I don't know what signals you think you're going to see, but I'll guarantee A: yes, you will indeed see them because there's a lot of degrees of competence in the use of these systems (I still periodically get spams where the sender sent out their template rather than filling it in with a mail merge), but B: those will just be the ones you notice, not the entire population.
Rather than a binary decision, I like to measure with "how much information does it take for me to detect a forgery". For instance, things like modern architecture renders can absolutely fool me as being real at standard definition, but at 4K they still struggle (too precise, too clean, even when they try to muss it up on purpose). I need a few paragraphs of GPT text to identify it conclusively, and that's just the default tone it takes. Ask it to give you some text in a specific style and I don't know how much it would take. For all I know I'm already missing things. You're probably already missing things you didn't realize.
I know they can exist but from what I can get my hands on it seems saying something offensive is a way to prove you are human for the time being.
I'm looking more long term, like, a year out, five years out, and on that time frame I think it's very reasonable to expect to see ChatGPT-levels of quality being completely commoditized. OpenAI or whatever commercial entity may by then have the next level AI, but there's definitely a threshold where the combination of LLM + human ingenuity gets to the point where there are effectively no signals of humanity left.
Moreover, I'm running on the theory that detecting LLMs is in fact not an arms race, but clearly ends in a win for LLMs. Basically, there exists text that simply looks human. It doesn't matter what level of superhuman AI you throw at it, it simply is in the space of "texts that a human was likely to generate". Once that level is attained, it won't matter that there is something that can generate superhuman text quality or something.
(Or, put another way, if your clever "ride the AI wave" startup plan is "create a startup that can detect LLM usage", I advise you strongly to wargame out your 1 year and 5 year futures for that tech, especially in the face of deliberate attack. I would not be optimistic.)