LLMs Are This Close to Destroying the Internet
boehs.org
boehs.org
The real issue isn't that LLMs make it much easier to generate garbage content, it's that the filter mechanisms we have used for two decades are all failing, and have been massively degrading for a decade now. Just like google is dominated by SEO spam, most of the platforms I mentioned above are overrun by bots and click farms. Upvotes, hearts and reshares can be bought for pennies. Or just buy a preprogrammed bot that posts "subtle" hints to your Onlyfans everywhere.
Of course on an individual level there are strategies to avoid that content. And those won't really be changed by LLMs
(Not trying to be snarky or pedantic, "last rights" threw me for a loop until I figured it out.)
I don’t want to read clickbait. Whenever I see clickbait, I click off. That means for half my searches I need to kludge through many, many results to find something by a human who actually cares
AI producing and consuming its own output, and the resulting "model collapse", with human beings increasingly out of the loop in thus game of telephone.
I've noticed some "summaries" of content here on HN, posted as comments, with the telltale sign of LLMs. Adding no value whatsoever. Could some day be the case that a large number of posts even here are of AI bots interacting with each other?
I'm not sure about the "rebels" at the end of the article, how will they find each other in sufficient numbers, how will word of mouth spread in sufficient numbers?
I think the next step is actually a big leap backwards - human-curated knowledge bases, islands in the sea of AI-generated noise.
Also, the LLMs aren't necessarily doomed. Presumably, the humans running them are looking out for the problems of inbreeding or LLMs-eating-their-own-excrement, and will implement some combination of not feeding that crap, and/or freezing the model at a peak point where the results are still primarily based on human creations. Perhaps ChatGPT 4.3 will be the peak?
EDIT: typos
In that case the cure will be worse than the disease, in my opinion. I don't want a siloed internet. Silos exist and are bad enough now, imagine if in the future the only good, human-produced content is secreted away within them.
I think this will be a radical change for the better. Once the baseline quality and trust drops low enough, there is as an incentive for curation, walled gardens, and reputation based networks. I hope for the return of webrings, moderated forums, and quality standards for participation.
I also think it may finally break the hold of social media obsession. Outrage loses much of its luster when people realize the other end of it is likely a bot. This is a step toward asking "who is this person, and why do I care what they think."
I suspect many of our current societal ills are caused by too much information spreading too quickly, not forced to go through a multitude of brains or stand the tests of time.
Friends of friends networks, personal recommendations sent over email, communicators, Discord.
And even saying that sounds absurd...because it is absurd. We find ways to get new insights from essentially the same human minds from thousands of years ago interacting with each other to produce data.
Why would sufficiently advanced LLMs be any different?
1. Humans get input from real life, not just other humans. An AI can’t go outside and smell the flowers. As the % of AI input from other AI approaches 100% they will be cut off.
2. Current LLMs are not “sufficiently advanced”.
They even talked about #1 in the article, if you had bothered to read it
https://en.wikipedia.org/wiki/Georgian_numerals#Numeric_valu...
Practically speaking, it's arguably impossible to create a new internet, i.e., the physical infrastructure for one.
But it's possible to create alternatives to "the" web. And nothing requires alternatives to be anything like the web that LLMs are trained on.
As in, a page should only rank well if the page is linked to from websites that have good reputations.
So it's unlikely reputable news websites are going to link to and boost the search rank of sites full of fake articles and fake images, and any news website that did would quickly start to lose reputation/trust and their own ranking.
Before AI, the web was already full of unlimited junk, spun articles, clickbait and spam, with backlink farms trying to promote it. What's different here?
So the problem is not being able to detect link pyramids? Where is the link pyramid getting its PageRank reputation from and can't you ban the lot of them once the manipulation is discovered?
It's not like e.g. the BBC are going to start linking to AI generated articles without checking their accuracy. And if they did, the BBC would risk their PageRank influence and their own ranking.
I get backlink farms can work right now and Google doesn't strictly follow PageRank, but I'm asking if the PageRank concept used in the right way would help.
There has long been a similar situation for song lyrics. It doesn’t matter if you get a song’s lyrics exactly right. If you sing it in a little different way, that’s fine.
I’m also reminded of the situation with folk music where there are many variations on any given tune, where the tune came from is lost in history, and it’s fine.
This is how cultural evolution works. The remixes used to be done by people and now they’re increasingly done by machine. LLM’s are new, but Wikipedia is a remix, too, and so are social media and forums like this one.
When you really care about getting it right, you need to go into research mode. Follow the citations and find the primary sources. Stack Overflow is a useful source of hints, but also read the documentation, the source code, and do your own testing.
It’s more work. Most of the time we don’t need to do it because the hints are good enough, and when they aren’t, we can recognize that.
Search engines are still quite useful for research when you find the right keywords. You might need to switch to a more specialized search engine, though? They’re still useful even if they don’t have the enormous cultural impact of the default choice.
That is not dead which can eternal spam,
And with strange aeons death gives not a damn.1. The damage in the form of the destruction of news sources and replacement of writers with AI has happened and continues to happen. Writers are already jobless. News sources are already dead. But sure, these jobless people with no hope for their industry are "doomers". The article doesn't mention artists, but they are also getting hammered.
2. The worst players already have agency over their data choices, they are using that to build silos and destroy their competitors who don't have that ability.
3. Some social places have already become unusable. Twitter for instance. Threads may be able to push back on this, but they are already a dangerous silo.
Due to the volume of half-arsed news the current AIs can generate your average reader might not notice that some is missing, though I questions that. People are starting to notice that their newspapers doesn't actually have news anymore (and that's just do to cost optimization and competition from ad supported online news).
I fear that we're entering a world where some of us pay for news written by real journalists, while the masses consume garbage "news" which is more tailored to them clicking ads, rather that learning about the world.
We’re all copying other people’s homework, especially in social media. How much of what you know about the world comes from personal observation? Most people haven’t traveled to most places.
The news business are almost entirely an ad business now. Only a few sites do real investigative journalism and only with a few journalists.
Many news businesses are now owned or majority controlled by billionaires and have the expected editorial biases that also don't really support hard hitting journalism.
2. Open source data sets can have curation too, this seems like a silly strawman to make.
3. Twitter isn't just unusable because of AI. It was never more than barely useable to begin with, and many of users didn't help the situation.
I got it from your original post: you don't give a shit. I was just calling you out on it. At least you own it.
> 2. Open source data sets can have curation too, this seems like a silly strawman to make.
It's not a strawman, but "curated open source data sets" are a red herring: those data sets are not what is going to be controlling the online experience of the vast majority of those online.
News is going to break, and people who break it accurately with personality and a unique take are going to do well regardless if whether newspapers and other bastions of old world journalism continue to exist.
From a reader’s point of view, it seems like there’s plenty to read.
If Big Tech can't even catch pornography or explicit uses of the n-word, there's no way they'll be able to filter subtle LLM / art generator hallucinations.
Obviously generalizing over pictures of humans and pictures of porn will enable CSAM generation ...
Plus a LOT of things have been declared illegal pornographic material in various places: all porn (China), drawn CSAM, any kind of drawn porn with unclear ages, ...
If someone is going to suggest that there is "illegal child pornography in DALL-E's training data" I think it needs to be backed up - that's a pretty big accusation.
Your attitude is naive and competely at odds with (honest) AI safety research. The rational thing for users is to assume DALL-E 2 is contaminated by CSAM unless OpenAI can publicly demonstrate otherwise. Their leadership is too stupid and dishonest for people to take their word.
https://ieeexplore.ieee.org/abstract/document/9423393
https://arxiv.org/pdf/2110.01963.pdf?trk=public_post_comment...
https://www.theguardian.com/technology/2023/aug/02/ai-chatbo...
I find OpenAI's refusal to document what's in the training data infuriating, but I that doesn't lead me to assume they weren't able to filter out CSAM from it.
If they don't care about this issue at all, why are they exposing low paid workers in Kenya to these conditions?
For people writing on the internet for motivations other than ad-driven monetization, and those who read their stuff, there is basically no issue.
I doubt these users will not experience a drop in quality as most authorship will roll out LLM laden content which will add more noise to the signal.
This is the key, I think. But it's small--The Small Web. I think we'll be better off, though.
Have a day job and produce quality content in your spare time, is my motto. A few billion people producing content in their spare time is still more than we can ever hope to consume.
That seems like it might make the training process untenable.
Information production seems to be on an exponential curve, so it won't be that long before you're training on a minority of all data.
If you were to teach a LLM to play chess today - wouldn’t you use fully AI generated data because that’s by far the best data ?
The notion that you could just grab data off the internet completely unvetted, and then train a useful model off that data seems very unrealistic. To get anything useful out of real world crawl data, you need a lot of filtering and processing to select the good bits. This is true if you're training an AI or building a search engine.
In an ironic twist the gatekeepers of the semi-walled gardens/islands of email that arose are now harvesting it for training content. I do so despise those weekly machine generated synopses telling me how I am doing.
https://www.udio.com/songs/p66uVGEgifEBLdoR5Ttyue
It's not that it's perfect, it's cheesy and commercial and obviously lacking some coherence- but it also has some pretty good parts and... performances... and if passed on radio it would certainly attract attention. A casual listener would hardly imagine that it's AI-generated.
The scariest part are the lyrics: "lorem ipsum dolor sit amet...". Nonsense, just pure mindless nonsense, sung with feeling and expression. I think this is a good example of the things to come.
If you're a song maker you can probably use LLMs to put together a decent lyric or help in the process of turning your ideas into something that makes sense.
To think that you just press a button and something comes out of it and you're happy with that, eh, I mean it's possible but that's never what arts been about. You can get lucky, but you can have consistently decent results by just putting some more effort into the creation process.
Then, before deciding that we like a song that we like, we'll have to investigate whether it's real or not. If the emphasis and expression and tonal changes actually correspond to the intents and emotions of a performer or are just created from nothing.
I don't see this sort of thing as much of a threat to the mainstream music industry, since the mainstream music industry isn't really about selling music, it's about selling concert tickets.
Before, you actually had to pay people to compose songs and those people had to sit down at a piano and write the songs, then other people had to perform the songs.
Now, I can buy a certain number (?) of the latest NVIDIA GPUs, build a cluster, and then churn out songs literally 24/7. Endlessly, song after song after song, each minute a new song. Just press the button, and a new song it created -- _no other humans involved_.
Yes, slop, kitsch and mass produced music were always produced. Never at this scale.
See also the k-pop revolution (as stanned by people who don’t understand a word of Korean)