suddenly the confirmed quality of the scraped data will be at a premium.. "Scrape Engine Optimizers" ?
With AI, we have an exponential level of productivity. But what is being produced? 90%: garbage.
The problem is that what is being produced is essentially "garbage" generated by models trained on garbage. Quality knowledge is increasingly submerged and suffocated by spam and low-quality content.
The real challenge of the future will be filtering and cleaning up, on each level.
We already see this with synthetic training data that basically uses logic in form of math and code as constraint.
I've heard this argument before, but you don't need to think too hard to see the limitations of a machine with no senses.
There's a paper on it somewhere.
Only if you assume that people who train models are stupid.
And it's simply not reasonable for AI companies to have human hands read through individual comments everywhere from beginning to end to build their training data. There isn't enough time in the universe to advance AI while doing that and also being accurate. Something will always slip through.
Why would human review be the only possible way to remove enough of the tainted training data?
> Where are they supposed to find content to train AI on that isn't polluted with AI content that'll result in a feedback loop?
If nothing else: you could look for old data. At the moment, training assumes that input data is essentially without limit. But machine learning has lots and lots of old and proven techniques for what to do when your training data is limited.
You can also look into techniques for avoiding model collapse. Just because one group of researcher showed that this happens with some specific models, doesn't mean it needs to happen in general.
Assume it's exactly the arm's race you suggest. So OpenAI develops a new AI-content detection today, and tomorrow someone brings out something that defeats it. In a month OpenAI brings out another new technique that is unbeaten for only a day. Rinse and repeat.
Is that what you had in mind?
In any case, OpenAI knows exactly what it crawled when. So if it has a technique that only got beaten at timestamp X, then everything they crawled up to is good to use.
> And it'll get harder to validate it.
Sure, and that's true in general: advances become harder, because we pick the low hanging fruit first anyway. Nothing new about that.
The arms race I described is also only one approach. Model trainers will also want to investigate economising on training data, making their approaches more robust to model collapse, multi-modal training, and a million other strategies and tricks that I can't think of in thirty seconds.
The very comment you just replied to said so.
And even that needs to be curated because before AI tools there was bot content filling up the internet.
...and even without bots, a lot of human authored content are low value, poorly written, etc.
There are (probably) companies out there whose business is to create, curate and improve training sets.
Probably the only real way to validate content is real is building a validation system into devices. Confirm when a photo is taken and send an ID to a server, then when photos are shared, its ID is compared to the image on the camera/phone manufacturer's server. For text, validate every little key press. And there are still ways to game these systems, but I would not be surprised if they're introduced to mitigate AI diffusing everywhere.
I don't believe that the people who train models have a secret way of identifying and filtering out bot-generated content that no one else (email spam filters, search engines, etc.) have identified. I do believe that they feel their models need to have up-to-date information on a variety of topics that require regularly ingesting new data. So no, I don't think they have a good way to avoid their inputs rotting from their outputs.
And what do you mean about short term gains? If you are training a model, and you see model collapse, where's the short term gain? I don't get it.
The incentive for the person who trains a model, even in the short run, is for them to avoid model collapse.
> I don't believe that the people who train models have a secret way of identifying and filtering out bot-generated content that no one else (email spam filters, search engines, etc.) have identified.
Huh, why would they need a secret filter? A filter is only one way thing you can try. You can also look into using different models, different training, making your approach more resistant to model collapse; training multi-modal models, using approaches to economise on training data; and thousands of other ideas I can't think of in thirty seconds.
> So no, I don't think they have a good way to avoid their inputs rotting from their outputs.
You lack imagination. People can be remarkably clever if its in their (short term!) interest to find solutions.
Two very significant things happened in 1981.
After years of claiming the government couldn't help people, Ronald Reagan was elected and Republicans have been working hard to make that statement more true ever since. A big part of that was deregulation of the financial markets.
That same year, Jim Welch became Chairman and CEO of General Electric. He juiced the stock prices by selling off the company's prize jewels, real estate, and future, and for a while (before the utter collapse of the company) artificially raised the stock price so high that executives around the country copied him, and and an entire industry of vultures like Mitt Romney started private equity firms to cannibalize healthy companies for their personal profit.
Someone in the chain will be. Even the smartest people buy a lot of their training datasets. What happens when those get contaminated?
Filters are also not 100% infallible
> And you negotiate a contract where the seller bears some of that risk
So the training data will be polluted anyway, but "the seller will bear some risk"
Why would they need to be?
A small number of samples can poison LLMs of any size https://www.anthropic.com/research/small-samples-poison
--- start quote ---
In a joint study with the UK AI Security Institute and the Alan Turing Institute, we found that as few as 250 malicious documents can produce a "backdoor" vulnerability in a large language model—regardless of model size or training data volume. Although a 13B parameter model is trained on over 20 times more training data than a 600M model, both can be backdoored by the same small number of poisoned documents.
--- end quote ---
Poisoning is a completely different topic.
It's not a different topic. It's literally the topic of this branch of discussion.
> The internet becoming majority bot content basically guarantees this becomes a real problem for the next generation of models.
Only if you assume that people who train models are stupid.
--- end quote ---
And then literally everyone who commented on this, including me, was talking about issues with training data contamination. And you are the only one dismissing it as nothing important that can be easily fixed.
> The bigger concern is what happens when AI models start training on AI-generated content at scale. We're already seeing model collapse in research papers where output quality degrades when training data is contaminated with synthetic text. The internet becoming majority bot content basically guarantees this becomes a real problem for the next generation of models.
Model collapse.
Eg by filtering data, by procuring better data, by applying techniques for making do with more limited data (we used to have a lot of those, and they are still known), or you can also adapt your training process to be less vulnerable to model collapse. Just because some researchers have shown that this happened for the models they tested, doesn't mean it has to be a universal thing.
It's just harder when you cut all traffic to them, devalue their work and fill the air with AI noise.
We'll have the internet we deserve
Marx, Nietsche, Debord, Foucault, Baudrillard, Adorno - they already saw writing on the wall, or at least fragments of it.