“Imagine a newsroom where you have to produce a daily newspaper and you suddenly stop getting feeds from the outside world. Your earlier newspapers are the only source.”
This is what to me LLMs eventually would get to—same content being fed again and again.
If you just slurp all AI content sure, you get the collapse this paper talks about. But if you only ingest the upvoted conversations (which could still be a lot of data, and is also a moat by the way) what then?
The other reason I find this line of argument overly pessimistic is we haven’t seriously started to build products where this gen of LLMs converse with humans in speech; similar opportunities to curate large datasets there too.
Finally, there is no reason OpenAI cannot just hire domain experts to converse with the models, or otherwise build highly curated datasets that increase the average quality. They have billions of dollars to throw at GPT-5; they could hire hundreds of top tier engineers, mathematicians, economists, traders, or whatever, full time for years just debating and tutoring GPT-4 to build the next dataset. The idea that slurping the internet is the only option seems pretty unimaginative to me.
Considering that RLHF took GPT-3 from a text completion model to an instruction following chat bot, you could use expert feedback to fine tune the model in whatever domains you wanted or a mixture of domains to produce an even more generally capable model.
If it didn't work for cyc...
given how prolific bot farms/karma farms/etc are, you might still end up in the same spot with this criteria.
https://mwichary.medium.com/one-hundred-and-thirty-seven-sec...
My point was that this will not be necessary for the same reason it isn't necessary to filter out human-made content.
We’re just reviewing our prior stats and insuring they do not deviate too much such that the wrong people would be impacted.
At the very minimum, you can assume every piece of text data pre Dec-2022, and every image before Aug-2022 to be completely human made. That still leaves decades of pure human digital data, and multiple centuries of distilled human data (books) to be trainable on.
And we haven't gotten into videos yet, which is another giant source of data yet unexplored.
Never forget, humans train on human-generated data. There's no impossible theoretical reason why AI cannot train on AI-generated data.
Should be some type of nuke that would release weird isotopes but with minimal toxicity?
World War Two is rather convenient in that respect, as there are large quantities of steel that were left to sit around for several decades after those ships sank.
It’s been a while since I’ve done low-background gamma-ray spectroscopy, but I believe there were some setups that went even further, using lead that had been smelted by the Romans. That way, any contamination present at the time of smelting would have a few thousand years to decay away.