The Curse of Recursion: Training on Generated Data Makes Models Forget
arxiv.org
arxiv.org
"There is very little information available about OpenAI’s forthcoming successor to ChatGPT, GPT-4. But I’m going to make a prediction: when assembling the vast amount of text used to train GPT-4, the people at OpenAI will have made every effort to exclude material generated by ChatGPT or any other large language model. If this turns out to be the case, it will serve as unintentional confirmation that the analogy between large language models and lossy compression is useful. Repeatedly resaving a jpeg creates more compression artifacts, because more information is lost every time. It’s the digital equivalent of repeatedly making photocopies of photocopies in the old days. The image quality only gets worse.
Indeed, a useful criterion for gauging a large language model’s quality might be the willingness of a company to use the text that it generates as training material for a new model. If the output of ChatGPT isn’t good enough for GPT-4, we might take that as an indicator that it’s not good enough for us, either. Conversely, if a model starts generating text so good that it can be used to train new models, then that should give us confidence in the quality of that text. (I suspect that such an outcome would require a major breakthrough in the techniques used to build these models.) If and when we start seeing models producing output that’s as good as their input, then the analogy of lossy compression will no longer be applicable."
[1] https://www.newyorker.com/tech/annals-of-technology/chatgpt-...
All AIs built are trained on data from the Before Times, and even though they try to assimilate, the way a teenager tries to adapt to the local accent of a new town, there are always moments where they slip up and reveal their geography.
But we also know that kid who learned everything from books, pronounces the words wrong and uses definitions for them that nobody has used in decades (work with one of those now. I thought I could talk people to death, and he wears even me out.) those AIs will sound like out of touch nerds too.
Someone once ranted about people using big words, “having a crush on their high school English teacher they never got over.” I knew exactly what he meant. I spent my childhood hiding how smart I was and part of my 20’s reveling in it. It’s off putting. Wisdom comes from everywhere, and the smartest often have the least. Now I’m solidly in the Feynman camp: if you can’t explain your domain to college freshmen then you don’t know what you’re talking about (yet).
When I’m refactoring code and introduce a new concept that’s like another one but different rules, I jump straight to a thesaurus. The word I pick out of the air might be good enough, but I guarantee you there’s a better one out there. Not fanciest or longest word, the most concise one. (Example: kind vs type in some circles of Type Theory). Some people act like that’s a crutch, but I’ve reviewed or refactored their code so I know that opinion and $5 isn’t worth a cup of coffee.
Do you really think it is just a coincidence that you happen to exist at the very last point of high value pre-AI data? ;-)
High-performance error correcting codes have the property that the closer they operate to the Shannon Limit, the better they perform when below the limit but the more dramatically they fail when the limit is exceeded. Gut feeling says the same should be true for LLMs: as the model/compression ratio gets better, and a "Shannon Limit" is approached, they should perform better but fail more spectacularly if the limit is exceeded.
The link between neural nets and information theory is well known, but there don't seem to be many results out there for LLMs. No doubt there are rooms full of PhD students working on it?
https://medium.com/@chris_bour/bridging-information-theory-a...
Giving an AI the ability to self modify would just be a roundabout way of training it on itself. Repeatedly compress a JPEG and you don’t get the “enhance” effect from Hollywood. You get degraded quality and compression artifacts.
AI in a vat that can't do it is obviously useless. It's the ML equivalent of a computer running purely functional software: i.e. just sitting there and heating up a bit (though technically that is a side effect).
Conversely, any AI that's meant to be useful will be hooked up to real world inputs somehow. Might be general Internet access. May be people chatting with it via REST API. Might be a video feed. Even if the AI exists only to analyze and remix outputs of LLMs, those LLMs are prompted by something connected to the real world. Even if it's a multi-stage connection (AI reading AI reading AI reading AI... reading stock tickers), there has to be a real-world connection somewhere - otherwise the AI is just an expensive electric heater.
Point being, you can assume every AI will have a signal coming in from the real world. If such AI can self-modify, and if it would identify that signal (or have it pointed out) as a source of new information, it could grow based on that and avoid re-JPG-compressing itself into senility.
ChatGPT is clearly HEAVILY persuaded to respond in a particular stock style. It's being artificially hamstrung and constrained in a sense. So all of its output, even if it covers a variety of subjects, will often use very similar patterns and writing styles. "it's worth nothing....", etc.
So unless they unshackle these constraints, which is unlikely for obvious reasons, isn't this always going to be inevitable?
The problem for companies like OpenAI is that this isn’t worth their valuation without lots of further continued investment.
Enter Microsoft who is doing everything they can to feed the next training models by using users data without explicit permission.
As customers and competitors truly grok this, MS + OpenAI strategy will tested.
With "free data" from reddit gone, now the cost of symbiotic gen+human data will be even more expensive.
It's seems the best one could hope for is that recycling generating text into new training data would be not detrimental. But it's really difficult for me to imagine how this would ever be useful. It seems this would imply that the LLM had somehow managed to expand the dimension of vector space spanned by the original training data. Which sounds either impossible or like the model became sentient.
The number of dimensions? Well, not by itself I guess. But the span of output compared to training data? Sure, why not?
I think it's also worth pointing out there's a difference between text produced by an LLM looped on itself, which arguably may not contain any new information and would be like repeatedly recompressing the same JPG, and text produced by LLM/human interaction. The latter is indirectly recording new knowledge simply because people's prompts are not random. Even with human part of the conversation discarded, feeding such LLM output back into training data would end up selectively emphasizing associations, which is a good signal too (even if noisier than new human-created text).
Back when enthusiasts and researchers read mostly textbooks, papers, and Wikipedia (rather than blog posts, tweets, and READMEs) there was much more discussion around the 'InfoMax Criterion' -- quite elegantly demonstrated by Tishby et al. via his closely-related 'Information Bottleneck' studies -- which is just that idea: mutual information is maximized between input and output, subject to inherent limits of statistical processing of the realizations of such systems. What determines the asymptotic maximal value of the mutual information is the inductive bias of the system vis a vis the training set. This is all standard theory, perhaps so fundamental that it is obscured by all the application-oriented study and instruction.
It’s compression all the way down.
If there were many different models being trained and used widely it would, at minimum help mitigate this issue. Also, having multimodal models will likely change the balance. If models can train directly on "real world" data that can help fill in the entropy gaps.
Incestuous learning is pretty much guaranteed unless generated content starts being flagged or there is an explosion of entirely novel models.
Personally, I'm hoping we will see a Cambrian explosion of new LLM models and approaches. We've seen some beginnings of this in image generation so it's not entirely implausible.
Another thing the study doesn't capture: What is the effect of combined human + AI content? It's plausible that an explosion (due to lowered barriers) of new human guided/augmented ai content could counteract the effect.
“Imagine a newsroom where you have to produce a daily newspaper and you suddenly stop getting feeds from the outside world. Your earlier newspapers are the only source.”
This is what to me LLMs eventually would get to—same content being fed again and again.
If you just slurp all AI content sure, you get the collapse this paper talks about. But if you only ingest the upvoted conversations (which could still be a lot of data, and is also a moat by the way) what then?
The other reason I find this line of argument overly pessimistic is we haven’t seriously started to build products where this gen of LLMs converse with humans in speech; similar opportunities to curate large datasets there too.
Finally, there is no reason OpenAI cannot just hire domain experts to converse with the models, or otherwise build highly curated datasets that increase the average quality. They have billions of dollars to throw at GPT-5; they could hire hundreds of top tier engineers, mathematicians, economists, traders, or whatever, full time for years just debating and tutoring GPT-4 to build the next dataset. The idea that slurping the internet is the only option seems pretty unimaginative to me.
Considering that RLHF took GPT-3 from a text completion model to an instruction following chat bot, you could use expert feedback to fine tune the model in whatever domains you wanted or a mixture of domains to produce an even more generally capable model.
If it didn't work for cyc...
given how prolific bot farms/karma farms/etc are, you might still end up in the same spot with this criteria.
https://mwichary.medium.com/one-hundred-and-thirty-seven-sec...
My point was that this will not be necessary for the same reason it isn't necessary to filter out human-made content.
We’re just reviewing our prior stats and insuring they do not deviate too much such that the wrong people would be impacted.
At the very minimum, you can assume every piece of text data pre Dec-2022, and every image before Aug-2022 to be completely human made. That still leaves decades of pure human digital data, and multiple centuries of distilled human data (books) to be trainable on.
And we haven't gotten into videos yet, which is another giant source of data yet unexplored.
Never forget, humans train on human-generated data. There's no impossible theoretical reason why AI cannot train on AI-generated data.
Should be some type of nuke that would release weird isotopes but with minimal toxicity?
World War Two is rather convenient in that respect, as there are large quantities of steel that were left to sit around for several decades after those ships sank.
It’s been a while since I’ve done low-background gamma-ray spectroscopy, but I believe there were some setups that went even further, using lead that had been smelted by the Romans. That way, any contamination present at the time of smelting would have a few thousand years to decay away.
I wonder if opening GPT and DALLE to the public was partly intended to pollute subsequent data for anyone that gets into AI down the road. Suddenly a lot of publicly accessible data is worth less, leaving only players who've got a hoard of time-stamped data to compete with (like Google, Facebook). OpenAI almost certainly has the hashes of what it spits out too, so they'll be able to sort the wheat from the chaff for a while yet.
The market for data may be getting interesting.
Normal hashes are extremely fragile, so they'd have to use something more sophisticated. Scott Aaronson said in a podcast a few months ago that OpenAI has implemented such a system but at the time they had not decided to start using it.
The purpose being discussed at the time was to provide a tool for educators to detect cheating, but presumably it could also be used for filtering future datasets.
Make sure you have a representative training dataset, real or synthetic, it doesn't matter.
They could have replicated it here by having GPT-4 score the samples and throwing out most (but not all) of the bad ones. I have no idea what would happen if you e.g. throw out a majority of the bottom 70% and keep the top 30%. It's conceivable to me that it would end up improving or at least not getting much worse with each generation.
Even the best-looking JPEG (as judged by humans) is still lossy.
You gotta add something new to the mix. That's possible when the AI is part of a larger system. AlphaZero demonstrated that even self play could be a source of signal, as long as it gets the feedback of the game, the model can learn.
> Evolution through Large Models
To me, this research supports a hypothesis I've had for a while that we're going to get to truly excellent AI by using synthetic data to bias it towards excellence and away from mediocrity.
$20 says the next round of major model training is using synthetic data for training that was filtered through a discriminator trained entirely on human data.
The human data as a reference is certainly important to avoid polluting (and to its point there's an advantage for those already having it), but moving away from edge cases isn't necessarily a bad thing practically given edge cases can result in negative practical performance (as opposed to academic next token performance).
Once they can do that and produce a productively improved model, then that's really the start of self-improvement.
Isn't that interesting? The idea of "mental liquidity", or "strong opinions weakly held"? https://news.ycombinator.com/item?id=36280772
Well, aren't they? I believe any kind of reinforcement learning is supposed to be biased into the last training set.
The one thing I don't get and could have been missing in the past... a lot of the corporations and private things, like farms operate on debt. Now maybe it's a bit reductionist, but if you're a farmer operating on debt, if interest rates go up you need to increase prices to cover operating expenses. And this get compounded all the way up to the end consumer as every step in the supply chain marks up by a fixed percent, and because everything is getting more expensive decided lets mark up by a larger percent. So higher interest rates really could be contributing to inflation. And it's just creating a cycle. And with the current levels of debt never seen before in history, it's unlike other periods.
LLM garbage, on the other hand, would take up the whole model..
- Generated content added to LLM. QED.
- Generated content added to SERP. DEAD.
;-)
In self-play, the objective measure of the "truth" of a move can ultimately come out of the rules of the game which any machine can compute.
With an LLM, the machine is only aping, emulating, simulating the output the people have produced about the world - the machine has no access to actual "real world" that people are using language to describe. Human beings talking about the world is data that increases your own knowledge of the world - your own predictions of that talking, not so much.
Human based RL is used because humans know stuff about the real world and can sort language utterances by this. There's "self play" process that gives a system this sort of knowledge.
I would think that the human selection process would help to head off this conversion, both by selecting against incorrect results, and also by introducing variance outside of the model. On the other hand since a person can only act as a filter, I can also see how that would be of limited value long term.
So, while interesting for nailing down that phenomenon, the broader implications everyone wants to draw from it are not very good - very few people are using random GPT-3/4 or Stable Diffusion samples!
This paper has one huge hole in it: it assumes that content on the internet is not moderated and that the training dataset will never evolve to take rating into consideration. On social media, the form of moderation is # of likes. Once detected, bots that output bad data will be banned and content deleted.
The key issue I have with the paper is for good synthetic data it is impossible to tell it apart from human generated data.
Let's say your LLM can generate text with the same quality as the input 98% of the time, and the other 1% of the time, it's wrong. Each round of recursive training amplifies that error. 96% accuracy after the next round. 67% after 20 rounds.
There's no way for it to get better without more real human input.
Huh? If the LLM was only ever spitting back identical content straight from the training set, that would be a symptom of extreme overfitting, which everyone universally agrees is a bad thing — not a perfect thing.
You can see the whole process as a very high latency reinforcement learning system.
Also, the curation process counts as "real human input" itself, so I don't think it contradicts my point.
Worse than this, since the training set already takes rating into consideration. For example the reddit dataset includes only articles with more than 3 upvotes.
Now wait till the generated content is indistinguishable from human content (to humans) and it will be hard to figure out what's in your training set.
Certain skills interfere with each other. For instance playing chess makes you worse at go, and playing go makes you worse at chess. Certain pairs of languages are likewise hard to learn together - my son found that he could not study both Russian and Chinese at the same time.
But in general you just develop more and better memories.
He still plans to learn Russian, though.
Actually, not scary as it forces everyone to look for solutions both the in-box kind and the out-of-box kind.
1st summary: The text discusses the impact of large language models (LLMs) on online text and image generation. It argues that LLMs can cause model collapse, a phenomenon where the original diversity of data is lost due to feedback loops between models and data sources. It shows that model collapse can affect various types of generative models and warns that it poses a threat to the quality and reliability of online content. It suggests that data from real human interactions will become more valuable to prevent model collapse.
2nd: The text warns that large language models (LLMs) can cause online content to lose its original diversity and become less trustworthy. It explains that LLMs can create feedback loops that reduce the variety of data they learn from. It suggests that real human data will be more important to avoid this problem.
3rd: The text argues that LLMs can make online content more homogeneous and unreliable by creating feedback loops that limit their data sources. It recommends using real human data to prevent this issue.
4th: The text warns that LLMs can produce biased and repetitive online content by relying on narrow data sets. It suggests using diverse and authentic human data to avoid this problem.
5th: The text advises using varied and real human data to prevent LLMs from creating online content that is unfair and redundant due to limited data sources.
6th: The text suggests using diverse and authentic data to avoid biased and repetitive online content by LLMs.
7th: The text argues that LLMs need better data to produce more varied and realistic online content.
8th: The text claims that language and literature models (LLMs) require more diverse and accurate data to generate online content that reflects the real world.
9th: The text argues that LLMs need better data to create realistic online content.
10th, Bing started to hallucinate, I double-checked that it didn't search and the chat was new: The text claims that online content creators who use language models (LLMs) should have access to more reliable and diverse data sources to produce realistic and accurate content. It warns that LLMs can generate misleading or harmful content if they are trained on biased or low-quality data. It also suggests some ways to improve the data quality and availability for LLMs, such as using data verification tools, crowdsourcing data collection, and collaborating with other content creators.
11th: The text argues that online content creators need better and more diverse data sources to use language models (LLMs) effectively. It cautions that LLMs can produce harmful or inaccurate content if they rely on flawed data. It also recommends some methods to enhance the data quality and access for LLMs, such as verifying data, crowdsourcing data, and cooperating with other creators.