ChatGPT is trained on the internet corpus written by humans. Over time, more and more of this corpus will be written by ChatGPT.
Has anyone tested what happens when the output of a LLM is fed back as the training data?
Has anyone tested what happens when the output of a LLM is fed back as the training data?
Why was this possible? The game itself acted as an anti-bullshit filter. So you can train on your own generated data if you filter it.
Like this one: Large Language Models Can Self Improve