I thought training LLMs on content created by LLMs was ill-advised but this would suggest otherwise
I thought training LLMs on content created by LLMs was ill-advised but this would suggest otherwise
The exact training is proprietary but they seem to use a lot of GPT-4 generated training data.
On that note... I've often wondered if broad memorization of trivia is really a sensible use of precious neurons. It seems like a system trained on a narrower range of high quality inputs would be much more useful (to me) than one that memorized billions of things I have no interest in.
At least at the small model scale, the general knowledge aspect seems to be very unreliable anyways -- so why not throw it out entirely?
You can get diverse low quality data from the web, but for diverse high quality data the organic content is exhausted. The only way is to generate it, and you can maintain a good distribution by structured randomness. For example just sample 5 random words from the dictionary and ask the model to compose a piece of text from them. It will be more diverse than web text.
I agree if we are talking about maxing raw reasoning and logical onference abilities, but the problem is that the ship has sailed and people expect llms to have domain knowledge (even more than expert users are clamoring for LLMs to have better logic).
I bet a model with actual human “intelligence” but no Google-scale encyclopedic knowledge of the world it lives in would be scored less preferentially by the masses than what we have now.
Phi is another really good example but that's already covered from the article.
[0] - https://www.latent.space/i/146879553/synthetic-data-is-all-y...
Mode colapse theories (and simplified models used as proof of existence of said problem) assume affected LLMs are going to be trained with poor quality LLM-generated batches of text from the internet (i.e. reddit or other social networks).
Millions? Where are they? Where are they used?