is synthetic data a really big deal right now and LLM? if so, are there any take-home ideas that might apply to other areas, say analysis of MRI?
In other fields like computer vision, synthetic data is useful for generating ground truth data, like for depth masks.
Asking GPT-4 to create 50 new conversations from a chapter of a textbook creates higher quality data than most of “the pile” and can extend beyond what GPT-4 has in its existing dataset. This is exactly what MS has done with Phi.
In this case even though parts of the dataset are synthetic the bound is on the code not necessarily the 50 ways I got gpt4 to say “write me a script to do x” or modeled other interactions with that code data source.