WaveCoder: Enhanced instruction tuning with refined data generation
arxiv.org
arxiv.org
In other fields like computer vision, synthetic data is useful for generating ground truth data, like for depth masks.
In this case even though parts of the dataset are synthetic the bound is on the code not necessarily the 50 ways I got gpt4 to say “write me a script to do x” or modeled other interactions with that code data source.
Asking GPT-4 to create 50 new conversations from a chapter of a textbook creates higher quality data than most of “the pile” and can extend beyond what GPT-4 has in its existing dataset. This is exactly what MS has done with Phi.
https://nitter.net/TeamCodeLLM_AI/status/1747652471714144702