I’ve disputed that fact before. It depends what part of your data is generated from GPT4 if I have high quality code and I’m using gpt4 to synthetically generate variations on how some could ask for the code to be written it’s entirely possible because of the high quality code for the model to be better. It’s not all or nothing if you’re mixing synthetic data with quality sources.
In this case even though parts of the dataset are synthetic the bound is on the code not necessarily the 50 ways I got gpt4 to say “write me a script to do x” or modeled other interactions with that code data source.