If I were training a code model I'd take a snippet of code, have the existing LLM explain it. Then use the explanation and the snippet for the test data.
You want us to rely on models that are overfit to hallucinated LLM interactions.
* The LLMs have a sufficient "understanding" of the request and of how to write code to fulfill the request
* Have a way to validate the suggestion by actually executing the code (at least during training) and inspecting the output
From what I've seen we are still far away from that, Copilot and GPT-4 seem heavily reliant on very well-commented code and on sources like Stackoverflow
/e: sorry, sounds a bit stand off-ish.
Let me give an example: I was trying to find a way to clone a gorm query to keep the code clean. The documentation doesn't have anything (no, .Session isn't a solution) and the only place I had was issues discussing that. Apparently you can't. So I'll be ditching gorm and move to pgx in the near future. That's how it happens for me all the time. The documentation is lacking the hard part, always.