Where Are Large Language Models for Code Generation on GitHub?
arxiv.org
arxiv.org
However, it was good for automating boilerplate and for one-off utilities that were tedious to write. In those two categories, many new uses are similar to previous uses in the training data. So, extrapolation is easier for such jobs.
The abstract suggests that’s the kind of use they’re doing. That’s also why they’re not usually fixing the bugs. We can often ignore or work around bugs in such use cases.
It is also a tiny project, just a single microservice, although my upcoming open source project is fairly substantial and I suspect it'll be at least 30% LLM generated.
The fact is LLMs are really good at generating chunks of code, sometimes even entire files, but they aren't that good at making systematic changes to large swaths of existing code. They can be made to work for that scenario with tons of prompt engineering and prompt chaining, but that isn't where LLMs shine.
For example almost all of the HTML in that project was made by Claude. I then modified the code and re-uploaded it and asked Claude to add collapsible info boxes to parts of the page after the fact, and Claude completely failed at the task.
Still though I am seeing an incredible increase in my development velocity.
[1] https://www.github.com/devlinb/escaperoom -ugly UI, designed to help people create escape rooms for LLMs. A hosted version will be up later today!
"Specifically, we first conduct keyword searches, such as “generated by ChatGPT” and “generated by Copilot”, to locate GitHub code files that include such keywords, retaining only those files that contain GPT-generated code."
This seems like a pretty serious weakness to me; presumably there is a lot of code generated by LLMs that isn't annotated as such.
I don’t label my code when I use GitHub copilot.
I carefully read the code and use TDD where GitHub copilot generates code to match my test spec.
I also provide samples of already written code to ensure it follows a similar pattern…