CodeCompose: A large-scale industrial deployment of AI-assisted code authoring
arxiv.org
arxiv.org
In other terms, out of 4.5 million suggestions about 80% were off, yet there is 91% positive reception. That's 3.6 million rejected suggestions that potentially distracted programmers from doing their work. Yet users are happy. Is there a contradiction in these figures?
Though I guess the success rates when using Stack Overflow aren’t too dissimilar.
Meaning, how many rejected solutions were sufficient to give the engineer enough context to turn a 30m task into a 5m task because they generated a recommendation, got an idea, rejected it, and rewrote it more efficiently or more correctly?
There's a lot of "devil in the details" likely buried in here.
I am well aware that others are having a different experience with it.
Some folks may never accept AI code completion / suggestions (like some prefer vim over modern IDEs) but at least people working on this stuff can describe points known to focus on.
In this case it gave me 3 suggestions but I only accepted 1. I could see this taking 5-10 suggestions for an LLM to when it’s not something as straightforward as a function name. It’s still very useful despite this low acceptance rate
- anticipate when the suggestions are likely to be useless and not even bother
- scan the proposals to see if they are what you want in cases it's useful
It's a boilerplate generator and you're happy when it saves you tedious mental effort.
On the other hand the person trying to track down a subtle bug afterwards might be a little less happy at having to wade through oceans of boilerplate.
if(a) {...}
if(b) // here it predicts line by line from a once you start with similar logic
if(c) // here it will do a one shot generalization of a and b for c
Likewise method variations, enum mappings, inferring call parameters from names and signature, etc. - these things are trivial to check and test but take effort to type out. When you know what you want to do and someone suggests the solution you had in mind you're happy you saved half a min or min of typing.I'm as, if not more, likely to zone out on tedium and introduce the subtle bug myself.
AI code completion (like Github Copilot) is like this. Still a time saver overall, even with a low acceptance rate.
Using GitHub copilot daily I find it’s suggestions often nonsense but interesting to see regardless. Often for boilerplate it’s spot on and it saves me dozens of lines of typing. But it also suggests stuff on every key stroke many of which I just type through, similar to intellisense. Assuming Metas code thingy is better, I would find myself in that 91%, as I’m already there with what’s available to the general public.
My only gripe, fwiw, with copilot in vscode is it interferes with intellisense. Often I want to see the code completion from both, but copilot jumps in before intellisense and the intellisense never renders and I use it as an inline api reference. Sometimes it’s so frustrating I have to turn off copilot. But, copilot is generally useful enough that I reenable it once I’ve understood the api stuff I’m unsure of. There’s some escape backspace period dance I can do that sometimes let’s intellisense win. I’ve not dug deeply enough into vscode configuration to know if there’s some parameter to tweak the race conditions. I’d note that when intellisense renders first copilot still renders its suggestions but the other way doesn’t work.
> The final model was calibrated for a target precision of 50%. That is, we tuned the model and the suggestions filtering, so that 50% of suggested edits on our evaluation dataset are correct. In general, increasing the target precision reduces the number of shown suggested edits, and decreasing the target precision leads to more incorrect suggested edits. Incorrect suggested edits take the developers time and reduce the developers’ trust in the feature. We found that a target precision of 50% provides a good balance.
Also, it seems like if the suggestions are too good then they’ll be blindly trusted and if they’re too bad they’ll be ignored?
Where to set the balance likely depends on the UI. For a web search, how many results do you click on?
[1] https://ai.googleblog.com/2023/05/resolving-code-review-comm...
- As mentioned in this paper I definitely do not want the AI suggestion crowding out a suggestion generated directly from the type bindings
- I often do want the AI to write an entirely new block of boilerplate. To do this you have to write a comment string targeted at the AI, then delete this afterwards
- Sometimes I'd just like the AI to explain to me what some code does without writing anything
- This isn't something I always want on; I find myself turning the plugin on and off depending on the context
Overall I think we need a novel UX to really unlock the AI's helpfulness
Here are some chat transcripts that give a flavor of what it’s like to code with AI this way:
My tool is open source, and currently only works if you have a gpt-4 api key.
I wrote up some notes on one effective approach:
https://github.com/paul-gauthier/aider/blob/main/docs/ctags....
As for the getting a suggestion by writing comments, an "insert from prompt" action perhaps, or just a separate prompt pane/popup/whatever-you-prefer combined with using good ol' copy+paste would suffice.
If you want to know what some code does, just select it & hit a keyboard shortcut (or right click and choose explain from menu).
If you want AI to write code for you, write a comment starting with a specific word, it suggests the implementation and you can choose to accept & replace the comment with it.
You can select code and ask it to explain it to you, or ask it to generate some boilerplate / code and then insert it at your cursor position without adding any comments like you described for prompting the Copilot autocomplete.
It seems okay, but I don't really use vscode that much
Hallucination about API calls was reported as a problem. I've seen that one. There's an amusing, and seriously annoying, tendency for these systems to make up some plausible API call that does what you need, but doesn't exist. Maybe something should collect up such suggestions as proposals for new API calls.
Disclaimer: I am the person in the video.
As someone who writes quite a lot of Hack, I'm selfishly interested in whether you plan to open-source this work (not the weights, obviously, but everything else).
I'm thinking of scenarios where they might get an idea and then reject them and rewrite it better.
First time I got my hands on Github Copilot some time ago, I intentionally chose something "odd" (so no CRUD API in X, Kafka stuff, sorting algorithm) so I chose to re-create some Excel sheet alike UI in React & Canvas.
First everything proposed was misleading but after the first couple of lines, it astonished me by really proposing things like nested loops for iterating over a grid (surely, I used an according function name). I came quite far (for a quick test) with marking single and multiple cells and being able to drag the marked area and double-click for text input. Then I realized it was about time to invent a proper data model and stopped it there.
All in all, I'm using it for like 50% of daily coding, but it's frequently misleading in subtle tricky ways that I expect to be quite dangerous especially if using it for tests as well (need to double-check if it's not BSing me twice). Debugging code is already a big drag in our work and I'm somehow not happy with debugging AI-generated code.
https://austinhenley.com/pubs/Vaithilingam2023ICSE_IntelliCo...
Disclaimer: I'm one of the co-authors.
On M1/M2, it offers a convenient single binary deployment, thanks to Rust. You can find the latest release at https://github.com/TabbyML/tabby/releases/tag/latest
(Disclaimer: I am the author)
Arguably if you optimise the system to do precisely that rather than pretend to hold civilized conversation you could get quite a bit more bang for you gazillion parameter buck.
Then everything goes dark, a few companies will absorb and confuse the art or otherwise find ways to exclude everyone else from the process.
Feels like tech is making billions but is a little lost ?
"In this paper we present CodeCompose, an AI-assisted code authoring tool developed and deployed at Meta internally. CodeCompose is based on the InCoder LLM that merges generative capabilities with bi-directionality. We have scaled up CodeCompose to serve tens of thousands of developers at Meta, across 10+ programming languages and several coding surfaces."
Looks like the InCoder model it's based on can be downloaded here: https://huggingface.co/facebook/incoder-6B
* To raise your profile and reputation generally
* The specific publish or perish incentives in academia
* Because you really think you’ve done something interesting and novel, and want to share it with the world
Only the middle one is removed when going to industry.
For FB... It probably helps keep the team valued internally + helps with retention & recruiting. For PhD trained types, this kind of paper is almost table stakes.
Less obvious... FB has been laying off teams like this despite productivity ROI intuitions, so if I was there, I'd be careful to quantify current + future ROI - I'm sure there are key #'s not being shared.
A great example is https://www.uber.com/blog/research/keeping-master-green-at-s...
I think maybe Gitlab Merge Trains predate it, but it was definitely influential.
(disclaimer: this is AI generated, but grounded on contents of papers, with real references, so I'd say it is still constructive)
The state-of-the-art in code generation has seen significant advancements with the deployment of large language models (LLMs) in various code authoring tools. One such example is the study on GitHub Copilot, Amazon CodeWhisperer, and ChatGPT [1], which evaluates the code quality of these AI-assisted code generation tools. The study reveals that ChatGPT generates correct code 65.2% of the time, while GitHub Copilot and Amazon CodeWhisperer achieve 46.3% and 31.1% correctness, respectively. These results indicate that LLMs have made substantial progress in generating high-quality code, but there is still room for improvement.
Other research in the field has explored various techniques to enhance code generation and assistance. For instance, RepoCoder [2] focuses on repository-level code completion by integrating code generation and retrieval models in an iterative paradigm. This approach considers the repository-level context, including customized information such as API definitions and identifier names, to improve code completion suggestions. Serenity [3] leverages library-based Python code analysis for code completion and automated machine learning. The authors explore the potential of data flow analysis produced by Serenity to improve code completion when combined with neural models.
In addition to these advancements, the field has seen progress in incorporating contextual information into code completion models. The paper on enriching source code with contextual data [4] investigates the impact of incorporating contextual information on the performance of code completion models. The authors conduct an empirical study to analyze the effectiveness of this approach. These achievements, along with the advancements in LLMs, contribute to the ongoing progress in code generation and assistance. As the field continues to evolve, it is expected that AI-assisted tools will become increasingly sophisticated and effective in assisting developers with various aspects of the software development process.
[1] Evaluating the Code Quality of AI-Assisted Code Generation Tools: An Empirical Study on GitHub Copilot, Amazon CodeWhisperer, and ChatGPT - 2023: https://arxiv.org/abs/2304.10778
[2] RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation - 2023: https://arxiv.org/abs/2303.12570
[3] Serenity: Library Based Python Code Analysis for Code Completion and Automated Machine Learning - 2023: https://arxiv.org/abs/2301.05108
[4] Enriching Source Code with Contextual Data for Code Completion Models: An Empirical Study - 2023: https://arxiv.org/abs/2304.12269