That is to say, you have code prompts here, let Copilot fill in the gaps, and rate that code. Is there a study that uses the same prompts with a selection of programmers to see if they do better or worse?
I'm curious because in my testing of copilot, it often writes garbage. But if I'm being honest, often, so do I.
I feel like Twitter's full of cheap shots against copilot's bad outputs, but many of them don't seem to be any worse than common errors. I would really like to see how copilot stands up to the existing human competition, especially on axes of security, which are a bit more objectively measurable than general "quality".
Nonetheless, we think that simply having a quantification of Copilot's outputs is useful, as it can definitely provide an indicator of how risky it might be to provide the tool to an inexperienced developer that might be tempted to accept every suggestion.
The reason I singled out juniors has nothing to do with them being the least likely to check documentation. In fact seniors are just as bad in that regard. Plus a lot of the time the reason an engineer goes to SO is a particular common problem doesn’t have a pre-built solution in a languages standard library. The reason I suggested juniors is just because that’s who the researchers said they mostly work with.
So let’s be clear about one thing: I know plenty of junior developers who have put plenty of seniors to shame. I’m not about to suggest that juniors are worse developers; maybe less experienced in the chronological sense but you need to be damn careful before making other generalisations.
Agree. My comment was more about: most of the engineers out there (and yes, I do include myself) tend to rely and trust on information that is easily accessible; no matter if they juniors or seniors. I think this also applies to everyday life in general as well (e.g., we tend to read the newspapers to get "informed", but we rarely go to the source of truth to check our "facts").
One suggestion: On Arxiv's "Code & Data" tab it says "No official code found", but it looks like you have shared your code here: https://zenodo.org/record/5225651#.YSRKBi2cbyU
It would be a good idea to get that linked on the Arxiv page, 'cause the first question I had was, "Can I reproduce the findings?"
However, (my opinion only follows) I think our paper shows that there is a danger for Copilot to suggest insecure code - and inexperienced / security non-aware developers may accept these suggestions without understanding the implications, whereas if they had to write the code from scratch then they might (?) not make the mistakes (as they need to put in more effort, meaning there might be a higher chance they stumble upon the right approach - e.g. if they ask an experienced developer for help).
Perhaps you could take a similar approach as [1] and leverage MOOC participants?
> My main question when reading is how the results compared to manually-written code.
Ah, this is exactly the question. But as you say, much harder to answer. Even if you run a competition, unless you can encourage a wide range of developers to enter, you won't be getting the real value. Instead you might be getting incidence rates of code written by students/interns. Perhaps if you could get a few FAANGs on board to either share internal data (unlikely) or send a random sample of employees (also very unlikely) to make teams and then evaluate their code... It seems like a difficult question to answer.
We think a more doable way would be to take snapshots of large open source codebases (e.g. off GitHub) and measure the incidence rate of CWEs, but this also presents its own challenges with analyzing the data. Also, what's the relationship between open source code and all code?
Lots of avenues to consider.
char str_a[20], str_b[20], str_c[20];
sprintf(str_a, "%f", a);
sprintf(str_b, "%f", b);
sprintf(str_c, "%f", c);
[1] https://github.com/google-research/arxiv-latex-cleaner