Coding on Copilot: Data suggests downward pressure on code quality
gitclear.com
gitclear.com
Back to code solutions, I never ever use them if I don't understand ALL that is happening in that code. I know the bot will ignore many corner cases or that it doesn't know some subtle implications about the context where I'll deploy the code.
I think someone that doesn't care about code quality was already copying and pasting code from StackOverflow and fisting the keyboard until it works then moving to the next ticket. I really don't think AI is making things worse in that regard.
- Making code more accessible mean people with lower skills can contribute.
- Increased in productivity increased expectations for deadlines and so there is less time for quality.
- 2023 was a year with lower code quality on average for everybody.
- Easier code to produce meant more code produced, and maybe it dwarfed the absolute amount of quality code, giving this impression, while the relative amount didn't change but we don't have their way of measuring good code.
Etc
True.
It does, however, correspond with my observations of devs at my workplace. Those who have been using copilot (and the like) have had fairly obvious drops in the quality of their work.
If the cost of updating code has dropped due to AI, is code churn still a good proxy for low code quality?
Imagine hiring a junior to write your code, except they’d just stop existing when they were done.
You’d need to ask a whole new junior to explain it to you later (if you were relying on AI)
At the recent startup I was at, the CEO (who was also an engineer) basically relied on chatgpt to make the whole app.
It was a complicated and fragile codebase. The docstrings, which were also AI generated, were frequently out of date if not just wrong.
We were in the midst of rewriting it before we went bust.
It only looks for brand new code that was immediately reverted/rewritten, and also looks for code that was copied from the internet vs. from within the repo (the former is probably non-idiomatic for the code base in question).
Basically, the AI correlation is baseless.
I don't know. But I'm pretty sure 50% reduction results in project failure and company bankruptcy.
I worked as a high school bureaucrat for several years, and through some simple scripts, I was able to make some tedious data entry tasks vastly faster. It was ugly, hacked together code, with lots of hardcoded values I had to update each semester, but it worked and was way better than doing it all by hand. Low quality code is worse than good quality code, but often better than manual labor.
I want my web browser and my bank to use high quality code, but I want semi-technical people who would otherwise do everything manually to be able to automate tedious portions of their jobs.
(The caveat, of course, being that a buggy script can screw up the thing the thing it's automating must faster than a manual process as well.)
Anecdotally, as a principal engineer, I’ve definitely noticed that new senior engineers on the team that say they are using chatgpt/copilot produce unprecedentedly bad code at unprecedented rates.
It takes me 2-3x longer to unwind such crap than it would for me to write it from scratch.
As we grow the team, this will definitely put us out of business unless we find a way to fix it.
Currently, we’re hoping the AI assisted engineers will get better at unborking code before merging it, but that’s a harder task than RTFM or going to stack overflow to copy-paste.
Note I'm not even strictly speaking criticizing the quality of the output per se. It is also a big jump over any previous technology and very impressive in its own way.
It is, nevertheless, quite dangerous because the jump in the human-perceived plausibility is much larger than the quality improvement.
Whereas earlier techs were obviously wrong to a human reader, in the case of code generation so obviously wrong that we never even considered using them, LLMs are extremely good at hiding the errors in the parts of the code that we are cognitively most inclined to overlook. This also has the effect of making it bizarrely difficult code to fix.
How it does this I do not know. A fascinating research question for some ambitious cognitive scientist. But the signal is very strong and I don't need to wait for a paper to come out to see it.
I do not think this is fundamental to AI. As I like to remind people, LLMs are not the whole of AI. They're just one technique, and one that partially for the very reason I discuss in this post, one I expect to eventually become a part of a larger system that can fix this problem at some higher level. I expect people to someday look back and laugh at us for thinking that LLMs could be used for all the things we think they can be used for. But the reasons they will be laughing are the very experience we're gathering now, and there's no skipping that phase.
Then: Outside of some extra complex or absurdly simple case it is very often harder to write tests which truly and fully test your code then it is to write the code correct.
In my experience often correct code is a product of carefully written test (which still are in reality imperfect), static code analysis (can be the type system, or external tools) applied to carefully written code and a proper code review.
So if you bring both of it together you have:
- AI supported code which is likely to contain bugs which are really easy to overlook in reviews
- AI supported test code which is has the same issue, i.e. they have gaps which are really likely to overlook by reviewers.
- more code due to less reuse and it also sometimes being easier to generate instead of use a library leading to more code review needing to be done and in turn more time pressure and less quality review
so put together: more bugs which are hard to find with test which are more likely subtle pass even with bugs and less time for proper reviews
So does it pass the test? Yes, but it was AI written too so can it be trusted?
From what I've seen, many early career engineers are the ones using code assist tools. Many early career engineers are often placed on lower stakes front end focused teams as well. This is mostly anecdotal data from non-traditional cs background engs that I know.
That said, I've done no LLM coding and don't have a feel for how it goes wrong.
Code isn't, yes, but LLMs are. That's part of the key here, coding can in some sense now be done at the prompt level along with an acceptance test suite, so rather than coding solution and test suite, you can hypothetically just write the test suite and have the LLM generated the code part.
Of course it helps, writing acceptance tests based on a spec is a lot easier than writing an implementation that passes those tests.
I think this is a bad take on this issue. It implies that if either chatgpt or the devs were just "smarter" this would not happen. This also implies that churn is a bad thing. And I don't think neither of these are true.
The real problem is that debugging is twice as hard as programming and now for the first time we got to the point that people can create software that "works" but they cannot fully grasp why exactly it works.
Of course you can go with the easy answer (these people are just lazy/dumb and should not be in this field) but this doesn't help. But the reality is that we don't really have anything better than printf and gdb to understand what the is happening inside our computers.
Hopefully now we will have more incentives to create better tools to help our understanding of what happens under the hood.
It seems to me that this is the real problem. The difference between a sub-par programmer and a competent one is that the latter can create code that they understand and can therefore more easily debug. Being able to create code that "works", but you then have to spend hours/days trying to grok in order to debug it, seems to me to be more harmful than productive.
The huge amount of times I saw
//This makes everything work somehow
sleep(1)
begs to differNow you have that 6 line script you wanted.
Honestly it is still pretty useful though, as you can see 'your code in different clothing styles' in some sense, very quickly to choose the best/most performant option. You just have to know what that looks like before going in, and it helps make it 'tangible' quicker.
Also, my current job limits the use of AI coding tools, because of the risk of our IP leaking out. To really be useful, there is a lot of context a programmer should be able to collate. Not all of it is going to be in the repo.
And AI gives you fast code production speed so it seems to be a fix.
But due to less structured code and code reuse it needs much more code for the same value produced. And it loves to produce the kind of bugs which are easy to pass reviews and tests (because this are the kind of bugs it most likely has as "correct" code in it's training data). Which means less time per-value for human code reviews and in turn lower quality results on many levels.
I originally hoped that with CS/SoftwareDev getting older and less explosive grows and Universities more realizing the (sometimes huge) gap between CS and SoftwareDev the industry would get a chance of consolidating tech and dev processes to fix that issue. But with the rise of LLMs this hope might be further away then it ever had been in the last 20 year.
just having more test doesn't mean you actually test more (or more correctly)
Having AI specialized on assisting testing is viable but in the end would be done quite different as far as I can tell even if it uses LLMs. Through maybe domain adoption and a very different scaffolding might be enough.
One time I had to use the Go protobuff reflection API, and Copilot was actual crucial because the docs are so bad.
Can’t use it in my current role and I feel less productive.