Copilot is by design trying to give you something that _looks_ correct without caring whether it actually is - so it optimises for real-looking but subtly buggy code, which is the worst kind of broken code.
Copilot is by design trying to give you something that _looks_ correct without caring whether it actually is - so it optimises for real-looking but subtly buggy code, which is the worst kind of broken code.
So we pivoted the product to being something you would run on full auto, for situations where you didn't need a high level of quality. I'm not sure if that option is available to programmers, though.
I wrote a little about this shift to what I think of as “conversational programming”: https://jessmart.in/articles/copilot
Have the machine notify you when it thinks you’ve made a mistake.
Humorously, this is a similar problem to the one autonomous driving has. Being alert when something goes wrong randomly is more difficult than being alert all of the time.
Personally I think the more interesting angle is the trolley problem this creates. People will die in self-driving car accidents and bugs will exist in AI generated code. Those people and bugs are different than the people who will die in human caused accidents and the bugs in human written code. If the number and severity of the results are lessened by the computer, are we willing to forgive the damage directly caused by the AI that falls short of perfection?
I don't think that is the goal just like the goal of the current generation of self-driving cars isn't for you to be able to take a nap in the driver's seat.
Imagine you need some code that would have traditionally taken you and hour to write. I believe the goal of Copilot is to generate the code for you as a starting point. Maybe you don't understand that code immediately and it takes you 20 minutes to figure out what is going on. Then you spend another 20 minutes tweaking it for your exact purpose. If that results in code of similar quality to what you would have written alone, then Copilot makes you more efficient by saving you 20 minutes.
1) Understanding the data model and logic of the code that interacts with the component I'm working on
2) Refactoring existing code to accommodate my change gracefully
3) Writing and fixing tests
4) Working through the code review process
For a major new piece of functionality, add
5) Put together a design document and review it with relevant stakeholders
The part that is fast is actually writing the code, as once I've done steps 1 and 2 (and sometimes 5) writing the new code itself is near trivial. I don't see how copilot could possibly help me in a meaningful way on these kinds of tasks.
The work that seems most amenable to copilot help is things like utility functions for transforming data/calculating things from it, as in the "Easter" example from the article. But here I would rather use a well-tested library, or if one doesn't exist (or I can't use it), write well documented code that I understand thoroughly.
Put another way, the work that copilot seems most adept at is "junior developer" work performed by people operating at a junior level. But if they delegate "figuring things out" to copilot, they're just going to spend way more time in code review. Or worse, they're not going to spend that time, and will learn nothing/stagnate in their professional progression.
Ever since the advent of satellite nav I've become terrible at learning my way around cities. I'm okay with the loss, since I can generally rely on having nav when I need it, and navigating cities isn't one of my core responsibilities. Copilot is not reliable (it won't answer your question every time), and it automates something that is your actual job. A junior dev might be better served by spending the extra 20 minutes muddling through and building their skillset.
I think the issue is that the MVP from a customer perspective is, effectively, being able to take a nap in the driver's seat. From a research perspective there are obviously intermediate milestones, but that doesn't make it fit for what people would want to use it for. Same goes for Copilot.
Maybe that is a requirement for some users, but it isn't a universal one. Plenty of people see a benefit in assistive technology that isn't complete such as adaptive cruise control or boilerplate/scaffolding dev tools.
It also raises the ethical question of whether these creators are responsible for the misuse of their products. Is it enough for them to say "This is how this product should be used. You are on your own if you use it outside these settings."? Holding developers responsible for the misuse of their software could create an actual slippery slope. Where is the line drawn? Do we start punishing people who create encryption algorithms because someone used the encryption to hide evidence of a crime?
I don't think you have to answer the ethical question to address the level of readiness that Copilot or self-driving cars are at. It definitely raises the question, but you don't have to answer it to talk about suitability for use cases.
As you say, it might address the requirements of some specific people. My argument is that Copilot is not good enough yet for the bulk of imagined use cases, whether or not you call that MVP, and I think the post makes a good argument about why.
Better code should largely be easier to understand.
By contrast, Copilot doesn't necessarily have any idea what you're trying to do. So it can, to an approximation, pattern match on what you've already written, and spit out valid code that is "inspired" by things it's seen in the past. But it doesn't actually know what you're trying to do. It doesn't know what your acceptance criteria are, or what invariants you're trying to maintain, or anything like that. And, at least in the places I've worked, most the interesting bugs (by which I mean, the ones that managed to cause trouble in production) happen when the programmer writing the code didn't have a firm idea of what they were trying to do. So, that's what worries me - I would fear that the spots where Copilot can't even theoretically be expected to do a good job happens to be exactly the kind of things for which people would tend to rely on it the most.
Maybe I'm being overly pessimistic? But that's kind of my job - I work in an area where "move fast and break things" is pretty antithetical. But it would still be a lot more compelling to me if I could see a paper demonstrating that a team using Copilot has fewer production defects than a team that's doing exactly the same work but without Copilot. Or alternatively, if it were repackaged as something that's a bit like a smarter version of IDE refactorings. "Hey, it looks like you're about to spit out a big old mess of boilerplate. Let us get that for you." Or, "Hey, some functions you called can fail, how about I go ahead and suggest a catch block so you don't forget to write one?" Basically, give me something that's a bit more smart cruise control and a bit less Autopilot.
This is an extremely good analogy -- in both situations, the human will become lazy and stop paying attention (regardless of whether they're supposed to keep their hands on the wheel, literally or metaphorically), and it will be possible to have a net result worse than either human or AI acting alone.
What's the point of striving to write better, more correct code, being a safer driver, if all we ever do is rely on the status quo to train models to be average?
There is a uniqueness requirement, but it has nothing to do with length. A unique one liner would have a copyright.
Oh this is going to make teaching intro to computer science sooooo 'interesting'.
It wouldn't be so bad if the students looked at the generated code and understood it, but experience tells me most of them will not.
Of course it will not write complete functions correctly.
It looks to you like it should work, but it doesn't, and you can't figure out why.
That's not "mostly working," that's a frustrating waste of time. It's hard enough to notice when you accidentally swap `i` and `j` -- why would you want to make your life even more miserable by spending your time finding all of the instances where a pattern matching robot has done something similar in an unfamiliar block?
And if you do happen to get "mostly working" code, but only want it to stay together long enough for you to fundraise, you're basically stating that you plan on foisting this technical debt onto the poor sod you happen to hire.
Attitudes like yours are the reason this dogpile scares me.
If that's the case, I agree with your assessment, that GitHub Copilot isn't delivering on its promise and I would not be using it.
My understanding, however, was the GitHub Copilot does produce functioning code. If you're saying, as I think you are, "No, GitHub is lying about Copilot." I find that claim fascinating (How the hell did we get to the point where a software company could release a product that literally does not do even the most basic version of what it says it does, and only a few people notice?), but I'd need more specific information from you before I'd believe it.
All I really need is for the product to work well enough that I can fundraise and hire someone who's better at programming than I am, someone who hopefully doesn't write comments about how unimportant other people's work is on HN.
GitHub Copilot, if it can create "mostly working" code, is a huge value-add to an early business trying to find product/market fit because it cuts down on the time it takes to prototype and generate early versions.
Not every piece of software has to work perfectly every time, and I don't think that's the standard to which we should hold any automated coding tool.
I could see Copilot used in such a way. I think the interaction would have to change though: force the user to give it the tests as input, not give it some basic instruction, have it generate code, and then I try to write tests after. The tests should be the spec that Copilot uses to generate its output.
Right now, I'm not excited about Copilot. Like you say, understanding what Copilot spits out is difficult and I suspect more error prone than just writing it yourself (since we often see what we want to see and can overlook even glaring mistakes). I'm also not excited about them ignoring the licenses of the code they trained on. But I can imagine a future iteration that generated code to pass some tests that I could get excited about.
In this exploration stage, total correctness doesn't matter since you're just getting a feel for the data. Copilot might help a lot with the associated boilerplate.
> The programmer does not even have to be exact in his own ideas‑he may have a range of acceptable computer answers in mind and may be content if the computer's answers do not step out of this range. The programmer does not have to fixate the computer with particular processes. In a range of uncertainty he may ask the computer to generate new procedures, or he may recommend rules of selection and give the computer advice about which choices to make. Thus, computers do not have to be programmed with extremely clear and precise formulations of what is to be executed, or how to do it.
From: https://web.media.mit.edu/~minsky/papers/Why%20programming%2...
Having offline coding interviews to find Software Engineers will become even more important.
The funny part will be when all the human programmers who steal code will get doxed as a side effect. It shine a light on lots of skeletons in the closet.
do not ignore the elephant in the room. copilot is stealing code from projects with open licenses.
Developers who not use this (or similar tools) will not be hired, or only in particular niche domains where correctness matters.