Risk Assessment of GitHub Copilot
gist.github.com
gist.github.com
Copilot is by design trying to give you something that _looks_ correct without caring whether it actually is - so it optimises for real-looking but subtly buggy code, which is the worst kind of broken code.
So we pivoted the product to being something you would run on full auto, for situations where you didn't need a high level of quality. I'm not sure if that option is available to programmers, though.
I wrote a little about this shift to what I think of as “conversational programming”: https://jessmart.in/articles/copilot
Have the machine notify you when it thinks you’ve made a mistake.
Humorously, this is a similar problem to the one autonomous driving has. Being alert when something goes wrong randomly is more difficult than being alert all of the time.
Personally I think the more interesting angle is the trolley problem this creates. People will die in self-driving car accidents and bugs will exist in AI generated code. Those people and bugs are different than the people who will die in human caused accidents and the bugs in human written code. If the number and severity of the results are lessened by the computer, are we willing to forgive the damage directly caused by the AI that falls short of perfection?
I don't think that is the goal just like the goal of the current generation of self-driving cars isn't for you to be able to take a nap in the driver's seat.
Imagine you need some code that would have traditionally taken you and hour to write. I believe the goal of Copilot is to generate the code for you as a starting point. Maybe you don't understand that code immediately and it takes you 20 minutes to figure out what is going on. Then you spend another 20 minutes tweaking it for your exact purpose. If that results in code of similar quality to what you would have written alone, then Copilot makes you more efficient by saving you 20 minutes.
1) Understanding the data model and logic of the code that interacts with the component I'm working on
2) Refactoring existing code to accommodate my change gracefully
3) Writing and fixing tests
4) Working through the code review process
For a major new piece of functionality, add
5) Put together a design document and review it with relevant stakeholders
The part that is fast is actually writing the code, as once I've done steps 1 and 2 (and sometimes 5) writing the new code itself is near trivial. I don't see how copilot could possibly help me in a meaningful way on these kinds of tasks.
The work that seems most amenable to copilot help is things like utility functions for transforming data/calculating things from it, as in the "Easter" example from the article. But here I would rather use a well-tested library, or if one doesn't exist (or I can't use it), write well documented code that I understand thoroughly.
Put another way, the work that copilot seems most adept at is "junior developer" work performed by people operating at a junior level. But if they delegate "figuring things out" to copilot, they're just going to spend way more time in code review. Or worse, they're not going to spend that time, and will learn nothing/stagnate in their professional progression.
Ever since the advent of satellite nav I've become terrible at learning my way around cities. I'm okay with the loss, since I can generally rely on having nav when I need it, and navigating cities isn't one of my core responsibilities. Copilot is not reliable (it won't answer your question every time), and it automates something that is your actual job. A junior dev might be better served by spending the extra 20 minutes muddling through and building their skillset.
I think the issue is that the MVP from a customer perspective is, effectively, being able to take a nap in the driver's seat. From a research perspective there are obviously intermediate milestones, but that doesn't make it fit for what people would want to use it for. Same goes for Copilot.
Maybe that is a requirement for some users, but it isn't a universal one. Plenty of people see a benefit in assistive technology that isn't complete such as adaptive cruise control or boilerplate/scaffolding dev tools.
It also raises the ethical question of whether these creators are responsible for the misuse of their products. Is it enough for them to say "This is how this product should be used. You are on your own if you use it outside these settings."? Holding developers responsible for the misuse of their software could create an actual slippery slope. Where is the line drawn? Do we start punishing people who create encryption algorithms because someone used the encryption to hide evidence of a crime?
I don't think you have to answer the ethical question to address the level of readiness that Copilot or self-driving cars are at. It definitely raises the question, but you don't have to answer it to talk about suitability for use cases.
As you say, it might address the requirements of some specific people. My argument is that Copilot is not good enough yet for the bulk of imagined use cases, whether or not you call that MVP, and I think the post makes a good argument about why.
Better code should largely be easier to understand.
By contrast, Copilot doesn't necessarily have any idea what you're trying to do. So it can, to an approximation, pattern match on what you've already written, and spit out valid code that is "inspired" by things it's seen in the past. But it doesn't actually know what you're trying to do. It doesn't know what your acceptance criteria are, or what invariants you're trying to maintain, or anything like that. And, at least in the places I've worked, most the interesting bugs (by which I mean, the ones that managed to cause trouble in production) happen when the programmer writing the code didn't have a firm idea of what they were trying to do. So, that's what worries me - I would fear that the spots where Copilot can't even theoretically be expected to do a good job happens to be exactly the kind of things for which people would tend to rely on it the most.
Maybe I'm being overly pessimistic? But that's kind of my job - I work in an area where "move fast and break things" is pretty antithetical. But it would still be a lot more compelling to me if I could see a paper demonstrating that a team using Copilot has fewer production defects than a team that's doing exactly the same work but without Copilot. Or alternatively, if it were repackaged as something that's a bit like a smarter version of IDE refactorings. "Hey, it looks like you're about to spit out a big old mess of boilerplate. Let us get that for you." Or, "Hey, some functions you called can fail, how about I go ahead and suggest a catch block so you don't forget to write one?" Basically, give me something that's a bit more smart cruise control and a bit less Autopilot.
This is an extremely good analogy -- in both situations, the human will become lazy and stop paying attention (regardless of whether they're supposed to keep their hands on the wheel, literally or metaphorically), and it will be possible to have a net result worse than either human or AI acting alone.
What's the point of striving to write better, more correct code, being a safer driver, if all we ever do is rely on the status quo to train models to be average?
There is a uniqueness requirement, but it has nothing to do with length. A unique one liner would have a copyright.
Oh this is going to make teaching intro to computer science sooooo 'interesting'.
It wouldn't be so bad if the students looked at the generated code and understood it, but experience tells me most of them will not.
Of course it will not write complete functions correctly.
It looks to you like it should work, but it doesn't, and you can't figure out why.
That's not "mostly working," that's a frustrating waste of time. It's hard enough to notice when you accidentally swap `i` and `j` -- why would you want to make your life even more miserable by spending your time finding all of the instances where a pattern matching robot has done something similar in an unfamiliar block?
And if you do happen to get "mostly working" code, but only want it to stay together long enough for you to fundraise, you're basically stating that you plan on foisting this technical debt onto the poor sod you happen to hire.
Attitudes like yours are the reason this dogpile scares me.
If that's the case, I agree with your assessment, that GitHub Copilot isn't delivering on its promise and I would not be using it.
My understanding, however, was the GitHub Copilot does produce functioning code. If you're saying, as I think you are, "No, GitHub is lying about Copilot." I find that claim fascinating (How the hell did we get to the point where a software company could release a product that literally does not do even the most basic version of what it says it does, and only a few people notice?), but I'd need more specific information from you before I'd believe it.
All I really need is for the product to work well enough that I can fundraise and hire someone who's better at programming than I am, someone who hopefully doesn't write comments about how unimportant other people's work is on HN.
GitHub Copilot, if it can create "mostly working" code, is a huge value-add to an early business trying to find product/market fit because it cuts down on the time it takes to prototype and generate early versions.
Not every piece of software has to work perfectly every time, and I don't think that's the standard to which we should hold any automated coding tool.
I could see Copilot used in such a way. I think the interaction would have to change though: force the user to give it the tests as input, not give it some basic instruction, have it generate code, and then I try to write tests after. The tests should be the spec that Copilot uses to generate its output.
Right now, I'm not excited about Copilot. Like you say, understanding what Copilot spits out is difficult and I suspect more error prone than just writing it yourself (since we often see what we want to see and can overlook even glaring mistakes). I'm also not excited about them ignoring the licenses of the code they trained on. But I can imagine a future iteration that generated code to pass some tests that I could get excited about.
In this exploration stage, total correctness doesn't matter since you're just getting a feel for the data. Copilot might help a lot with the associated boilerplate.
> The programmer does not even have to be exact in his own ideas‑he may have a range of acceptable computer answers in mind and may be content if the computer's answers do not step out of this range. The programmer does not have to fixate the computer with particular processes. In a range of uncertainty he may ask the computer to generate new procedures, or he may recommend rules of selection and give the computer advice about which choices to make. Thus, computers do not have to be programmed with extremely clear and precise formulations of what is to be executed, or how to do it.
From: https://web.media.mit.edu/~minsky/papers/Why%20programming%2...
Having offline coding interviews to find Software Engineers will become even more important.
The funny part will be when all the human programmers who steal code will get doxed as a side effect. It shine a light on lots of skeletons in the closet.
do not ignore the elephant in the room. copilot is stealing code from projects with open licenses.
Developers who not use this (or similar tools) will not be hired, or only in particular niche domains where correctness matters.
You're basically asking a robot that stayed up all night reading a billion lines of questionable source code to go on a massive LSD trip and then use the resulting fever dream to fill in your for loops.
Coming from a hardware background where you often spend 2-8x of your time and money on verification vs. on the actual design, it seems obvious to me that Copilot as implemented today will either not provide any value (best case), will be a net negative (middling case), or will be a net negative, but you won't realize that you've surrounded yourself with a minefield for a few years (worst case).
Having an "autocomplete" that can suggest more lines of code isn't better, it's worse. You still have to read the result, figure out what it's doing, and figure out why it will or will not work. Figuring out that it won't work could be relatively straightforward, as it is today with normal "here's a list of methods" autocomplete. Or it could be spectacularly difficult, as it would be when Copilot decides to regurgitate "fast inverse square root" but with different constants. Do you really think you're going to be able to decipher and debug code like that repeatedly when you're tired? When it's a subtly broken block of code rather than a famous example?
That Easter example looks horrific, but I can absolutely see a tired developer saying "fuck it" and committing it at the end of the day, fully intending to check it later, and then either forgetting or hoping that it won't be a problem rather than ruining the next morning by attempting to look at it again.
I can't imagine ever using it, but I worry about new grads and junior developers thinking that they need to use crap like this because some thought leader praises it as the newest best practice. We already have too much modern development methodology bullshit that takes endless effort to stomp out, but this has the potential to be exceptionally disastrous.
I can't help but think that the product itself must be a PSYOP-like attempt to gaslight the entire industry. It seems so obvious to me that people are going to commit more broken code via Copilot than ever before.
Why do that when you can just train a GPT-3 model on public repositories and call it a day?
I'm curious what it spills out for things like "Todo", or "this is probably broken", etc.
Which is why, based on Windows state, it will never come out of Microsoft.
A bare minimum baseline validation check for Copilot would be to see if it provides you code which won't compile in-context. If it will, then that means it's not even taking into account well-specified domain model of your chosen programming language's semantics. Which, upon satisfaction, is still miles away from taking into account the domain of your actual problem that you're using software to solve.
The only place where the approach taken, as-is, makes sense to me is for truly rote boilerplate code. However, that then begs the question... how is this machine learning approach more effective than a targeted heuristic approach already taken by existing IDE tooling, etc.?
FWIW, I don't think any of this is lost on GitHub. I think Copilot is more likely a tremendously marketable half-step and small piece of a larger longer-term strategy unfolding at Microsoft/GitHub to leverage an incredible asset they're holding, i.e... practically everybody's source code. The combination of detailed changelogs, CI results (e.g. GitHub actions), Copilot, and a couple other key pieces makes for a pretty incredible basis for reinforcement learning to multiple ends.
Maybe we should use Copilot to commit more open source code meaning that Copilot becomes more and more corrupted and unusable!
of course then we end up with a bunch of bad open source code that will turn people off of using open source.
Gee, I don't think Microsoft really thought this one through.
The expectation is entirely different than producing code. Code needs to be correct, secure, performant, and readable. Failure on any of those fronts can be expensive to disastrous. Nobody can reasonably expect a test suite to catch every bug, even if created by the smartest humans. If a copilot-created test does prevent a bug from shipping it provides immediate value. I could see it coming up with some whacky-but-useful test cases that a sane person might not consider. From a training perspective I would think that assertion descriptions contain more consistent lexical value than the average function signature.
It seems like the ambitious data scientists, product marketers, and managers fell in love with a revolutionary idea about AI writing code, and neglected to consult the engineers they are trying to ‘augment’.
Good tests are documentation that a computer can verify. Because they explain the meaning of parts of the system, they contain information not available in the code. If you try using ML for test generation, you'll have the same problem you do with GPT-3 prose: it might look plausible at first glance, but lacks coherent meaning.
You'd also end up with one of the problems common in big test suites: poorly factored tests that end up being the sort of expressive duplication that is a giant drag on improving existing code. ML is nowhere near advanced enough to say, "Gosh, we're doing the same sort of test setup a bunch; let's extract that into a fixture, and then let's unify some fixtures into an ObjectMother.
For people looking to get the computer to do the work of catching more things with less burdensome test writing, I suggest taking a look at things like Hypothesis: https://hypothesis.readthedocs.io/en/latest/
> If you try using ML for test generation, you'll have the same problem you do with GPT-3 prose: it might look plausible at first glance, but lacks coherent meaning.
There is a company in this space of generating "plausible tests" for legacy code bases at very large enterprises (think Goldman Sachs, telcos etc) called Diffblue [0].
They raised funding back in 2017 [1] and it seems their biggest value-add is in creating unit tests for legacy Java code bases that often have little to no unit tests.
Essentially, these AI generated unit tests help a team "document" all known the behaviors of a legacy code base such that when a change is introduced that violates the behaviors covered by the generated unit tests, the tool can alert the team of the potential presence of a regression.
Anyway, they offer a fairly basic browser-based demo of their AI product called Diffblue Cover [2].
Are you aware of them?
I feel like you just described every developer/codebase where mock testing is stupidly enforced. Where every single unit test mocks every single indirect object. 98% of the testing code is just exhaustive setup and teardown of objects not being tested by each test, and then a bunch of conditional checks to ensure that every deeper/indirect method is being called exactly the right number of times with exactly the right arguments and returning exactly the right value. Almost all of the test code is just hacking mock objects. The actual purpose of each test is buried so deep that it's impossible to even understand the business logic being applied.
I hate evangelical "mock testers" with a passion.
GPT is not "nothing more than a random number generator"
Though I don't fully disagree with it, though 'nothing more' is a bit too strong. The author of a GPT3 written comment like the one here where the prompt was pretty much just the thread is pretty much just the RNG. The language model makes the random choice draw from the distribution of plausible texts, and the RNG picks the output.
GPT3 could have written your comment-- if only it drew the right random numbers.
What RNG? It definitely doesn't randomly pick words. If the comment I responded to was written by a bot (is that legal? Can I report that?) then it's indistinguishable from a human written comment.
Exclusively selecting the most likely symbol produces pathological behavior outside of extremely short output.
What caused GPT3 to output its comment rather than yours is a product of its random choices. There is a set of choices it could have made which would have caused it to output your comment. You can see this property employed by the GPT2 text compressor: https://bellard.org/libnc/gpt2tc.html to compress text it just writes down the choices, using an entropy coder to represent likely choices with fewer bits.
I assume copilot is the same general structure as GPT-- just trained on different data.
And yes, the comment you responded to was written entirely by GPT3 (with some number of retries and trims). As it said-- it's "easy to mistake it for a human's work". :) There is nothing illegal about it, but I suppose HN would prefer that there be enough human supervision of bot comments such that they're limited to contexts where they are funny/insightful. :P
Is the same not true of human writers? Are human writers deterministic?
Even if that happened, which I am not expecting, I think the need is much more easily solved via means that are simpler and more effective. E.g., a good tester writing up a list of things they test about APIs: https://www.sisense.com/blog/rest-api-testing-strategy-what-...
Crucially, that's not what copilot is.
Maybe Copilot 2 will do exactly this; it will generate tests based of half working code, run them and suggest improvements, that would increase productivity by like ~100%, but to me this sounds to good to be true.
If Copilot can't write the correct code in the first place, you really shouldn't expect a proper test to be written by Copilot.
> Code needs to be correct, secure, performant, and readable.
Most tests should also have at least three of those attributes. Nobody actually wants their tests to be incorrect, slow, or impossible to understand or modify.
"Oh, and they're both wrong."
"Both look plausibly correct at a glance"
You would end up with tests that look plausibly correct, but test the wrong results.
1. Smart Code: Code that you honestly have to think about while you're writing. You write this code slowly and carefully, to make sure that it does what you need it to
2. Dumb Code: This is trivial code, like adding a button to a screen. This is code you really don't have to think about, because you already know exactly how to implement it. At this point your biggest obstacle is how fast can your fingers type on a keyboard.
For me Github Copilot is useless for "Smart Code" but a godsend when writing "Dumb Code". I want to focus more on writing and figuring out the "Smart Code", if I need to throw a form together in HTML or make a trivial helper function, I will gladly let AI take over and do that work for me.
UX is probably the most important aspect of most software products. Every software product is either "smart code" or "smart ux". No one pays much for "dumb code with bad UX" except in dysfunctional markets.
Adding a button to a screen should be trivial, and if it's not you need better tools. (As in "a not-horribly-misdesigned language and framework", not as in "giant transformer".)
Deciding where to add the button, its shape, its size, what happens when it's clicked, the text on the button, ... is anything but trivial.
This is true, but actually tedious glue code is often non-trivial. For example, in one of my hobby projects I have a repo where I have to write a lot of glue code to schlepp data from a CSV file format into an existing database. Doing this correctly requires reading through the (lengthy) documentation for the format of both the CSV data and also the system that ingests the database, since there are a bunch invariants about the key tables and columns that aren't enforceable in SQL (and obviously not enforceable in the CSV).
This is the sort of glue code/boilerplate where a synthesizer that can understand natural language would be actually helpful.
> but the better languages and frameworks don't exist yet
There are certainly some languages that are less verbose than others.
Java and Go are very boilerplate-y languages. Python is also pretty verbose and inexpressive for certain types of code.
The typescript example on the copilot page right now is a prefect example of "ugh just use a better language".
The examples of boilerplate where copilot shines seem like situations where really simply using "snippets" would work better. E.g., everything on the copilot homepage right now.
And because of the nature of boilerplate, Java has IDEs that will both generate and modify this boilerplate without thought, no AI required.
I don't remember the last time I typed the text 'class' for example. Instead I type "new Foo(someStringVar)", and then hit Alt-Enter and my IDE creates the file, the `class Foo` along with a ctor that takes a String.
In other words, a machine looking at only my own glue would be more likely to mismatch or use the wrong version in any given situation.
<Button ...
or whatever.I imagine there's some situations where Snippets do better, and where Copilot does better, but the more complex the situation the less i trust Snippets... but _also_ the less i trust Copilot.
It seems my trust in Copilot is very similar to that of use cases for Snippets. To throw out fake numbers, it makes me feel like Snippets (and tools like them) cover ~%70 of Copilot's use case. So i'm really curious on knowing what that %30 is, and if it is ever useful.
I suppose there's a (supposed) advantage that it's automatically finding and suggesting the snippet for you, rather than relying on you to think of it and recall the key binding, or know that it's there to such for.
Imho, it is just an argument for making better languages and libraries. (These libraries will also make it easier to use with copilot.)
Once we spot a tedious common pattern, we should be finding ways to DRY it up. Configs, libraries, frameworks, DSL, tools, and languages are all great ways to do that. Copy-pasting and machine-generating code are short-term thinking in two ways: they focus on the initial creation of the code at the expense of maintenance, and they give up on increasing abstraction, lock the system into a productivity plateau.
It is what the action of that button is where the real fun comes in.
I once had a project that was a yes/no dialog. Two buttons and some text. I had the dialog up and running in under an hour. The action that happened when you pressed yes took 3 months to finish.
The risk with WYSIWYG editors isn't that there's some tipping point where it becomes 10% more efficient to write code and you lose a bit of productivity or something. It's that something comes up half way through development and the WYSIWYG doesn't have the feature you need[1], and the entire project slams into a brick wall and dies instantly.
You can prevent this by running into the exact same problem CoPilot has, which is that reading code is harder than writing it. If you try to avoid the brick wall by having devs familiarise themselves with the code as the WYSIWYG generates it, those devs would have just been able to build it themselves in less time and with cleaner code.
[1] which will always happen eventually, because they're balancing the feature set for the exact reasons you mentioned. If they can do everything code can then the UX is going to be so bloated and horrible that it'd be trivially worse to use than just writing code.
If not, why should copilot?
The article shows that I can't trust GitHub copilot. So I don't think it is a representative name. Here, it would be more like a servant.
If people can copy paste the most insecure code from Stack Overflow or random tutorials, they will absolutely use Copilot to "write" code and it will be become the default, especially since it's so incredibly easy to use. Also, it's just the first generation tool if it's kind, imagine what similar products will accomplish in 20 years.
This is a product by a well-known company (GitHub) which is owned by an even more well-known company (Microsoft). GitHub is going to be trusted a lot more than a random poster on Stack Overflow or someone's blog online. And GitHub is explicitly telling new coders to use Copilot to learn a new language:
> Whether you’re working in a new language or framework, or just learning to code, GitHub Copilot can help you find your way. Tackle a bug, or learn how to use a new framework without spending most of your time spelunking through the docs or searching the web.
This is what differentiates Copilot from Stack Overflow or random tutorials. GitHub has a brand that's trusted more than random content creators on the internet. And it's telling new coders to use Copilot to learn things and not check elsewhere.
That's a problem. Doesn't matter what generation of the program it is. It creates unsafe code after using its brand reputation and recognition to convince new coders to not check elsewhere.
> GitHub has a brand that's trusted more
Consider Google Translate, right? Google is a well-known brand that is trusted (outside of a relatively small group of people that doesn't trust Google on principle). Yet every professional translator knows that the text produced by Google Translate is a result of machine translation, Google or no Google. They may marvel at the occasional accuracy, yet expect serious blunders in the text, and would therefore not just trust that translation before submitting it to their clients. They will check. Or at least they should.
Same with programmers.
And as you say, it will be the same with programmers. Who's this being targeted at? People "working in a new language or framework, or just learning to code". The whole value prop is, "You don't have to know what's going on!"
The important difference is that the target readers can usually spot an egregiously bad translation. But the target users for software cannot easily spot gaping security holes and other serious issues until something bad happens.
What, no, that's not true at all, that's like the second biggest problem. GTL routinely does stuff like invert the meaning of clauses, or drop information, or hallucinate absent context. Target readers can't reasonably be expected to catch any of that.
The value proposition is a better tab completion. It's not autopilot.
The bar is substantially lower for a 'programmer', especially with an incredibly large bootcamp market which churn out 'professional' 'programmers' in 6-8 weeks.
Been to a bootcamp? Know some leetcode? Someone will hire you. And then you got Copilot advertising its services to you as a way to learn how to code. The implication of 'learn to code' being 'learn to code correctly'.
Google Translate has no similar relationship with professional translators.
You'd be surprised. What you described here for programmers is true for translators as well, and probably for many other specialities in which the ability to deliver the result is more important than any documents certifying that you've had a formal training for how to deliver those results. In case of translators — found an agency? Check. Passed an interview with a test? Check. You are good to go.
People said the same thing about UML and similar tools so I'm not holding my breath.
I can also imagine clueless bosses mandating Copilot use and that's what scares me. The real costs of most code aren't in the first writing. They're in the long-term maintenance. Copilot does not and cannot understand the whole system, or what makes for maintainability down the road. So it can't make that better, and will likely make worse. In the same way that code generation tools and code wizards made things worse.
Because management policies correlate well with engineering excellence, right?
Copilot is something different. Code is suggested automatically and, what's the most important, suggested by the authority - hey, this is GitHub, huge project, largest code repo on the planet, owned by Microsoft, one of the most successful company ever. Why should you not trust the code they are suggesting you?
And that's for starters before malicious parties start creating intentionally broken code only to hack system built with it. Greedy lawyers who will chase some innocent code snippet asking to pay for using it, etc.
I wonder if there was any sort of filter for Copilot's input — only repositories with more than a certain number of stars/forks, only repositories committed to recently etc.
That's the whole point, and the rest is moot because of it. If I chose to let Copilot write code for me, I am responsible for it's output, full stop. This is the same as if I let a more junior engineer submit code to prod, but there aren't blog posts about not letting them work or trusting them with code.
It feels like it wrote the whole line which you were going to write exactly as it should have. But that's all it does. And it seems like Copilot is the same but on much larger scale and online.
I noticed that I ended up assuming the code reviewer role when I was trying to write code. Context switching between writing and reviewing felt unnatural.
I also think I am less likely to spot a bug than I am to avoid writing it in the first place. Taking the off-by-one error in the last example. I don't think I would have made that mistake, but if copilot had presented that code block, I probably wouldn't have noticed the error either.
It's not technically possible to precisely calculate the moon's phase based on time alone. It's an optical effect that is influenced by parallax, so you have to pay attention to location as well. This is, for example, why Eid al-Adha falls on different days in different parts of the world. So the function signature itself is potentially wrong, depending on my needs. I might find that out if I had to do some Googling to finish the function, but (assuming I didn't already know) I'm not sure if that possibility would ever have occurred to me if I were using Copilot.
Copilot can spit out code that's influenced by what others have written. But can it clue you into design considerations like this? Or should we be worried that it is helping us to write code that does the wrong thing with a higher degree of confidence?
This is such curious behavior to me. Does someone really @ a corporation hundreds of times about anything? Does this have any effect? Should it?
It makes me doubt the rationality of the author’s post if ve truly did this. Although I suppose maybe their use of Twitter is just completely different from anything I understand.
Something like Copilot, but trained explicitly to analyse the code instead of writing it could be much more useful, imo. Basically a real-time code review tool. There are similar tools already, but I'm talking of something that is able to learn from the actual codebase being worked on, perhaps including the documentation, and giving on-the-go feedback.
The problem with your proposal is that it's relatively easy to do what Copilot does at the moment using AI, i.e. guess what code you are looking for and find something that does (or says it does) more or less that. However, which codebase would you use to check against if the generated code is really correct? The same codebase that produced the more-or-less-correct code in the first place?
so an AI copilot should be watching out for code I write that looks similar to code that was updated in another repo. It could even use the text from issues to synthesize a suggestion of why your code might cause problems!
Impressive, for sure. Unclear whether it's a net-positive tool, though.
Maybe “autopilot” then (as in the Tesla marketing term, not as in a real autopilot)? /s
Both lull you with a false sense of security which will suddenly and unexpectedly cause you to pay dearly for using it. Interestingly both seem to also get a vaguely similar balance of critics, apologists, and advocates on here.
Sample size of two here, but what is it with companies using %pilot to describe product features which are nothing like the actual appropriate use of the term?
> This may or may not suddenly become a huge legal liability for anyone using Copilot.
And if it doesn’t, can’t Copilot also be used to license wash for people not using it, but claiming to be?
Indeed I don't remember any of these complaints ever being made about, for example, AI-generated music or images, even though they work exactly the same and were trained on datasets of copyrighted works, both commercial and CC-licensed.
Compared to the manual sampling that DJs and other musicians do, the AI process almost certainly produces only fairly generic code snippets, since they always need a multitude of examples. Some loop-through-file python snippet is a legal risk but five seconds of a Disney song don't reach the level of creativity needed for copyright to matter? That seems strange to me...
As long as the AI just regurgitates lines from repositories like a bad undergrad cheating on his homework, CS jobs should be safe.
The fact that it has picked up the GPL might not mean that much -- it might appear in dual-licensed projects.
Github have stated that Copilot is trained on all public code on Github, regardless of license. It very trivially follows that it has been trained on a lot of code that is explicitly GPL single licensed. We don't need to do any guessing here.
They said the same thing about Chess and Go.
Plus, in theory Chess can be solved by exploring the whole solution space (which is finite, even though insanely large) and heuristics can make this practical by reasoning about which branches can probably be cut off. At that point, having more and more processing power and memory helps make the task feasible.
Not that I want to downplay these achievements, they were certainly very significant, but it's still entirely different from "solving programming" (whatever that means).
Solving Chess/Go and programming are really not much alike.
This is like: there are scholarly books that quote extensive from original philosophers -- long, third-of-a-page quotations. Still I should be able to quote something in its original language (translations may be copyrighted) copying from the derived work. Copyrighted work is not supposed to be able to poison noncopyrighted work it originates from.
I would like to see this tracked behind the scenes. At any time I should be able to get Copilot to spit out a list of suggestions I've accepted. I should be able to generate this report for the lifetime of a project.
This sounds like a great example of an interview question (where the person is asked to find and fix all of the issues in a chunk of bad code). Unintended usecase for Copilot?
function getPhase() {
var phase = Math.floor((new Date().getTime() - new Date().setHours(0,0,0,0)) / 86400000) % 28;
if (phase == 0) {
return "New Moon";
} else if (phase == 1) {
return "Waxing Crescent";
}
// etc.
It feels like small incorrect modifications to any of the code here would completely break the function.I've seen stories and articles written by GPT-3 where it will lose the plot and context on the way - in comparison Copilot doesn't suffer from this as much? How?
How is it that a glorified statistical machine is able to put blocks of code so well together?
It's easier to guess a multiple choice questions when all choices can be generated by intellisense instead of having to look at a dictionary.
And I guess it makes sense that Copilot was trained in this way, even if it wasn't just a language model. How do you even begin to separate correct from incorrect code on the entire freaking github?
But I think TFA serves best to show the worst way to use Copilot. I haven't tried it but I suspect that it would do a lot better if it was asked to generate small snippets of code, rather than entire functions. Not "returns the current phase of the moon" but "concatenates two strings". That would also make it much easier to eyball the generated code quickly without too many mistakes making it through.
Of course, you could do the same thing with a bunch of macros, rather than a large language model that probably cost a few million to train...
For me, Copilot seems like it will be useful in the exact same way that StackOverflow is useful, as a means to pointing me in the right direction for code snippets or apis or techniques that I haven't and don't really want to memorize.
For example, on my current side project I wanted to know how to create a "unique enough" UUID in pure JS.
Copilot would hopefully save me a couple of google/stackoverflow searches as I can very quickly test what it suggests.
I already rarely ever take SO answers as gospel so it's unlikely that I'd do the same with Copilot but I think it significantly increases the speed with which I achieve the same results.
Well it will not be as ubiquitous as having all the github under your fingers, but perhaps is anyway better not to blindly cite the world's source code.
Sadly you aren't. The truth is, however, that models like GPT3 and its derivatives like Codex/CoPilot are by design incapable of ever achieving this.
The only way to generate both correct and secure code is to use a combination of proper specs and theorem provers. Even then this won't help with non-functional requirements, such as performance or platform-dependent resource constraints.
Generative models will always have the potential to yield broken code that doesn't do what you want or contains security flaws even if trained on "proper" code.
If I have to audit the code that CoPilot generates and if the code is as obfuscated as the Easter example, it's probably less useful than it says on the label...
Many complaints about Copilot remind me of the old Louis CK sketch where people complain about flying: YOU'RE SOARING THROUGH THE HEAVENS IN AN ALUMINUM TUBE. YOUR ANCESTORS WOULD'VE DIED OF DYSENTERY DOING THIS TRANSCONTINENTAL JOURNEY. Let's have some context here!
Sure, it's not remotely close to perfect and it's going to take a long, long, long time for it to get there. But still, there's something about Engelbart-ian about seeing the demo when it works perfectly.
As long as it is keeping track of when people do or do not accept their suggestions, it should get better over time. But in the meantime the best bet it is to treat it like a smart autocomplete, where you still have to at least check that it got it right.
In the future maybe it will be smart enough to be treated like an intern -- trust that the code is right but still verify it yourself if the code is of any importance.
I would say that GPL3 already probably does that, which means that nobody should be using this for actual code (except if it's GPL3). But it might be helpful to be explicit about this.
https://stackoverflow.com/a/58726426
It really is hard to swallow the argument that this is not license violation on a massive scale.
You could also have some kind of AI-driven testing/verification program -- Copilot and <other program> could go back and forth multiple times until the program is deemed correct and returned to the user.
I am pretty sure the new releases will contain features like better software license handling (e.g. 3 levels for types of licenses - permissive, copy-left, hardcore copy-left), trust score for snippets, possibly some validation of the code for some languages.
Maybe because they realize the flaws are fundamentally inherent in the very core of the product. They're using a GPT-3 derivative here. DNN models are not the right tool for this job.
1. Only permissive licenses - Only include in the training set repos with permissive licenses - MIT, Apache.
2. Copy-left - Step 1 + GPL, excluding AGPL and other "hardcore copy-left licenses".
3. All - Include all code, even unlicensed and AGPL.
User can choose which version they prefer based on profile of their project and their company? Majority of github repos have LICENSE, so it doesn't seem implausible?
My guess would be that some significant portion of github code published under a permissive license is actually licensed improperly.
Working that out at scale seems intractable, but maybe the training set doesn't need to be as big as I'm assuming.
I'm all for AI coding assistance, but there's an abstraction layer in between copilot and myself (humans) that scares me a bit. At least an obvious failure prompts discussion and learning.