I’m Not Sure That (If?) GitHub Copilot Is a Problem
michaelweinberg.org
michaelweinberg.org
I'm not going to comment on whether he's right or wrong, lest I fall into the same trap I just outlined, as I'm also not a lawyer, but, unrelated to this particular case, I see this kind of thing often on the internet (as well as real life), where people who have no knowledge or experience about a particular topic go on to comment as if they were an authority on the subject. I see it on /r/legaladvice and other subreddits, and on HN here in many cases, including on certain lab leak theories. It reminds me of the Gell-Man Amnesia Effect.
FTFY
I've edited my comment above to show that he's also president of the Open Source Hardware Association so I think he knows how open source code works.
I don't care at all about his view on ethics, in this case I'd prefer a Disney lawyer.
This matches with my experience. I have never used Copilot to generate a large body of code or even anything more than a few trivially inspectable lines.
To name a very specific example, earlier this morning I was working on a VS Code extension and needed to figure out how to get it to prompt the user to select a file. I started typing out a Google search, but realized that Copilot might already know what to do. So I entered a comment like the following:
// Prompt the user to select a file using VS Code's extension API
And, like magic, Copilot generated the correct incantation, including auto filling the parameters to the API call to single select a file (vs multi select or directory select). It saved me probably four or five minutes of searching the web and figuring out the right API call and options to get the behavior I wanted.
To me, Copilot's value is as this sort of labor-saving device. It can indeed be induced to generate larger amounts of code, much of it questionable, but the tool can be a huge time-saver when used responsibly.
The advertised usecase where you describe a function with a comment then it spits out a whole function is (in my experience) pretty bullshit most of the time. It's not good enough to produce good results enough of the time, so you have to always triple check what it spat out anyway. I find it's useally easier to actually write (or start writing) the function myself.
But where copilot really shines (in my experience) is as a really advanced auto-complete. I just write my code as normal and there is a actually a pretty decent chance copilot will spit out the next few lines of what I was planning to write, allowing me to write code faster.
When you are using copilot in this way, it's near impossible for it to do copyright infringement. It's generally only outputting a few lines, and I'm only accepting them when it's reasonably close to what I was planning to write anyway.
It's really good at generating boilerplate code. Or code that interacts with API functions.
Where it really shines is converting unformatted data into code.
For example, You can copy/paste a list of words into a comment, then start writing the code to initialize an array with those words, and it will very quickly cotton on to what you want to do and complete the whole transformation for you.
Sure, the same thing could have been done with a clever macro, or some vim commands I don't know. But I just find working with copilot to be more intuitive. And it can do more advanced transformations too.
When I'm using copilot in this way, it becomes pretty clear that copilot is not simple a code search engine, or a fancy interface Infront of stack overflow. It's pretty clear that it's outputting code that couldn't be found in any database.
I get a little annoyed when everyone (including githubs own marketing) focused on the "whole function from comment" usecase. Not only does it paint it in a bad light from a copyright perspective, but (IMO) it's really not it's best or even most technically impressive mode.
which suggest, loosely, that copilot has been trained on some of your code.
But it's not like I spend my life simply writing my past code again, and it hasn't been trained on the code I'm about to write, because it doesn't exist yet.
Besides, is that not the selling point of copilot? It's been trained on everyone's code, so it should be able to predict for everyone (for some definitions of everyone)
I've never seen anyone put a license on code they post on stack overflow, so I imagine you agree to release code in your answers under a permissive license (MIT, e.g.) when signing up on stack overflow.
On github however, it's common to put a "LICENSE" file in a repository. Github even automatically shows you what license it is when its content is a known, widely used license. They do that automatically.
But somehow, with all their fancy AI magic, it's too hard for github's employees to use the same automated licensing code to decide whether or not they should share licensed code. They could also slap on a comment saying: "WARNING: this code is licensed under the <license name> license." They could even expand that, saying: "Your repository has the <license name> license, so in order to use this code you'd have to change your license."
Github could fix this issue with a hash-map and an if-statement. By not doing so they're willfully letting people commit copyright infringement, and disrespecting everyones intelligence in the process.
> Content contributed before 2011-04-08 (UTC) is distributed under the terms of CC BY-SA 2.5.
> Content contributed from 2011-04-08 up to but not including 2018-05-02 (UTC) is distributed under the terms of CC BY-SA 3.0.
> Content contributed on or after 2018-05-02 (UTC) is distributed under the terms of CC BY-SA 4.0.
Basically it says you may share the code and you may use it for any purpose, but you must give attribution and you must license your derived work under the same license.
If you don't do those things then you are basically relying on some kind of fair use or "this is too trivial to be copyrightable". But yeah, there really is a license.
The fact is that it consumes a large corpus of work and launders it in way that removes all attribution should be a big no. That goes against the entire spirit of copyright law. It must be opt-in, or make Microsoft wait 95 years to incorporate code in public repos for their for-profit enterprise.
Didn't think about the impact on sites like Stack Overflow.
Seems more like a disadvantage of Copilot because it reduces interaction on Stack Overflow.
An imperfect analogy comes to mind between textbook authors, and people who write small commercial things under someone's direction.
The wailing and outcry here is to protect skilled work from being raided (like a crucial piece of code that solves something difficult or special), or the income of people who do skilled work (authors that are committers to non-trivial libraries), not so much about low-cost javascript writers or in-house forms generation code for business processes.
I feel like there are 3 categories in that idea:
1. mapping local vars to function call (and related, how to feed the output of a function into another function from the same library)
2. coming up with a minimal runnable example of a given function
3. semantic search for which function to use for a task
docs shipped with a library often fail at this, but there are things languages can add to get better at all 3. In particular, syntax for automatically consuming local vars in function calls, and in the package manager, extracting function comments and offering global indexes.
As for "actual lawyer": The linked-in profile of the author includes no work for any law firm or on relevant case law; it has some public-interest and open-source advocacy and his own "firm" consulting with startups, and some legal work for a startup as GC. It's good that he's been interested in this domain, but his work is not obviously validated. It may be wise for him not to attempt a legal assessment, or for readers not to grant him too much deference.
With no disrepect intended to the author, I would encourage readers to make their own assessments before accepting Copilot as benign.
I appreciate satvikpendem's caution about non-lawyers, but both discussion and the law are probably best served not by deferring entirely to lawyers. The law should make sense, and non-lawyers should be able to use it to reason about the legality of practices, at least to validate legal opinions.
Much of what matters in law is the process and posture, especially establishing who has the burden of action and proof. With Github copilot, Microsoft is daring (forcing?) people to invest a lot of money adjudicating novel issues for uncertain damages to recover on intangible benefits like community. If Microsoft/Github want to establish favorable case law, it's best to find weak opponents with few incentives. Even if there were a large company with code copied from Github, they'd have to be willing to anger Microsoft (who could likely fork their business on a whim).
The article setting the expectation that there's no problem or harm only helps Microsoft. In most cases, the executive director of any nonprofit, even when tied to the university, is largely responsible for fund-raising, and generally pursues the goals of the funders, or telegraphs opinions likely to raise funds. Better sources would be tenured professors with publication histories on point, or practitioners who have adjudicated some of the case law. Corporate lawyers working in the domain have the expertise, but likely not the incentives.
Fair use, per <https://fairuse.stanford.edu/overview/fair-use/what-is-fair-...>: "limited and “transformative” purpose, such as to comment upon, criticize, or parody a copyrighted work. [...] Most fair use analysis falls into two categories: (1) commentary and criticism, or (2) parody"
Copilot is decidedly not commentary, criticism, or parody. It's extracting code in whole and remixing, just as Google does in extracting text to its search listings. The difference is that there's a benefit to the listed site leading it to permit this use. Copilot offers zero benefit, and may indeed reduce interest in the source code.
Also remember fair use is an exception to copyright, a defense. The burden is on the infringer to establish the exception in case law. To date, no one has.