So everything generated also GPLv2?
So everything generated also GPLv2?
How would we classify that legally when it comes to training and generating code? Would you argue the machine is just picking up best practices and patterns, or would you say it has gained specifically-licensed or proprietary knowledge?
More generally, keep in mind that the legal world, despite an apparent focus on definition is very bad at dealing with novelty, and most of it will end up justifying a posteriori existing practices.
Even a search engine is not merely a "compilation of facts". A trained model is the result of analysis and reasoning, albeit automated.
I believe you're referring to Clean Room Design[1].
That said, copyright is typically generally assigned to the closest human to the activation process (it's unlikely that Github is going to try to claim the copyright to code generated by Copilot over the human/company pair-programming with it), but since copyleft in general is a pretty domain-specific to software, afaik the way that courts interpret the legality of using code licensed under those terms in training data for a non-copyleft-producing model is still up in the air.
Obligatory IANAL, and also happy to adjust this info if someone has sources demonstrating updates on the current state.
[1] https://techcrunch.com/2021/01/12/ftc-settlement-with-ever-o...
imagine if the output was ruled as being GPLv2, then having to go through a proprietary codebase trying to rip out these bits of code
it would be basically impossible
That's not even remotely the same thing.
Take a look at https://karpathy.github.io/2015/05/21/rnn-effectiveness/
At the same time, I would not be surprised if there are outputs that do correspond to the source training data.
[1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,....
edit: from https://copilot.github.com/
> What is my responsibility when I accept GitHub Copilot suggestions?
> You are responsible for the content you create with the assistance of GitHub Copilot. We recommend that you carefully test, review, and vet the code, as you would with any code you write yourself.
Well, that solves that question.
And yes, of course github is not going to take responsibility for things you do with their tools.
It's only infringement if the portion copied is significant either absolutely or relatively. A line here or there of the millions in the Linux kernel is okay. A couple of lines of a haiku is not. Copyright is not leprosy.
Wouldn't that imply that a person who learned to code on GPLv2 sources wrote writes some more code in that style (including "long strings of code", some of which are clearly not unique to GPL) is writing code that is "born GPLv2"?
I don't think it currently works that way.
"[...] the vast majority of the code that it suggests is uniquely generated and has never been seen before. We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set"
Anyways, GitHub is Microsoft, and Microsoft has really good lawyers so I guess they did everything necessary to make sur that you can use it the way they tell you so. The most obvious solution would be to filter by LICENSE.txt and only train the model with code under permissive licenses.
This line of thinking applies to the code generated by the model, but not necessarily to the model itself, or the training of it.
I was wondering the same thing, especially with MS being behind both.
edited: or this? https://docs.microsoft.com/en-us/visualstudio/intellicode/cu...
I suspect there could be issues on the training side, using copyrighted data for training without any form of licensing. Typically ML researchers have a pretty free-for-all attitude towards 'if I can find data, I can train models on it.'
Note also that it's not just a concern for copyright, but also privacy. If the training data is private, but the model can "recite" (reproduce) some of the input given an appropriate query, then it's a matter of finding the right adversarial inputs to reconstruct some training data. There are many papers on this topic.
> In 1994, the U.S. Supreme Court reviewed a case involving a rap group, 2 Live Crew, in the case Campbell v. Acuff-Rose Music, 510 U.S. 569 (1994)... It focused on one of the four fair use factors, the purpose and character of the use, and emphasized that the most important aspect of the fair use analysis was whether the purpose and character of the use was "transformative."
It has some neat examples and explanation.
[0] https://www.nolo.com/legal-encyclopedia/fair-use-what-transf...
Almost certainly not everything.
But possibly things that were spit out verbatim from the training set, which the FAQ mentions does happen about .1% of the time [1]. Another comment in this thread indicated that the model outputs something that's verbatim usable about 10% of the time. So, taking those two numbers together, if you're using a whole generated function verbatim, a bit of caveat emptor re: licensing might not be the worst idea. At least until the origin tracker mentioned in the FAQ becomes available.
[1] https://docs.github.com/en/early-access/github/copilot/resea...
[2] "GitHub Copilot is a code synthesizer, not a search engine: the vast majority of the code that it suggests is uniquely generated and has never been seen before. We found that about 0.1% of the time, the suggestion may contain some snippets that are verbatim from the training set. Here is an in-depth study on the model’s behavior. Many of these cases happen when you don’t provide sufficient context (in particular, when editing an empty file), or when there is a common, perhaps even universal, solution to the problem. We are building an origin tracker to help detect the rare instances of code that is repeated from the training set, to help you make good real-time decisions about GitHub Copilot’s suggestions."
Not a critique on your point, which a was just about yo bring up myself.
2) if I write a program that copies parts of other GPL licensed SW into my proprietary code, does that absolve me of GPL if the copying algorithm is complicated enough?
https://docs.github.com/en/github/site-policy/github-terms-o...
This is the same reason it doesn’t matter if you put up a license that forbids GitHub from including you in backups or the search index.
I think this is definitely a gray area and in some way iParadigms winning (compared to all the cases decided in favour of e.g. the music industry), shows the different yardsticks being used for individuals and companies.
I'm sure we will see more cases about this.
[1] https://www.plagiarismtoday.com/2008/03/25/iparadigms-wins-t...
There is no such thing as "fair use" as we have in copyright law?
There're definitely cases when devs avoid even looking at implementation before creating their own