Is GitHub a derivative work of GPL'd software?
drewdevault.com
drewdevault.com
To be sure, the infectious nature of the GPL is likely to be a bigger deal, but permissive licenses will be violated by missing acknowledgement just as much as the GPL will be violated by not licensing your code GPL.
If this license-washing doesn’t work, then you can’t just say “we’ll exclude GPL sources”, because you’ll still be breaking the rules on any copyrighted work that doesn’t have very specific public-domain-like do-whatever-you-like-and-don’t-acknowledge-me licensing.
The whole thing truly does depend on the fair-use classification of ML models. If that falls apart, the entire thing’s done (and a lot of other ML stuff will be, too). And that’s where regurgitation becomes so problematic, because it directly undermines that fair-use exemption.
More importantly, just because a chunk of code is published under CC0 on github that does not give you any guarantee that such code is not already breaching an existing copyright.
It happened in various occasions that closed source code was leaked into FLOSS projects and if you reuse it you are in breach of the license as well.
For example MIT, one of the most permissive licenses there is, still requires you to include the MIT license and copyright notice "in all copies or substantial portions of the Software".
It sounds like you meant something more like a public domain dedication like CC0, which is about the only way you can go more permissive past the likes of MIT/Apache2/CC-BY.
Creative commons specifically calls CC0 an "alternative to our licenses" [1]
https://creativecommons.org/share-your-work/public-domain/cc...
- produces an output
- that output is the exact same thing - text, code, image, et c - that went in
- that thing is licensed
- fails to fulfill the requirements of the license
than it's breaking the license. I don't see how it can be wiggled around.
The black boxes that produce entirely different outputs, or the ones that slurped billions of photos without even blinking at license but never vomits actual images are, at least, not reproducing the source material.
1. A hypothetical employer asks an employee to find a really fast algorithm to find an inverse square root. The employee has a perfect memory and can regurgitate fast inverse square root from Quake 2.
2. A hypothetical employer asks an employee to find a really fast algorithm to find an inverse square root. The employee uses this ML algorithm to generate the code which essentially regurgitates the only previously seen inverse square root function it was trained on.
At some point in time in the future id Software sees this is happening at our hypothetical company and sues. In neither situation is the "actor" held accountable since the "actor" was employed by the company which allowed for practices that enabled this infringement to take place.
I think an important legal question is: Do we consider a ML algorithm being trained on data the same as a human reading prior art for an innovation?
It seems pretty obvious that it is the same thing to me. Instead of the human reviewing the prior art themselves the human has written a program to do so, but this can't change the acceptable outputs produced by the human's creative process (that includes the ML in this case).
There's a line somewhere, but it's somewhere, not at either edge, and it's a complex, odd-shaped line we haven't yet defined. I don't know where it is. However, it's not where you're putting it, and there isn't a slippery slope to "IP is over." If you use five lines of my code in a different system, IP hasn't ended.
The amount of copying does not matter for a copyright claim. If you copy a single character from a codebase that could could get nowhere but there, and lawyers could prove it, that could go to court. This is a hypothetical but this is entirely possible. There have been copyright cases fought over single sentences especially in the music industry.
> Are you trying to launder, or is it unintentional?
I'm not a lawyer but in the US, Copyright is a strict liability statute which means intent does not matter.
Clearly that's not how copyright works.
> that could could[sic] get nowhere but there, and lawyers could prove it
Also, the letters in this case are not representative of the creative work itself. If you used the title font to write your post (hypothetically, I know HN does not allow this), you'd be in bad water already.
>....you'd be in bad water already.
Not true either, writing my message in harry potter font would do absolutely nothing. the internet is full of images which use these or similar looking fonts. Its called fan art or whatever, copyright laws are not applicable to this.
Please explain how it "is entirely possible".
Bear in mind that there must be creativity in that choice of character to qualify for copyright protection. If only one character is correct, there's no creativity.
A sentence in a song has far more discretion for creativity than a single character in a software program.
As best as I can tell, you would need to construct some sort of super-APL where "a single character" had significantly more information content than any Unicode glyph, to exceed any de minimis standard applied to a programming language.
(Think Prince's "Love Symbol".)
> that could go to court
I think you mean "that could be found infringing." Going to court is trivial, even if there is no infringement.
This used to be something cartographers did: https://en.wikipedia.org/wiki/Trap_street
Even if you didn't have the street's name, if you copied the road on the map (no words, just the shape) they would know you stole their IP. I do not think there is literally a character that would satisfy this condition. The point of what I was saying was the "size" of the copying does not matter at all.
The Substantial Similarity is explicitly about this phenomenon. For something to be substantially similar, from a software perspective, you could hypothetically see the modules and data structures contain roughly the same data types and that the flow of logic is the same. You could also point to a function and say "these are the same variable names". Substantial similarity is not the only component of a legal analysis of a copyright claim. Another important component of copyright is if someone had access to the material. If I had code on my laptop that I showed to no one, uploaded nowhere, and at some point I find another developer who is doing some code that is character for character identical to my work, I likely cannot do anything.
The law is complex here, I am not a lawyer, but I think the important component here: The length of copying is not a factor at all. If for some reason your algorithm requires `magic_code.seed(20154)` and your code does this by doing `magic_code.seed(ord('人'))` and you find someone else doing the same thing you'd definitely want to investigate what's going on there.
The entire copyright doctrine is created to help protect the creativity of an author. If there is something creative in your code that is copied, you likely have a copyright claim you could argue.
Not "roughly the same data types", "same variable names", "magic_code.seed(20154)", etc. I have no question with those. Those are enough characters that they may contain creative - and thus copyrightable - content.
That is, I question your assertion that 'the "size" of the copying does not matter at all' by asking you to come up with an entirely possible scenario for why there is no de minimis case in software, even down to a single character, when there is in every other area of copyright.
BTW, the WP page you pointed to notes that "Trap streets are not copyrightable under the federal law of the United States." Thus, "stole their IP" has no meaning - they have no copyright, trademark, patent, etc.
Where in copyright law for any area is there an established minimum "length"? I've never heard this nor heard lawyers claim of such a thing.
OTOH, your original context for an "entirely possible" scenario at https://news.ycombinator.com/item?id=27730596 was "The amount of copying does not matter for a copyright claim", which implied your scenario was not that hypothetical example I posited.
(A 1E-20 probability * maximum expected lawsuit payment with successful lawsuit = don't worry about that scenario.)
We do know that the courts have decided many cases are de minimis non-infringing use of materials otherwise under copyright.
Do you think there is no acceptable de minimis argument in software? If not, why is software somehow special compared to other areas of copyright?
A "de minimis analysis ... usually focuses on the amount of the copyrighted material that is copied." - quoting http://patentarcade.com/2020/04/nba-2k-avoids-tattoo-copyrig...
If 'de minimis' use exists in software, what realistic scenario are you thinking of where a single character is enough?
To be clear, I think zero shared characters can still show copyright infringement, for reasons you mentioned about abstraction-filtration-comparison.
But your "entirely possible" example had no other shared similarities beyond a single character, and I can't see how any court wouldn't think that was a de minimis use of that character - assuming it had copyright protections in the first place!
>The amount of copying does not matter for a copyright claim.
This is incorrect. The amount of material used is the third factor in a fair use test. There are other ways it comes up as well (e.g. damage calculations).
>> Are you trying to launder, or is it unintentional?
> I'm not a lawyer but in the US, Copyright is a strict liability statute which means intent does not matter.
In this context, this is incorrect as well. People seem to be misinterpreting laws pretty badly, so instead of explaining how to apply here, I'll give a simpler, analogous context:
https://en.wikipedia.org/wiki/Online_Copyright_Infringement_...
In this context, it depends on a lot of things, such as direct versus contributory infringement. This would likely be a contributory infringement case, where intent is almost a requirement. https://www.legalmatch.com/law-library/article/what-is-contr...
Keep in mind these are all factors. They matter, so they'll help swing a case one way or the other, but they're not something you can bank on in isolation. That's what will make for interesting case law.
The 3rd factor of the fair use test deals with how substantial the copied material is in relation to the entire work. You can make a copy of 100% of the original work and have it still be fair use while in another situation copying 0.002% of an original work would not be fair use. The measurable amount (bytes, seconds, square inches) you are copying does not matter here. Instead the "importance" of what you are copying to your criticism is being defined in this prong.
For instance:
I am a movie buff and:
1. Make a copy of the entire movie of Citizen Kane 2. Remove the audio and replace it with a commentary track going over every framing/camera trick through the movie 3. Upload this to an educational youtube channel that teaches viewers how to shoot a movie
I would likely be allowed to use fair use as a defense to a copyright claim. Here, if an expert could say all 100% of Citizen Kane's film, contained relevant film techniques that I was actively commenting on, I would have copied the correct amount for my usage.
If I instead owned a movie review channel and I said "Citizen Kane is my favorite movie" and I displayed the entire movie after that I would likely not be using an appropriate amount of the film.
It's also important to note that fair use is a very narrow defense that covers a very limited subset of uses. It is not a generally applicable, or even a guideline, for how to skirt around copyrights.
If I am doing code autocomplete, and copying three lines out of your program, those three lines are exceptionally unlikely to be "substantial ... in relation to the entire work."
Where you're confused is lots of places, but the biggest one is you're mixing up the prongs. Most of the reason your Citizen Kane example might work (it probably wouldn't) is Factor #1, character of use. Commentary and education are favorable.
I can come up with examples where 0.002% of an original work would not be okay, but they're pretty contrived.
For your benefit, a random link: https://guides.lib.utexas.edu/fairuse/fourfactor
Size definitely matters. It's not a hard-and-fast rule, but a factor.
It absolutely does for fair use analysis, see 17 USC § 107: “[…] In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include—[…]the amount and substantiality of the portion used in relation to the copyrighted work as a whole”
How are these statements not contradictory? Do you know where the line is, or not?
The trouble is verbatim copy of original copyrightable code, with an attached license that doesn't allow it. Github is most likely in violation because it's doing just that.
However that specific square root square example may or may not be a violation. There's a difficult question of what constitutes original copyrightable work, a trivial well defined algorithm isn't, a larger code with comments could be.
This is very myopic because there's a big difference between "memorizing an algorithm" and "memorizing an implementation of an algorithm".
> There's a difficult question of what constitutes original copyrightable work, a trivial well defined algorithm isn't
I don't think copyright law agrees with this sentiment in the US. Even the spacing, variable names, and comments are non-functional attributes of this code. And, this also goes further, because the machine code it produces is also data that has copy rights. There's layers and layers of complexity here that is difficult to enumerate with forum posts.
No there's isn't. An implementation of an algorithm is not copyrightable. It's specifically cited as an exception in European directives.
I should say I speak more from a EU perspective. US may vary slightly. Both vary by country/state as well.
There is a challenge with determining the threshold for copyrightable original work and it's not the same in every jurisdiction. That we can agree on. (I personally think comments are a big problem because they're always freeform text). In my opinion the tool should consider doing an autoformatting and removing comments, to avoid these issues.
The gist:
>> "There is an observable trend in US law, based on fair use and older notions in US copyright law of the need for creativity, that judges give a looooot of leeway to “machines that read”. Copilot fits pretty squarely in that tradition."
If that's something I'd planned on and intended, and you own the copyright to the painting, you have all sorts of legal tools to go after me -- contributory infringement, collusion, and so on. Still, contributing to a legal violation is not the same as engaging in one.
I suspect an argument can be made that github now incorporates AGPL code generated by co-pilot, and so is AGPL. A fair use argument might be made as well; we've all copied one-liners from blog posts and tutorials, and that's okay. I'm having a harder time seeing a reasonable argument that an ML model is directly infringing just because it can produce a copyrighted work, though.
So I don't think it's a cut-and-dry legal question.
I will give a caveat: Coders tend to read laws much too literally, like computer code. When I was was an obnoxious teenager, I thought I'd found all sorts of contract / license / etc. loopholes in all sorts of legal documents, and I considered myself profoundly clever, thinking I'd outsmarted the lawyers.
Nope.
In college I took law classes, got into the real world, did a few startups, and saw a few legal cases. Court systems have technical rules you need to be aware of, of course (e.g. if you miss a deadline....), but interpreting the law is really grounded in common sense. Common sense is culturally situated, and ours is situated in hundreds of years of case law.
Thats very interesting. Because AFAIK Github is not "shipping" copilot to anyone. They run it and make it available only to those that use their online editor.
If that's the case, then there will not be any users of Copilot except those that unknowingly open themselves to liability. I suspect there will be few.
I would be okay with that situation. But I'm not okay with the situation as it currently stands.
Where it stands: We have a new technology. We don't really understand it. There aren't well-defined ethical or legal bounds. We're figuring them out, blumbering along as we go. People are trying things, and are making all sorts of arguments, some of which will stand the test of time, and some of which won't.
I'm more concerned when we do this with weapons, medicine, or social science, but an autocomplete tool seems like a relatively benign place to start figuring this out.
But let's fix this before it goes further, especially since mindlessly copying output from a machine will probably result in more bugs.
Not all networks are susceptible to this, but if you can get a 1:1 copy, that's no different from bittorrent, limewire, or KaZaA. That doesn't seem implausible for text.
Napster was an easy litigation mostly because of intent. There was a document trail a mile long that the primary goal of Napster was to facilitate copyright infringement, although on paper, it could be used to share anything.
Cameras, xerox machines, and audio recorders can all be used for infringement, but aren't illegal since that's not their primary purpose. For even finer lines, see the Betamax case. You can also look at cases about different sorts of transformations and transmissions:
- I buy a DVD. It is transferred via USB cable from a DVD player connected to my computer, temporarily copied to computer memory, and then shown on my monitor.
- I buy a DVD. It is transferred via the internet from a DVD player in my data center, via the Internet, temporarily copied to your computer memory, and then shown on your monitor.
Can you see how these have nearly identical technology, yet different legally? If you ever take a law class, there are a lot of discussions about how to draw such lines (disclaimer: IANAL, but I've sat in on law classes before).
To be clear, I'm not taking a stance about where this ought to land legally. I don't know. But the analogy to bittorrent, limewire, or KaZaA is a faulty one.
Plot twist: Depending on the jurisdiction, you are responsible for distributing means specifically designed to circumvent copyright, which makes you liable for a much greater punishment, and liable for every reproduction done by your users.
There's a reason that painting robots are sold as general painting robots for various industrial usage, they're not advertised to and shipped with templates to reproduce copyrighted paintings. ;)
How is that different than sending someone a jpeg of a copyrighted painting? Or the jpeg and the executable that can render the jpeg. Or the jpeg and the executable and the computer it runs on. Each of those gets closer to what you said.
What case did this proposed legal test come from?
Github contains plenty of software uploaded without any license, which, by default, means closed source. Similarly, if use significant portions of those you violate copyright law.
There is no concept of "infection" and "virality" in copyright law.
But I think you know what I’m referring to by “the infectious nature of the GPL”: that such licenses as GPL and the CC-SA family require that derived works be under a similar license. And the fact of the matter is that people are likely to care about that a bit more than missed attribution, whether they should or not. But, as I say, it’s still violation.
With other licenses if you distribute derived work, the distribution is illegal.
With GPL in the mix the derived work becomes GPL as well.
For something like MIT/Apache2, the burden is that your derived work must acknowledge the copyright and license of the original work.
For something like GPL, the burden is that your derived work must acknowledge the copyright and license of the original work, and must be under a similar license to the original work.
So yeah, it requires more of you, but it’s fundamentally the same proposition of “obey the terms or it’s illegal”.
Correct. On top of that, GPL is very lenient and provides multiple escape mechanisms. Also it does not require any damage compensation.
As with any other software you must either stop distribution or come to an agreement with the party owns the copyright to the code. The GPL merely provides you a get out of infringement free card whereby you may release all of the code under the GPL. This is different only insofar as it is more permissive not less compared to proprietary software.
The idea that the GPL infects your code is a fantasy.
Meanwhile companies will more or less blanket-ban GPL and AGPL software, even if the terms would not impact them at all.
I’m conflicted on this, on one hand I think FFmpeg is a brilliant piece of engineering and I want the creators to receive credit, but on the other I think Audacity and its built in components as listed in its “third party licenses” document are brilliant pieces of software that do receive credit. Further, I believe the one of the primary purposes of technology should be to enable the masses to create, and this licensing places an arbitrary yet powerful restriction on that for folks who aren’t already computer-proficient.
Audacity and FFmpeg are both under the L/GPL, compatible versions even. The reason Audacity doesn't include FFmpeg is because they are afraid of software patents. GPL has nothing to do with it. At all.
> Because of software patents, Audacity cannot include the FFmpeg software or distribute it from its own websites. Instead, following the links below to instructions to download and install the free and recommended FFmpeg third-party library.
If you incorporate closed source software into your work you may stop distributing your software or rip it out possibly while paying a massive fine.
If you incorporate GPL licensed software into your work you may stop distributing your software, rip it out, OR release your software as GPL. Furthermore while being discovered using someone else's proprietary software in yours will almost certainly result in at best a threat of near instant lawyer assisted corporate suicide using GPL software the same way has historically led to a strongly worded letter and months to years to fix your problem.
This is to say that nearly all copyright is inherently viral and GPL is slightly less so than many because you have additional options. Hearing people, many of whom metaphorical develop software Ebola for a living describe the GPL as viral is in a word, amusing.
Also aficionados of permissive licensing seem to be frequently not only OK with proprietary licensing but approve of the fact that code can be used in such a fashion. It's almost as if the attitude is everyone's else's work MUST be free for ME to use in MY proprietary code else they are somehow taking away MY freedom!
And it's still an incorrect and arbitrary use of term.
If you release software that is breaching GPL due to a dependency you have many options to rectify the issue.
Under no conditions your software "catches" a GPL "infection" and magically changes its license while you are not looking. There is such concept in copyright law.
EDIT: you can do better than silent downvotes.
Maybe it’s easier to see an issue when the GPL specifically lays out the creator’s intent, compared to the nebulous intent and law that underlies unlicensed content. GPL focuses the discussion on a concrete license rather than the entire body of IP law around fair use.
Or maybe it’s just easier to see an issue when it’s “our” community’s content being used.
Perhaps somebody with an AGPL'd project on GitHub and deep enough pockets could force them to open source Copilot...
(No, I don't think that's actually plausible, but the scenarios are fun to think about)
Big tech uses other public or "free" CC, etc knowledge bases to train ML models. Like Wikipedia as a training set for a GPT or word2vec model. Or using WikiData to seed a knowledge graph. Or training based on a common crawl dataset. All of these resources are used to make the assets that major tech companies use to make their products a lot smarter. They're built on the free labor of the rest of society.
Big tech reaps tremendous, society altering benefits from the free labor of others. And "training a model" seems to whitewash away the licensing concerns.
Shouldn't that be disconcerting whether we are developers or wikipedia editors?
Instead, we just see everyone using the “all rights reserved” license, and the commercializer saying who cares it’s fair use.
I think we have overwhelming data at this point. What you rather have: Linux or Windows? The WWW or American Online? Wikipedia or Encyclopedia Britannica? Sqlite or Oracle?
Copyright law hinders the creation of ideas, not helps. It's time for Intellectual Freedom.
Someone should make a movie about a world without copyright.
We don’t need copyleft in a world without copyright.
I don't think that that's a realistic fear, though; this is basically already how it works with FOSS code being used in remotely-hosted SaaS systems and there being no effective legal recourse for FOSS authors.
It just doesn't officially say it yet in a law.
I mean, I disagree with your conclusion, I think "aboloishing" copyright is a bad idea.
But I even more strongly disagree with how you phrase this. "Overwhelming" data? What are you talking about? There's a valid debate to be had, but if you're starting form the assumption that it's completely clear that your way is right, I think it's going to be pretty hard to convince anyone else of this since almost everyone disagrees with this idea.
Well first I am a big believer in always building more datasets, and I have been involved with some efforts to built datasets on the copyright issue, and expect there to be many more. But this is a "pebbles" and "planets" type of situation, where the value of the public domain creations dwarf those of the copyrighted ones (American Online vs The World Wide Web; Microsoft Windows vs Linux; Patented medicines vs public domain ones). It's not that public domain innovations are a little bit better, they are OOM better. In the long run it's no contest.
So I think building bigger and better datasets across all domains is important, but this is not going to be a close call. I thought I might have a blind spot in medicine, but then spent a few years in biomedical research and realized that nope, the copyrighted and patented stuff in there is mostly junk compared to the public domain innovation.
Edit: Thanks to those who have pointed out that Nat Friedman has publicly commented on this (on HN!). My mistake.
>In terms of the permissibility of training on public code, the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use.
As long as enforcement and punishment of corporate lawbreakers doesn't hurt them at all, they're gonna continue to find new ways to abuse the system and get away with it. Only way to change it is to make their crimes really hurt them when they're called to task over it.
And yes, I do agree that some tech companies push the boundary of information and commercial law and wait to be tested by someone with deep pockets.
I'm not sure whether or not I agree that is happening here, but I've spent enough time in big companies to develop a healthy skepticism of their decision-making processes. So much groupthink.
I still don't understand how there's so much armchair quarterbacking around legality of Copilot, from people with zero legal skills and just anti-big-tech hot-takes.
Because Microsoft's legal department (and the company itself) has a spotless reputation for ethical and legal compliance.
Regardless, if you’d rather discuss morals - I don’t believe what they’ve done is immoral or wrong either. I use Linux, and write FOSS code and contribute to the FOSS community quite regularly.
Devs can ignore legal for a very short time, but the moment you have a team working on it, you need legal approval to continue. And if there's going to be a public launch, you have an entire legal, data, and privacy team to vet it first.
Put simply, there is no universe wherein MS's legal team hasn't cleared Copilot before launch, and I'm disappointed HN groupthink believes otherwise.
I’m saying FAANG+MS definitely have this process, you can ask any employee how it works there.
There is 0% chance that this wasn't run past their legal team.
>In terms of the permissibility of training on public code, the jurisprudence here – broadly relied upon by the machine learning community – is that training ML models is fair use.
Anything that is licensed has some level of exclusivity on it, and therefore it is not Public Domain.
What are you basing that? Especially when this is what is on the Copilot FAQ:
>GitHub Copilot is powered by OpenAI Codex, a new AI system created by OpenAI. It has been trained on a selection of English language and source code from publicly available sources, including code in public repositories on GitHub.
Absolutely zero mentions of public domain.
2) If an author did grant license to his code, then somebody, but not anybody, can use the code according to license.
3) If a somebody refuses to comply with license, then he permanently loses the right to use the code under such license, so then see (1).
>In its most general sense, a fair use is any copying of copyrighted material done for a limited and “transformative” purpose, such as to comment upon, criticize, or parody a copyrighted work. Such uses can be done without permission from the copyright owner.
https://fairuse.stanford.edu/overview/fair-use/what-is-fair-...
this is false. Fair Use is not about the license, it is a set of exemptions from copyright. for example I do not need a license to reproduce and publish a snippet of book for the proposes of review / criticism. The author may not like what I have to say, and did not grant me a license but I certainly can use the snippet under copy right law.
The Legal Experts of ML are relaying I believe on the "education" exemption in Fair use for their "training of ML models is fair use". That will be an interesting case for sure
GitHub's CEO posted this:
> In general: (1) training ML systems on public data is fair use (2) the output belongs to the operator, just like with a compiler.
Not sure I understand what fair use means in this context though and I can't believe that it's an argument that works for every conceivable model.
My clipboard is an ML model that trains on exactly one sample at a time. Ctrl + C is a hotkey for doing a training run on currently highlighted data, and Ctrl + V is a hotkey for generating output from the previous run.
What do you think the 'P' in GPL stands for?
Regardless:
1) You don't need a license to look at things. Copyrights reserve very specific rights: Reproduction, creating derivative works, distribution (limited to first sale), performance, display, and in some cases, transmission
2) And data can't be copyrighted.
US IP law is based on a clause in the US Constitution, and is fundamentally bounded in reach.
What do you think the 'L' in GPL stands for?
GPL code ≠ public data
I'm not going to engage in a flame war here. I clearly laid out references to relevant parts of the law: (1) What a license is and is not needed for under US copyright law (2) purpose of the GPL, and (3) distinction between copyright and data rights.
If you have any sort of citation for this nonsense, please post it. Otherwise, I'm checking out here.
When you say that a computer program "looks" at my data, because my code is for public, it means that the owner of the computer program use my code to make derivative work without obeying of my license, which is explicitly forbidden by copyright law.
For example, you cannot pretend that your smartphone is just looking at a copyrighted film in a cinema. Nobody will ever bother to ask your smartphone, what it's doing. Stop pretending that your smartphone can "look at data".
While correct, using protected works in an AI model requires some form of reproduction. In Canada, even the loading of information into RAM is considered a reproduction (although such reproductions may be excluded from infringement under s. 30.71 of the Copyright Act, RSC 1985, c C-42).
> 2) And data can't be copyrighted.
True, but data is often an abstraction of other expressions that are protected by copyright. Additionally, compilations of data and database structures attract different protections in various jurisdictions. E.g. sui generis database rights in the EU.
Based on his assertion, can I take a copy of windows (or github for that matter) and train an ML on it to replicate its behavior (assuming I have such a powerful AI)? It seems to me that is essentially what follows. In fact I'd argue that this is more akin to clean-room reverse engineering.
It also seems to fly in the face of case law for using soundsnippets in songs. Isn't what the model is doing really just a sophisticated algorithmic rearrangement of samples?
Second, code is not data, it is a copyright-protected creative work.
It’s really hard not to see this as him erasing the rights of individual developers for his own benefit.
IANAL. I don't know if that distinction exists in law. But it doesn't make sense to me. How do you distinguish code from data? Isn't the whole point of code that it is treated (by the machine) as both code and data?
Suppose I make a "copyright-protected creative work", for example a melody, and then encode that as a database table?
Suppose I have a list of population statistics? Suppose those population statistics have been ingeniously contrived (e.g. by encoding, ordering or whatever), so that given the right interpreter or compiler, the logic of a copyright software work can be replicated exactly, using the list as code?
It's hard to imagine a list of statistics that can be read by a human as easily as he can read code; but that doesn't matter. It's hard to read machine-code too; but copying machine-code without permission is infringing.
It doesn't help that copyright law has drifted over the years, and that despite international conventions, the law varies from place to place. USA and EU see copyright differently. Commenters here are not declaring whose copyright law they are referring to.
It does. Data is not copyright-able but source code is. There is so much legal precedence around this it’s nearly impossible to contest.
Calling code data is a way to minimize its status as a copyright protected creative work, so as to defend ignoring the protections it has under the law.
But the moment I redistribute (1) the model and that (2) model reproduces the original copyrighted training data that is an obvious violation.
The reality is there is zero case law that fits training of ML models enough for any legal expert to claim jurisprudence makes it legal. ML is too new of a technology for that claim.
If ML model is reproducing verbatim entire sections of code, I fail to see how that is "fair use" and not reproduction. I also fail to see how any legal expert would make the claim that it is.
The response of "oops it should not have done that" would also not hold up well in a licensing / copyright dispute
He's correct that AI can train, but it's not relevant.
(1) The issue is with CoPilot outputting copyrighted code AND attaching a separate license to it (looks like it's hiding the initial license and can generate another license automatically). This is reproducing copyrighted code and misrepresenting the license, neither is allowed.
(2) He posted that the output belong to the operator. It's factually wrong. Original copyrighted code belongs to their original writer, not the operator. In effect Github Copilot cannot "launder" license and it is very wrong for their CEO to claim otherwise.
(3) Looks like he's trying to waive responsibility of Github by stating that the developer receiving the code is responsible? Wrong, GitHub is responsible for their actions (vary with the jurisdiction). There's a complex matter of who's responsible when some proprietary code will end up in a company product and they get sued. The company is in violation, they can turn against GitHub for providing copyright code, GitHub was responsible for it. There's a complex chain of responsibility, any lawyer worth their salt would cringe at the claims and responsibility that GitHub is exposing itself to.
This whole thing could be done above board, and it's pure sloppiness on the part of github that they've chosen not to.
That’s an understatement. There is a deep lack of understanding of IP issues as demonstrated by their public representatives. Surprising for a company like GitHub where software IP should be a part of their core competency. They have made defenses like “it’s public data” which means nearly nothing in regard to copyright enforcement. The Avengers is public “data” (it’s not data in terms of the law), that doesn’t mean you can use the content for whatever you want. It shows a total lack of respect for the work that the FOSS community does and their rights. This has irreparably changed my perception of GitHub forever in a strongly negative way.
> It shouldn't do that, and we are taking steps to avoid reciting training data in the output.
Seems like he agrees that __reproducing__ copyrighted code is not permissible.
Also seems likely to me that Nat etc were probably misled / incurious (no excuse) about precisely what CoPilot would do in practice.
Actually, they should not be able to, as the whole point is that CoPilot does not just paste code, but instead understands it and applies knowledge, not finished snippets.
The fact that it does paste snippets is the problem, not the authorship of the snippets, as these simply should not exist.
On the other hand, a wholly independent system to search the training set based on AST matching and text similarity could solve this problem in the majority of cases.
Considering that their entire business model from the start has been "capturing the open-source ecosystem onto a closed and proprietary platform that doesn't interoperate", I'm more inclined to assume malice here than 'sloppiness'.
Edit: By which I mean that their approach most likely just couldn't work well enough for end users, in terms of hassle, if they actually did it above board and included licensing information. So not caring about that was the profitable option.
Seems like he agrees with you that __reproducing__ copyrighted code is not permissible. Doesn't answer the broader point about training with that code.
If it can then there's no problem unless CoPilot actually did produce code that was a close copy of GPL'd code and GitHub used it. Their own analysis says that that is really unlikely and easy to prevent anyway.
Almost certainly. But what people tend to forget is that "working out the correct legal interpretation" isn't actually the job of a lawyer; it's the job of a judge. The job of a lawyer is to devise a legal strategy that works for the client.
In practice that will quite often mean "yes, this is illegal, but we will publicly claim that it's legal and ride the wave, and the consequences will cost us less than we can profit off it".
Even if it's not explicitly stated like that and couched in vague "legal risk" handwaving, the lawyer ultimately works for their client and that's who their allegiance is to, not to the law.
That's the way it works in practice; in theory, lawyers are "officers of the court", and are supposed to serve the law.
The pessimist in me believes that this fundamental conflict will be used by big established players to get rid of smaller competition, and otherwise won't benefit anyone.
Ultimately, this question also involves one of humans. When I read a book on programming, and use some of the tricks I've just learned in a project of mine, do I violate the publisher's copyright? Do humans posess a "license to launder" that silicon training circuits lack?
The author and the book were published for the exact this reason, you learning stuff .
But in this case situation is different, someone creates a blackbox program and puts in all of your code that is under license X . Then they sell this blackbox program, and depending on what input people give to the blackbox it will output parts of your stuff. If your stuff is original then the chance is that most of it will be inside the blackbox and it will be spit out with the right inputs, and copyright and licenses are now washed.
Personally I am surprised on how amateurish this seems to be, seems like they used mostly text as input and not something more advanced , if MS did not put the properietary code they own in the blackbox then I am also suspicious.
I know several programmers working at multi million/billion dollar companies that use GPL/AGPL libraries within completely closed source codebases. Some of these products are shipped as DLLs/binaries rather than hidden behind web services, and even still, nothing has ever come out of it as far as license enforcement. Your company’s legal department will not look for these things proactively.
Copilot may be the Napster moment that changes this, though.
However This is one of the complaints that is leveraged by many in the Linux community, as Linux Foundation seems to take the stance you have outlined. Under no circumstance it seems will they enforce the licensing for the projects under their banner, that is sad but also not shocking since they are not Business Organization with some of the largest closed source companies in the world backing their existence...
> I know several programmers working at multi million/billion dollar companies who use GPL/AGPL libraries within completely closed source codebases. Some of these products are shipped as DLLs/binaries rather than hidden behind web services, and even still, nothing has ever come out of it as far as license enforcement. Your company’s legal department will not look for these things proactively.
There are companies (e.g., those producing mainly FOSS themselves) and organizations who do care, for example the SFC: https://sfconservancy.org/copyleft-compliance/enforcement-st...
You can also allow them to pursue violations on your, or a project, behalf: https://sfconservancy.org/copyleft-compliance/
Even if it may sound naïve, I still hope & believe they take all those proprietary leeches down, ideally bleeding those out, that still not want to comply.
By using GPL where you shouldn't you're putting your whole company at risk.
It's only a matter of time before lawyers come along smelling blood.
I think the offical position of most tech companies is to simply hand wave away any concerns the community(ies) may have about the legality, ethics, or anything else related to these disruptive technologies.
They want to push them as fast as possible into the market to make them impossible to economically remove "fixing" the issues later.
Apparently this doesn't prevent them publishing things in the Code Vault, or via this; that I might have wished to remove from their custody.
User agreements that amount to "you agree to be used," I guess.
If whatever stalker that I offended becomes unhappy that they might not have gotten my code off the net completely, are they gonna blame GitHub for it, or continue to try getting others to take actions against me? If some court sided with them and decided I had liability for publishing that code; would GitHub's possible resurrection of it be their responsibility, or mine?
EDIT: previously stated "after they offered me the option to do that or adopt a Code of Conduct on a single contributor project of mine."
Perhaps I'm wrong about the source of that; but I did delete my account and expected that they would no longer publish my code as such; others had forked things and that was as expected.
As in: they reached out to you directly and stated "you must adopt this CoC or we'll ban you from GitHub"? I can't find any other instances of people encountering this...
Can you please elaborate on this?
You were able to get pretty much that same output from an Enterprise install.
Thats probably the difference that explains the different reactions. At the end of the day, Google News gave traffic to publishers websites. Copilot is very different.
I've always thought the riskiest part of Google News are the tiny thumbnail images used to illustrate each story section. I wonder if they pay for those; I think they might.
My comment assumed that GPL meant the GNU Public License, where you are allowed to use the code, but are restricted to distribute binaries.
AGPL is a different license, and it's probably not legal that code with that license should have been used in their training set.
eevee’s tweet actually suggests the opposite: that even derivative works of GPL-licensed software must also use the GPL.
Thread on the tweet in question https://news.ycombinator.com/item?id=27687450
But as a temporary legal fix, couldn’t they put all the attributions in one large file and after each output of copilot write: “results based on the work of these people [link to attributions.txt].“
would that still be problematic?
Given that GitHub’s code runs on their own servers, doesn’t that protect them from needing to provide source code to anybody? I’ve always assumed running GPL code in the data center is fine.