Just because you’ve made something cool doesn’t give you the right to harm others in the process.
If MS or OpenAI don’t think this is the case then they should have also included their private repositories.
Just because you’ve made something cool doesn’t give you the right to harm others in the process.
If MS or OpenAI don’t think this is the case then they should have also included their private repositories.
Is fair use on a massive scale still fair use? Courts generally think so, otherwise Google would have been out of business a long time ago.
"Microsoft does it, therefore it must be right" does not a sound argument make.
No, businesses are — not business. Not necessarily...
They are doing this to make sure that any lawsuit can be easily dismissed. It has nothing to do with the legality of the action (which sounds like fair use as the parent described it), and everything to do with the expense of a potential lawsuit compared to the cost effectiveness of simply telling developers “don’t do that”.
Most people think that the law has two shades: lawful vs unlawful. But the more practical distinction is expensive lawsuit vs dismissed lawsuit. This is the lens through which corporate lawyers see copyright and it might explain why so many programmers think that copilot is “obviously” breaking the law and “stealing” their code.
Questions of derivative works and fair use come up fairly frequently even in the open source world. This isn't solely a question of corporate lawyer posturing. I don't know any copyleft authors that would be okay with someone copying & pasting their code, making trivial changes, and saying it isn't a derivative work. Of course, their understanding of the law may be flawed. You'll get to find out in court.
You're right. A lot of this boils down to how much you want to spend in court proving your usage is just under fair use. We've moved beyond the question of ethics if you're intentionally violating a project's source license and relying on fair use to do whatever you want with the code. If you want to poke someone with a stick, you can't be surprised when they hit back. I contend what the OP described isn't clearly fair use (note I'm not saying that it clearly isn't fair use either). It ultimately doesn't impact me because I'm just not going to copy & paste code from projects without attribution and following the license, but I'd be worried about anyone reading that comment as objectively true.
It reminds me of the Google vs Oracle case. Apparently a court found that copying even a small amount of code in breach of its licence is not permitted. [0]
[0] https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_.... (ctrl-f for de minimis)
Courts have held that it doesn't apply to music, why do you think different rules apply to code?
Songs are different from code, in that the “hook” that makes the money may be only a few seconds long. There are many creative choices that a songwriter/producer can fit into just a few seconds: the harmony, melody, rhythm, lyrics, timbre, effects, ...
Whereas for code, the space of creativity is limited by functional considerations. A creative choice is protected by copyright but not all choices that programmers make are creative. Often the choices are limited by the API/interface or by efficiency considerations and it turns out that there’s only one good way to do something.
A function may be very intricate, yes, while still containing almost no creative value (e.g. a Vulkan setup function). Music doesn’t have an equivalent to this - the placement of every note is a creative act.
Yes, I too, and probably many people will do.
> In fact sometimes I copy and paste short passages and then rework them.
This I usually don't unless I check the license first. (Everybody ought to be allowed, but sometimes the license might not be.)
On the other hand, if you copy from private repositories, it quickly gets into the territory of stealing trade secrets.
There are just minor deviances, not relevant to this case, such as how long Disney bullied the countries to protect a work.
Software is usually considered a work. The AI needs to know if has permissions to copy and use the code, and then offer derived work on the proper terms and conditions. copilot doesn't do that. It might copy GPL code into non-GPL code, thus violating the GPL license, thus being an extreme risk.
In the US there have only been two extensions of copyright terms since Disney came into existence.
The first was in 1976, as part of a major overhaul of US copyright law to update the previous law (from 1909) to take into account the large changes in technology since then, and to make US law work more like the rest of the world to pave the way for the US later joining the Berne Convention. The changes for Berne compatibility included longer terms.
I assume Disney did support this, but only because as far as I can tell it had pretty widespread support. It had enough support that it would have passed even if Disney had adamantly opposed it.
The second was in 1998, and that was specifically a term expansion (as opposed to a term expansion like that of 1976 that was a side effect of harmonizing US law with the rest of the world). Europe had expanded terms a few years earlier, so the 1998 change in the US might have been motivated at least in part by harmonization, but I don't think the differences in terms between the US and the EU would have been enough to get it passed without some major interests pushing for it, so it is probably fair to give Disney a good part of the credit or blame for this one.
Here you see the countries which did not extend it the 2nd time to 70 years. https://en.wikipedia.org/wiki/List_of_countries%27_copyright...
There was of course no widespread support for these extensions, as all its arguments were flawed and not only violated logic but also several constitutions. https://de.wikipedia.org/wiki/Copyright_Term_Extension_Act#G... (the en version is mostly cleaned on these counter arguments)
Also, the open source community has far less leverage to apply pressure to Google than it does to GitHub. We may be able to do something about this.
The whole point of fair use is that the license doesn't matter. You can have a license that says I'm not allowed to use what you wrote for any purpose ever and I can still use it under fair use.
That said, copilot itself is not a replacement for your open source project that it was trained on. The code it generates may or may not be, but that's probably not Github's problem as far as copyright law is concerned.
I don't really think this argument passes muster.
It's just automating the copying and pasting (and slight reworking) of boilerplate code that would normally take me much longer to do, especially when I am working with a language I'm less familiar with but is necessary for my stack. I've literally never seen it suggest code that is more or less almost exactly what I would have come up with given a lot more time. In essence, it eliminates tedium- exactly the point of all of programming: Work elimination.
I do think there are ethical questions around whether it's right for google to digitise physical books without the permission of the authors, and keep them on their servers and make money from them without recompensing the authors. That's something an individual would not get away with doing, so it seems wrong that it's OK for google.
I have two questions:
1. Why have licenses, then?
2. What if I just use leaked sources of closed source software and call it fair use?
The default under copyright law is that any substantial copy is infringement.
A license is a legal document that grants someone permission to use a work that they otherwise would not have had.
However the law also gives its own permissions to use a work - it defines what is unlawful infringement and what is lawful fair use.
The code snippets that copilot generates look more like fair use than infringement. They are small, adapted to the destination context, and usually not direct copies of one source but more of an average of many different sources. And usually the programmer does not keep the suggestion that copilot suggests unmodified - the programmer does their own editing of the snippet afterwards to further tune it to the surrounding context.
2. What if I just use leaked sources of closed source software and call it fair use?
As pointed out upthread, if it the source code is leaked then there may be trade secret protections. The GPL specifically allows the code to be posted online, so by design it is not secret.
The reverse maybe true. I may be GPL'ing a code to prevent a useful algorithm from being buried deep inside a commercial code with an incompatible license. What makes it a "trade secret" level code? I have a 25 line algorithm which is worthy of its own paper. What if I open its reference implementation with AGPLv3+?
I have no problems with you reading the paper, and implementing it. I don't obfuscate my papers, but I put the implementation out with AGPLv3+. You can't use that in a codebase with an incompatible license. I expect and want you respect the license of my implementation.
> The code snippets that copilot generates look more like fair use than infringement. They are small, adapted to the destination context, and usually not direct copies of one source but more of an average of many different sources. And usually the programmer does not keep the suggestion that copilot suggests unmodified - the programmer does their own editing of the snippet afterwards to further tune it to the surrounding context.
Emphasis mine. First, there's no consensus on fair use, yet. Second they may be direct copies of the code. Third, they're remixed with other code pieces, which makes it a derivative work of many code pieces at once, then lastly, programmer re-derives the derived work. Which is clearly a derivative of GPL code, which brings in GPL license with itself (if what copilot derives the code from GPL licensed repositories, which it does).
I have no problem with Copilot as a technology. I have no problems with other licenses, which are not breached when used by Copilot and derived and used. The point which makes my blood boil is copilot using this GPL corpus, and don't admitting it publicly, breaching the terms of GPL en masse, and outright ignoring it. Then feeding this GPL derived code to any and all projects which pay for a copilot membership, and calling it a day.
The SCO vs IBM lawsuit was over only a few lines of code, after all.
I cant use a derivative of Mickey Mouse in my product, even if I change his colour and give him a hat, even if these changes were made by an AI. Why would it be different for code? I cab only use Mickey Mouse as fair use if its done for a specific barrow set of proposes (satire, news reporting etc).
I don't think that will happen but it might be interesting if it did.
This is the end of Microsoft's actual calculation.
I would like to point you to this: https://twitter.com/mitsuhiko/status/1410886329924194309 HN Comments at the time: https://news.ycombinator.com/item?id=27710287
That's not what copyright is protecting.
I believe github's lawyers would have had hundreds of hours of dicussion about this and at this point, they believe they are in the right, and anybody who disagrees should use the legal system to resolve the matter.
In the meantime, what it is and isn't doing wrt licenses seems to be poorly understood externally.
I don't think you can have your cake and eat it on this one.
It's kinda shocking that they think they can sell this, even providing it for free is extremely sketchy but at least complies with BSD/GNU/CC licensed stuff I guess.
Is every product user liable when a vendor ships some stolen code?
No, but the difference is the users of a product are typically not making and distributing copies. That’s not the case if you use someone else’s code in your project.
The user would be unlicensed, and in lieu of the vendor resolving this then the user would need to purchase licences to continue using the software legally (ie if a vendor gives you a pirate version of photoshop, you can’t just use it forever just because someone sold it to you).
There are usually clauses in enterprise software agreements that attribute liability for unlicenced components to the vendor for this reason. But ultimately if there isn’t a contract or the vendor vanishes, the user will need to go get a licence.
If you want to test the theory, I’ll send you a few images to put on your website, and when you get a claim through from the copyright owner you can try to argue that I sent it across without a copyright notice so I am liable ;)
All they could do is filter by the LICENSE file in the repo.
Unfortunately for them, by law copyright and license are determined by the authors and merely represented by a LICENSE file, which could be lying about both.
The court isn't going to accept that excuse when this goes to trial.
It's hard enough for us human to find our way in this mess, I've little hope for an AI.
But maybe it's just the first step. The final step being able to sell an AI that understands Copyright management. I'm sure there is a big market for that.
1) Require each repository to opt-in to be learned from.
2) Require any source file used for learning to have an SPDX license heading.
3) Have a list of approved permissive licenses to avoid any proprietary or copyleft arguments.
Using SPDX headings as the explicit guide would solve the problem of different code content using a different license within a project. An example being QtWayland: the client pieces are Proprietary/LGPL/GPL, whereas the compositor parts are Proprietary/GPL. That's not something you'd know from the license files at the root of the project (and post-6.3 they use SPDX instead of the prior license template heading).
Granted, this doesn't solve the problem of the chain of trust (is the individual publishing the code truly the copyright owner), but I think it would be a basic start for a program like this. The opt-in nature would make things... difficult, but I think that's a fair trade-off for something like this.
But until lawyers push for a standard that would make this part of their work irrelevant, I can't see how it could happen :)
Permissive or not doesn't matter. Public Domain or not is what matters. Permissive licenses still require you to propagate the copyright notice, which Copilot strips.
Unfortunately not. It's really stupid.
https://jtip.law.northwestern.edu/2021/05/28/copyright-issue...
However, even if infringement occurs during machine learning, training AI with copyrighted works would likely be excused by the ‘fair use’ doctrine.[ii] For example, in Authors Guild v. Google, Inc.[iii], Google had scanned digital copies of books and established a publicly available search function. The plaintiffs alleged that this constituted infringement of copyrights. The Second Circuit held that Google’s works were non-infringing fair uses because the purpose of the copying was highly transformative, the public display of text was limited, and the revelations did not provide a significant market substitute for the protected aspects of the originals.
The Second Circuit's tests listed in your citation specifically fail in this case. It's not highly transformative since it's just regurgitating snippets to be used in other competing works rather than applying the body of works to a different domain. And it's specifically to provide a market substitute for the protected aspects of the original works.
Additionally, none of this says 'its all great and it's on the user to figure it out'.
Lots of companies do not put their code in public repositories, granted I understand the perspective of violating a license, but the point is if you don’t want your code used by someone else (even with the risk of not getting credit, don’t know why that matters) then don’t make your repo public period.
To that point, what’s to stop GitHub from making a policy that states: “All public repositories will be utilized in AI training”?
The point is that it’s not respecting the license, not just that it’s not giving “credit”. If I release code under a GPL license, I damn well don’t want someone using that code under a license that’s not GPL-compatible, no matter how it got there.
def average(*numbers):
return sum(numbers)/len(numbers)
if not is it because it is too small? what’s the minimum line number that ownership kicks into? what if I change the function name and the variable names?That's as simple as that.
The point I'm trying to make is if something is under a copyleft license, you can't copy and paste it verbatim to something non-copyleft. It's what the license says.
Also, to be pedantic, the function I'm commenting on is pure maths, and you can't license/patent mathematics.
On the other hand, if there's some magic sauce of doing something, let it be 25 lines, what will you say? It's just 25 lines, so you can't license it? To be more pedantic, I actually have an algorithm, which is around 25 lines and does something novel. I've published a paper on it.
If I license the reference implementation with AGPLv3+, and you use it and close it, and if I can't go after you, what's the purpose of the license?
You can read the paper and try to implement it. It's free in that regard.
Isn’t there already precedent in other forms of IP, such as chord progressions in music, sentence length in literature, etc?
Is it again too trivial?
I don't know. That's my function's length.
> do you count comments?
No comments, no blank lines.
> can I codegolf a few lines to get below the limit?
You bet. But, if you copy my reference implementation, you need to get the license as well.
However, the research is on the open. Read it, implement it. That's no problem.
But, CoPilot is not reading my paper. It's reproducing my function verbatim, which is under a license which has share-alike mechanics.
There are no rules about the form of the code itself that governs whether or not someone owns it. Common sense applies. Sure you could "steal" very small, common, code snippets and get away with it; but that doesnt make it less wrong.
When a commercial entity explicitly does it, however, some times we can catch them. Like if they do it through algorithms that we more or less know how they work - i.e. the algorithm is using advanced control flow logic to copy and paste from it's training data set and copyrighted material is in that data set
Copyright really is not only concerned with what exactly is on the page, but also how you got there, and where the knowledge came from to get you there.
What if I read your codebase, and then years later while programming for myself I inadvertently use solutions you came up with while thinking I came up with it myself?
There really are no hard set rules, and this is something that is handled on a case-by-case basis based on whether or not a convincing argument can be made that you copied a novel idea from someone else and claimed it as your own.
We can argue the semantics of it all we want, but the subject area is an active battleground. Typically it only matters when money starts to get involved, since no one usually presses the issue or gets involved with random personal projects. So when an enterprise level company leverages that lack of caring into a proprietary pay-to-use project that operates by copying and pasting code from copyrighted material, then it seems like a case might be able to be made for it.
OSS knights: THE LICENSE.
MS: Aight, I guess we have a few lines of hq src to help out with…
Github: Same.
Other OSS people: We really don't care one way or the other.
As long as the word of the lincense was upheld for another 2 weeks before it ceased to matter for the rest of all time.
Jesus fucking christ. People. I get that oss licensing is dear to the collective hn heart – but, at best, it's completely irrelevant in regards to where this will inevitably lead, regardless of current questions/issues with license violations. You can (if all the repos of MS and Github are not enough to train this thing on, which is a laughable idea) even fucking buy additional source code if that's what it takes to strengthen Copilots legal foundation. The cost is insignificant. People will be happy to sell for super cheap. It's a non issue.
Why do you wilfully choose to be distracted instead of facing and thinking about the future together?