Microsoft want court to toss lawsuit accusing them of abusing open-source code
reuters.com
reuters.com
https://www.zdnet.com/article/linux-developer-abandons-vmwar...
https://www.zdnet.com/article/vmware-sued-for-failure-to-com...
One extreme is AI is allowed to spit out copyrighted code verbatim as long as it technically goes through an AI first. Of course that defeats all open-source languages by adding a backdoor around them.
The other extreme is that AI is not allowed to spit out a single line of copyrighted code, in which case we'll have endless lawsuits to figure out if CodeGPT used a GPL-licensed fast inverse square root or if it used the public-domain fast inverse square root.
I think we'll land somewhere in the middle: If an AI regurgitates a "substantial" number of lines of code, then it's creators can be held liable (a.k.a. the "we'll know it when we see it" standard.)
Microsoft never changes. Always looking for a dishonest buck. Does 'Embrace, Extend, and Extinguish' ring a bell for younger players? Thought not.
> OpenAI, Microsoft want court to toss lawsuit accusing them of abusing open-source code
Seems pretty obvious to me but we'll see how it goes in the court.
> Held: 2 Live Crew's commercial parody may be a fair use within the meaning of § 107. Pp. 574-594.
The ruling states explicitly that commercial usage can be a determining factor in determining whether usage is fair or not, but that it does not in and of itself make the use "unfair".
To me, the latter scenario is much closer to what Github is doing with Copilot, which is one of the things the plaintiffs are alleging violates open source licenses.
Here's a good link explaining the Fair Use test: https://copyright.columbia.edu/basics/fair-use.html
> The fair use of a copyrighted work ... for purposes such as criticism, comment, news reporting, teaching (including multiple copies for classroom use), scholarship, or research, is not an infringement of copyright. In determining whether the use made of a work in any particular case is a fair use the factors to be considered shall include ...
(1)the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2)the nature of the copyrighted work; (3)the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4)the effect of the use upon the potential market for or value of the copyrighted work.
It's commercial, they use all of the code to built the model and the original code loses its value because you can get through Copilot. 3 of 4, depending on the original license it's 4 of 4 against fair use
That's not what "portion used" means. If you summarize a book then the portion used is <1%, not the entire book.
> the original code loses its value because you can get through Copilot
That's not even remotely true. You might get a fragment or two but you have to rebuild a program from scratch to replace it.
And as far as "nature of the copyrighted work" it's a completely different beast. It's a programming tool instead of whatever code was fed into it.
Only commerciality is a clear mark against it, and that factor is far from decisive by itself.
If you ask a AI to write 1000 unique summaries of the same book, each including a unique 0.1% copy of the book, what you have is a convoluted copying protocol. Asking the AI to write 1000, or 10000, or 10^infinitive number of fragments won't change the fundamentals of what is being done.
In the end you have to ask the question what a judge and jury will say. In the BitTorrent protocol you split a file into tiny fragments of between 32 kB and 16 MB, and in the beginning people did make the claim that such small fragments could not possible be copyrightable. A 32 kB portion of a 10GB movie is so small that it has no significant relationship with the original work. Courts disagreed and people went to jail.
Was the creator of the BitTorrent protocol liable? or the person transmitting the file?
Is Xerox liable for the copy made? or the person using the copier?
If Xerox would allow to copy money they would be liable.
Lobby organizations for rights owners presented their own theories, one was that concept of "making available" a copyrighted work. They argued that the upload was equally if not more guilty of infringement than the downloaded. They also accused the website owners for facilitating and enabling.
Then came the pirate bay case and a glaring issue struck the pirate movement. If technology can make creative solutions using code to bypass copyright, courts can in turn make creative solutions around law. The law that was used to charge the founders of the pirate bay was originally intended to combat bike bars when those places was used as headquarters for illegal gangs, a far step away from a website hosting files which enable two people on the internet to transfer files.
So we can go around and blame the AI for doing the infringement, or even the researcher who invented the math that created AI, but as with any creative technical solution around copyright we have to ask what creative solutions the lawyers and judges will make.
Copilot is basically a single interactive summary of all of github.
I would not currently worry about the threat of someone making a thousand wildly different copilots that deliberately memorize different fragments. Especially because I expect the real copilot to be tuned over time to reduce the number of fragments it picks up. But if such a person emerges, it's clear that they are the problematic actor.
The original license of the code used to train on doesn't really matter to the fair use question, that only matters once the fair use defense fails and the court has to decide a remedy.
Also, there is a difference here. A fair use quote from another work is a reference. It's not the thing, it's referring to the thing.
When copilot takes a chunk of code from another work to put into yours, it's using the thing directly, or rather, you are by using it.
It lacks citation which a quote would have, and instead of being a quote to discuss the other work "Dr Foo once said <remarkable genius insight>" copilot is like you writing a novel and simply copying Dr Foo's remarkable genius insight.
The size of the snippet doesn't matter, it's the usage and the lack of citation.
Even a cheap-ass totally doable collective citation like getting all the contributors to agree to have their works included, and then having a big list of all contributors somewhwere, and then each user just needs to say "includes code from copilot collective" They don't even have that, which would be good enough.
> Additionally, the district court determined that the commercial nature of Google's use weighed against its transformative nature. Although Kelly held that the commercial use of the photographer's images by Arriba's search engine was less exploitative than typical commercial use, and thus weighed only slightly against a finding of fair use, the district court here distinguished Kelly on the ground that some website owners in the AdSense program had infringing Perfect 10 images on their websites. The district court held that because Google's thumbnails "lead users to sites that directly benefit Google's bottom line," the AdSense program increased the commercial nature of Google's use of Perfect 10's images.
> In conducting our case-specific analysis of fair use in light of the purposes of copyright, we must weigh Google's superseding and commercial uses of thumbnail images against Google's significant transformative use, as well as the extent to which Google's search engine promotes the purposes of copyright and serves the interests of the public. Although the district court acknowledged the "truism that search engines such as Google Image Search provide great value to the public," the district court did not expressly consider whether this value outweighed the significance of Google's superseding use or the commercial nature of Google's use. The Supreme Court, however, has directed us to be mindful of the extent to which a use promotes the purposes of copyright and serves the interests of the public.
---
I will also draw attention to:
> The fact that Google incorporates the entire Perfect 10 image into the search engine results does not diminish the transformative nature of Google's use. As the district court correctly noted, we determined in Kelly that even making an exact copy of a work may be transformative so long as the copy serves a different function than the original work.
So, if the courts find in Microsoft and OpenAI’s favor (which remains to be seen despite the many armchair lawyers here), your license would mean jack squat.
They probably didn't rigorously track the licensing issue, but I'm pretty sure training a LLM is completely acceptable use of source under Freely licensed code. It would be somewhat amusing though if CoPilot is forced to spit out the license for every piece of code used to develop the derivative work, along with copyright notices and whatever else the licenses may require.
Copyrighted content can be used without the holder’s permission under “Fair Use”.
Don’t assume all code can be copyrighted. Purely functional expressions are not copyrightable. Code is math.
There’s a lot here to unpack.
There's a more interesting question about the copyright status of the code it outputs, since the language model is sort of like a compiler, but also not like a compiler since the output is based on other people's copyrighted code.
I feel a lot of people get caught up on the output code and completely ignore the fact that copilot itself is likely a massive copyright violation.
Copilot breaks the assumptions about the lossy nature of human memorization, so a lawsuit challenging the merits of the activity is at least warranted.
https://www.cnn.com/2021/04/05/tech/google-oracle-supreme-co...
> Writing for the Court, Breyer said that while it is difficult to apply traditional copyright concepts in the context of software programming, Google copied “only what was needed to allow users to put their accrued talents to work in a new and transformative program.”
> A world where Oracle was allowed to enforce a copyright claim, Breyer added, “would risk harm to the public” because it would establish Oracle as a new gatekeeper for software code others wanted to use.
The fair use tests that were used in the SCOTUS case, I believe, would fall on the side of "developers using GPT or Copilot to generate code do not generate substantial parts of the code and are below the amount of work needed to show sufficient creativity in writing it."
The example is https://horstmann.com/unblog/2010-11-15/NodePolicyImpl.html
If that is not a copyright violation and considered to be fair use, then the code generated by GPT or Copilot likely also falls in the the same bucket.
I don't necessarily agree with that, but that's my reading of the tea leaves.
The ability to copyright-launder via an API will lead to some interesting consequences for sure: I wouldn’t want to be elastic search or mongodb relying on source-available licensing if it comes about.
At least that's how compaq beat IBM and started all this monkey business.
Granted I could be full of it, I wasn't alive yet.
The same party can do it and not have it count as a copy, but it could be a copy if the party was not careful. So a company that wants to avoid potentially being sued will not allow one party to do both. So jimmaswell is correct to my understanding but a company may want extra legal armor/padding.
If that's the future people want, that's fine - but everyone should play by the same rules.
Would love to see this being done on decompiled proprietary code. Training done on it. And released into the wild.
But the amount of data necessary, and computing power to do it might not be available for the common person.