Judge dismisses DMCA copyright claim in GitHub Copilot suit
theregister.com
theregister.com
If I, a human, were to:
1. Carefully read and memorize some copyrighted code.
2. Produce new code that is textually identical to that. But in the process of typing it up, I randomly mechanically tweak a few identifiers or something to produce code that has the exact same semantics but isn't character-wise identical.
3. Claim that as new original code without the original copyright.
I assume that I would get my ass kicked legally speaking. That reads to me exactly like deliberate copyright infringement with willful obfuscation of my infringement.
How is it any different when a machine does the same thing?
That’s why I think the opposite of what you claim is true: if you were to do this, absolutely nothing would happen. When they do it, they will get sued over and over until the law changes and they can’t be sued, or they enter some mutually-beneficial relationship with the parties who keep suing.
Read up on the DMCA and the impact it has on e.g. nintendo emulators and the developers thereof
I'm skeptical Github Copilot reproducing a couple functions potentially used by some random Github project is going to be a threat to another party's livelihood.
When AI gets good enough to make full duplicates of apps I'd be more concerned about the source. Thousands of smaller pieces drawn from a million sources and being combined in novel ways is less worrying though.
The death of Citra wasn't really a deliberate action on the part of Nintendo, it was collateral damage. Citra was started by Yuzu developers and as part of the settlement they were not able to continue working on it. Citra's development had long been for the most part taken over by different developers, but the Yuzu people were still hosting the online infrastructure and had ownership of the GitHub repository, so they took all of it down. Some of the people who were maintaining Citra before the lawsuit opened up a new repository, but development has slowed down considerably because the taking down of the original repository has caused an unfortunate splintering of the community into many different forks.
There is some speculation Nintendo was involved with the death of the Nintendo 64 emulator UltraHLE a long time back, but this was never confirmed. If indeed they did go after UltraHLE, then this would just like Yuzu be a case of them taking down an emulator for a console they were still profiting from, as UltraHLE was released in 1999.
The most famous example of companies going after emulators is Sony, which went after Connectix Virtual Game Station and Bleem!. Both were PS1 emulators released in 1999, a period during which Sony was still very much profiting from PS1 sales. Sony lost both lawsuits and hasn't gone after emulators since.
In 2017, Atlus tried to take down the Patreon page for RPCS3, a PS3 emulator. However, Atlus only went after the Patreon page, not the emulator itself, which they did because of their use of Persona 5 screenshots on said page. The screenshots were simply taken down and the Patreon page was otherwise left alone. Of note is that Atlus is a game developer, so they were never profiting from PS3 sales. However, they were certainly still profiting from Persona 5 sales, which had only released in 2016.
These are the only examples I can remember. Did I miss anything?
> There is some speculation Nintendo was involved with the death of the Nintendo 64 emulator UltraHLE a long time back, but this was never confirmed.
iirc it got c&d but a case was never filed in court, the source code turned up eventually anyways.
Locks keep people honest. Unfortunately, software lockpicks have the unfortunate reality of being as easily distributable as the software itself.
I’m quite well read on the DMCA but admit you probably know far more about how Nintendo wields it.
Still, I suggest that it’s a lot more likely that GitHub is going to get sued than you or GP.
Finally, I believe using the legal system to bully independent software developers is, in legal terms, super lame. We are probably in the same side here.
You are probably more likely to be on the wrong end of a dmca take down request as a poor person since you dont have the resources to fight it, and its not about recovering damages just censorship.
Because intent matters in the law. If you intended to reproduce copyrighted code verbatim but tried to hide your activity with a few tweaks, that's a very different thing from using a tool which occasionally reproduces copyrighted code by accident but clearly was not designed for that purpose, and much more often than not outputs transformative works.
Training a model doesn't involve reproducing a copyrighted work, preparing a derivative work, distributing that work, or performing that work.
Fair use isn't required because none of the exclusive rights afforded by copyright apply.
That is the whole purpose and mechanism by which they operate.
Also the intent does not matter under law - not intending to break the law is not a defense if you break the law. Not intending to take someone's property doesn't mean it becomes your property. You might get less penalties and/or charges, due to intent (the obvious examples being murder vs manslaughter, etc).
But here we have an entire ecosystem where the model is "scan copyrighted material" followed by "regurgitate that material with mechanical changes to fit the surrounding context and to appear to be 'new' content".
Moreover given that this 'new' code is just a regurgitation of existing code with mutations to make it appear to fit the context and not directly identical to the existing code, then that 'new' code cannot be subject to copyright (you can't claim copyright to something you did not create, copyright does not protect output of mechanical or automatic transformations of other copyrighted content, and copyright does not protect the result of "natural processes", e.g 'I asked a statistical model to give me a statically plausible sequence of tokens and it did'). So in the best case scenario - the one where the copyright laundering as a service tool is not treated as just that, any code it produces is not protectable by copyright, and anyone can just copy "your work" without the license and (because you've said if you weren't intending to violate copyright it's ok) they can say they could not distinguish the non-copyright-protected work from the protected work and assumed that therefore none of it was subject to copyright. To be super sure though they weren't violating any of your copyrights, they then ran an "AI tool" to make the names better and better suit your style.
I am so sick of these arguments where people spout nonsense about "AI" systems magically "understanding" or "knowing" anything - they are very expensive statistical models, the produce statistically plausible strings of text, by a combination of copying the text of others wholesale, and filling the remaining space with bullshit that for basic tasks is often correct enough, and for anything else is wrong - because again they're just producing plausible sequences of tokens and have no understanding of anything beyond that.
To be very very very clear: if an AI system "understood" anything it was doing, it would not need to ingest essentially all the text that anyone has ever written, just to produce content that is at best only locally coherent, and that is frequently incorrect in more or less every domain to which it is applied. Take code completion (as in this case): Developers can write code without essentially reading all the code that has ever existed just so that they can write basic code, because developers understand code. Developers don't intermingle random unrelated and non-present variables or functions in their code as they write, because they understand what variables are and therefore they can't use non existent ones. "AI" on the other hand required more power than many countries to "learn" by reading as much as possible all code ever written, and then produce nonsense output for anything complex because they're still just generating a string of tokens that is plausible according to their statistical model - the result of these AIs is essentially binary: it has been in effect asked to produce code that does something that was in its training corpus and can be copied essentially verbatim, with a transformation path to make it fit, or it's not in the training corpus and you get random and generally incorrect code - hopefully wrong enough it fails to build, because they're also good at generating code that looks plausible but only fails at runtime because plausible sequence of tokens often overlaps with 'things a compiler will accept'.
Intent frequently matters a great deal when applying laws.
In the specific area of copyright law, it doesn't itself make the use non infringing, but it can absolutely impact the damages or a fair use argument.
I concluded that it was just completely impossible for a properly trained stable diffusion model to reproduce the works it was trained on.
The SD model easily fits on a typical USB stick, and comfortably in the memory of a modern consumer GPU.
The training corpus for SD is a pretty large chunk of image data on the internet. That absolutely does not fit in GPU memory - by several orders of magnitude.
No form of compression known to man would be able to get it that small. People smarter than me say it's mathematically not even possible.
Now for closed models, you might be able to argue something else is going on and they're sneakily not training neural nets or something. But the open models we can inspect? Definitely not.
Modern ML/AI models are doing Something Else. We can argue what that Something Else is, but it's not (normally) holding copies of all the things used to train them.
Thinking in terms of compression, the compression in generative AI models is lossy. The mathematical bounds on compression only apply to lossless compression. Keeping in mind that a small fraction of the training corpus is presented to the training algorithm multiple times, it's not absurd to suggest that these works exist inside the algorithm in a recallable form. Hence the NYT's lawyers being able to write prompts that recall large chunks of NYT articles verbatim.
$ ollama list
NAME ID SIZE MODIFIED
yi:34b ff94bc7c1b7a 19 GB 7 days ago
mistral:latest 61e88e884507 4.1 GB 2 months ago
mixtral:8x22b bf88270436ed 79 GB 2 months ago
llama3:70b be39eb53a197 39 GB 2 months ago
phi3:latest a2c89ceaed85 2.3 GB 2 months ago
dolphin-mistral:latest 5dc8c5a2be65 4.1 GB 2 months ago
yarn-mistral:7b-128k 6511b83c33d5 4.1 GB 2 months ago
yarn-mistral:latest 8e9c368a0ae4 4.1 GB 2 months ago
llama3:latest a6990ed6be41 4.7 GB 2 months ago
For comparison, here's some stable diffusion checkpoints. ComfyUI/models/checkpoints $ du -h .
6.5G breakdomainxl_v03d.safetensors
6.5G dreamshaperXL10_alpha2Xl10.safetensors
6.5G sd_xl_base_1.0.safetensors
5.7G sd_xl_refiner_1.0.safetensors
...
And I seem to recall there are some theoretical lower bounds on even lossy compression. Some quick back of the envelope fermi estimation gets me a hard lower bound of 5TB for "all the images on the internet"; but I'm not quite confident enough in my math to quite back that up right here and now.I'm not sure what your math is coming from and it seems trivially wrong. A single black pixel is a very lossy compression of every image on the internet. A picture of the Facebook logo is a slightly-less-lossy compression of every picture on the internet (the Facebook logo shows up on a lot of websites). I would believe that you can get a bound on lossy compression of a given quality (whatever quality means) only if you assume that there is some balance of the images in the compressed representation. There are a lot of assumptions there, and we know for a fact that the text fed to the GPTs to train them was presented in an unbalanced way.
In fact, if you look at the paper "textbooks are all you need" (https://arxiv.org/pdf/2306.11644) you can see that presenting a very limited set of information to an LLM gets a decent result. The remaining 6 trillion tokens in the training set are sort of icing on the cake.
I think you'll agree that it would be a bit absurd to threaten legal action against someone for storing a single black pixel.
OTOH Someone might be tempted to start a lawsuit if they believe their image is somehow actually stored in a particular data file.
For this to be a viable class action lawsuit to pursue, I think you'd have to subscribe to the belief that it's a form of compression where if you store n images, you're also able to get n images back. Else very few people would have actual standing to sue.
The black pixel won't get you sued, but the Facebook logo example I used could get you sued. Specifically by Facebook. There is an image (n = 1) that is substantially similar to the output of your compression algorithm.
That is sort of what Getty's lawsuit alleges. Not that every picture is recallable from an LLM, but that several images that are substantially similar to Getty's images are recallable. The same goes with the NYT's lawsuit and OpenAI.
I do realize the benefits of the 'compression' model of ML. Sometimes you can even use compression directly, like here: https://arxiv.org/abs/cs/0312044 .
I suppose you're right that you only need a few substantively similar outputs to potentially get sued already. (depending on who's scrutinizing you).
While talking with you, it occurred to me that so far we've ignored the output set o, which is the set of all images output by -say- stable diffusion. n can then be defined as n = m ∩ o .
And we know m is much larger than n, and o is theoretically practically infinite [1] (you can generate as many unique images as you like) , so o >> m >> n . [2]
Really already at this point I think calling SD a compression algorithm might be just a little odd. It doesn't look like the goal is compression at all. Especially when the authors seem to treat n like a bug ('overfit'), and keep trying to shrink it.
That's before looking back at the "compression ratio" and "loss ratio" of this algorithm, so maybe in future I can save myself some maths. It's an interesting approach to the argument I might try more in future. (Thank you for helping me to think in this direction)
* I think in the case of the Getty lawsuit they might have a bit of a point, if the model might have been overfitted on some of their images. Though I wonder if in some cases the model merely added Getty watermarks to novel images. I'm pretty sure that will have had something to do with setting Getty off.
* I am deeply suspicious of the NYT case. There's a large chunk of examples where they used ChatGPT to browse their own website. This makes me wonder if the rest of the examples are only slightly more subtle. IIRC I couldn't replicate them trivially. (YMMV, we can revisit if you're really interested)
[1] However, in practice there appear to be limits to floating point precision.
[2] I'm using >> as "much greater than"
I'm not aware of evidence that support that claim. If I ask ChatGPT "Give me a recipe for squirrel lemon stew" and it so happens that one person did write a recipe for that exact thing on the Internet, then I would expect that the most accurate, truthful response would be that exact recipe. Anything else would essentially be hallucination.
You can certainly try to hit a nail with a screw driver, but that doesn't make the screw driver a hammer.
Of course, if I ask a question that isn't as well served by the corpus, it has to do its best to interpolate an answer from what it knows.
But ultimately its job is to extract information from a corpus and serve it up with as much semantic fidelity to the original corpus as possible. If I ask how many moons Earth has, it should say "one". If I ask it what the third line of Poe's "The Raven" is, it should say "While I nodded, nearly napping, suddenly there came a tapping,". Anything else is wrong.
If you ask it a specific enough question where only a tiny corner of its corpus is relevant, I would expect it to end up either reproducing the possibly copyright piece of that corpus or, perhaps worse, cough up some bullshit because it's trying to avoid overfitting.
(I'm ignoring for the moment LLM use cases like image synthesis where you want it to hallucinate to be "creative".)
And you probably wouldn't want to - if I ask if donuts are radioactive and one person explicitly said that on the internet you probably aren't going to tell me you want it to spit out that answer just because it exactly matches what you asked. You want it to learn from the overwhelimg corpus of related knowledge that says donuts are food, people routinely eat them, etc etc and tell you they aren't radioactive.
If your analogy was you were a human who memorized every variation of a problem (and every other known problem) and there was a tiny perctange of a chance where you reproduced that exact varation of one you memorized, but then added an after the fact filter so you don't directly reproduce it...
It's more like musicians who basically copy a bunch of music patterns or chord progressions before then notice their final output sounds too similar to another song (which happens often IRL) then changes it to be more original before releasing it to the public.
This is mere assumption. AI is supposed to work like that, but that's a goal, and not the result of current implementations. Research shows that they do memorize solutions as well, and quite regularly so. (This is an unavoidable flaw in current LLMs; They must be capable of memorizing input verbatim in order to learn specific facts.)
> and there was a tiny perctange of a chance where you reproduced that exact varation of one you memorized
This is copyright infringement. Actionable copyright infringement. The big music publishers go after this kind of accidental partial reproduction.
> but then added an after the fact filter so you don't directly reproduce it...
"Legally distinct" is a gimmick that only works where the copyright is on specific identifiable parts of a work.
Changing a variable name does not make a code snippet "legally distinct", it's still copyright infringement.
This is Github Copilot after all. I use it daily and it autocompletes lines of code or generates functions you can find on stackoverflow. It's not letting giving you the source code to Twitter in full and letting you put it on the internet as a business under another name.
To just see how much they disliked it, youtube copyright strikes is basically a trained AI to detect music patterns to identify sound with slight variations or copyrighted songs and take videos down. Generating slight variations was one of the early method that videos used to bypass the take down system.
https://www.wardandsmith.com/articles/supreme-court-announce...
https://easlerlaw.com/software-computer-code-copyrighted#:~:...
The legal term you're looking for here is the "Abstraction-Filtration-Comparison" test; What remains if you subtract all the non-copyrightable elements from a given piece of code.
Even in the countries other than USA where algorithms have become patentable, that happened only due to USA blackmailing those countries into changing their laws "to protect (American) IP".
It is true however that there exist some quite old patents which in fact have patented algorithms, but those were disguised as patents for some machines executing those algorithms, in order to satisfy the existing laws.
In Zenimax vs Oculus they basically argued that a bunch of really abstract yet entirely generic parts of the code were shared, we are talking some nested for loops, certain combinations of if statements, and due to a lack of a qualitative understanding of code, syntax, common patterns, and what might actually qualify for substantively novel code in the courtroom, this was accepted as infringing. [1]
Point is, the legal system is highly selective when it comes to corporate interests.
[0] https://en.wikipedia.org/wiki/Substantial_similarity
[1] https://arstechnica.com/gaming/2017/02/doom-co-creator-defen...
I don't even think it's that. In recent cases like Oracle v. Google and Corellium v. Apple, Fair Use prevailed with all sorts of conflicting corporate interests at play. The Zenimax v. Oculus case very much revolved around NDAs that Carmack had signed and not the propagation of trade secrets. Where IP is strictly the only thing being concerned, the literal interpretation of Fair Use does still seem to exist.
Or for a more plain example, Authors Guild. v. Google where Google defended their indexing of thousands of copywritten books as Fair Use.
Arguably, the AI platforms have an even stronger case as their nominal goal is not to have their systems reproduce any part of the works verbatim.
AI platforms that replaces and directly compete with authors can not use the same argument. If anything, those suing AI platforms are more likely to bring up Authors Guild v. Google as a guiding case to determine when to apply fair use.
The more recent Warhol decision argues quite strongly in the opposite direction. It fronts market impact as the central factor in fair use analysis, explicitly saying that whether or not a use is transformative is in decent part dependent on the degree to which it replaces the original. So if you're writing a generative AI tool that will generate stock photos that it generated by scraping stock photo databases... I mean, the fair use analysis need consist of nothing more than that sentence to conclude that the use is totally not fair; none of the factors weigh in favor it.
To the extent that a user of co-pilot could induce it to produce enough of a copyrighted work to both infringe on the content (remember that algorithms are not protected by copyright) and substitute for the original by licensing in lieu of, I would expect the courts to examine that in the ways it currently views a xerox machine being used to create copies of a book. While the machine might have enabled the infringement, it is the person using the machine to produce and then distribute copies that is doing the infringing not the xerox machine itself nor Xerox the company.
Specifically in the opinion the court says:
>If an original work and a secondary use share
>the same or highly similar purposes, and the secondary use
>is of a commercial nature, the first factor is likely to
>weigh against fair use, absent some other justification for
>copying.
I find it difficult to come up with a good case that any given work used to train co-pilot and co-pilot itself share "the same or highly similar purposes". Even in the case of say someone having a code generator that was used in training of co-pilot, I think the courts would also be looking at the degree to which co-pilot is dependent on that program. I don't know off hand if there are any court cases challenging the use of copyright works in a large collage of work (like say a portrait of a person made from Time Magazine covers of portraits), but again my expectation here is that the court would find that while the entire work (that is the magazine cover) was used and reproduced, that reproduction is a tiny fraction of the secondary work and not substantial to its purpose.
Similarly we have this line:
>Whether the purpose and character of a use weighs in favor
>of fair use is, instead, an objective inquiry into what use
>was made, i.e., what the user does with the original work.
Which I think supports my comparison to the xerox machine. If the plaintiffs against Co-Pilot could have shown that a substantial majority of users and uses of Co-Pilot was producing infringing works or producing works that substitute for the training material, they might prevail in an argument that co-pilot is infringing regardless if the intent of github. But I suspect even that hurdle would be pretty hard to clear.
But in any case, Authors Guild is not the final word on the subject, and anyone trying to argue for (or against) fair use for generative AI who ignores Warhol is going to have a bad day in court. The way I see it, Authors Guild says that if you are thoughtful about how you design your product, and talk to your lawyers early and continuously about how to ensure your use is fair and will be seen as fair in the courts, you can indeed do a lot of copying and still be fair use.
Or put differently, if the Warhol image had used Goldsmith's image as a reference for a silk screen portrait of Steve Tyler, I'm not sure the case would have gone the same way. Warhol's image is obviously and directly derived from Goldsmith's image and found infringing when licensed to magazines, yet if Warhol had instead gone out and taken black and white portraits of prince, even in Goldsmith's style after having seen it, would it have been infringing? I think the closest case we have to that would have been the suit between Huey Lewis and Ray Parker Jr. over "I Want a New Drug"/"Ghostbusters" but that was settled without a judgement.
I do agree that Warhol is a stronger argument against artistic AI models, but it would very much have to depend on the specifics of the case. The AWF usage here was found to be infringing, with no judgement made of the creation and usage of the work in general, but specifically with regard to licensing the work to the magazine. They point out the opposite case that his Campbell paintings are well established as non-infringing in general, but that the use of them licensed as logos for soup makers might well be. So as is the issue with most lawsuits (and why I think AI models in general will win the day), the devil is in the details.
Substantial similarity refers to three different legal analyses for comparing works. In each case what the analysis is attempting to achieve is different, but in no case does it operate to prohibit similarity, per se.
The Wikipedia page points out two meanings. The first is a rule for establishing provenance. Copyright protects originality, not novelty. The difference is that if two people coincidentally create identical works, one after another, the second-in-time creator has not violated any right of the first. (Contrast with patents, which do protect novelty.) In this context, substantial similarity is a way to help establish a rebuttable presumption that the latter work is not original, but inspired by the former; it's a form of circumstantial evidence. Normally a defendant wouldn't admit outright they were knowingly inspired by another work, though they might admit this if their defense focuses on the second meaning, below. The plaintiff would also need to provide evidence of access or exposure to the earlier work to establish provenance; similarity alone isn't sufficient.
The second meaning relates to the fact that a work is composed of multiple forms and layers of expression. Not all are copyrightable, and the aggregate of copyrightable elements needs to surpass a minimum threshold of content. Substantial similarity here means a plaintiff needs to establish that there are enough copyrightable elements in common. Two works might be near identical, but not be substantially similar if they look identical merely because they're primarily composed of the same non-copyrightable expressions, regardless of provenance.
There's a third meaning, IIRC, referring to a standard for showing similarity at the pleadings stage. This often involves a superficial analysis of apparent similarity between works, but it's just a procedural rule for shutting down spurious claims as quickly as possible.
I'd say this is extremely relevant in this case.
to clarify - I thought you just had to negotiate with the cover artist about rights and pay a nominal fee for usage of the song for cover purposes - that is to say you do not negotiate with the original artist, you negotiate with a cover artist and the whole process is cheaper?
Say you want to make a recording of "Valerie" by the Zutons. You need permission (a license) from the songwriters (the Zutons presumably) to do this. You usually get this permission by paying a fee. Having done that, you can do your recording. Whenever that recording is played (or used) you will get a performance royalty and they will get a songwriting royalty.
Say you want to use a cover of "Valerie" by the Zutons in your film or whatever. Say the Mark Ronson version featuring Amy Winehouse. You need permission (a license) from the person who produced that version (Mark Ronson or his company) and will need to pay them a fee, some of which goes to the songwriter as part of their deal with Mark Ronson which gave him the license to produce his cover in the first place.
The Zutons don't have the right to sell you a license to Mark Ronson's version so if that's the version you want you have to negotiate with him. Likewise he doesn't have the right to sell you a license like the license he has (ie a license to do a recording/performance) so if you want that you have to negotiate with them.
The closest I could get to a situation like that would be if I told Band B do a cover of Song A for my movie and I paid the licensing costs as part of my deal with Band B, but still not the same as the parent poster's description.
To be able to do so, these softwares build a syntax tree of what your code snippet is, and compare the tree structure with similar trees in open source software without being fooled by variable names. To speed up the search, they also compute a signature for these trees so that the signature can be more easily searched in their database of open source code.
As a matter of fact, the Eclipse Foundation requires every contributor to declare that every piece of code is their own original creation and is not a copy/paste from other projects, with the exception possibly of other Eclipse Foundation or Apache Foundation projects because their respective licenses allow that. Even code snippets from StackOverflow are formally forbidden.
If I am not mistaken, in the Oracle-Google trial over Java on Android, at the end Google re-implementation of Java API on Android was considered fair-use, because Google kept the original "signatures" of the Java SDK API and rewrote most of the implementation with the exception of copying "0.4% of the total Java source code and was minimal" [1] However the trial came to this conclusion after several iterations in court.
[1] https://en.wikipedia.org/wiki/Google_LLC_v._Oracle_America,_....
The plain fact is that you can claim copyright on plenty of stuff that isn't copyrightable.
Consider AI model weights at all: they're the result of an automatic process and contain no human expression; almost by definition, model weights shouldn't be copyrightable, but people are still releasing "open source" models with supposed licenses.
The distinction is pedantic but important, IMHO. AI doesn't explicitly copy either.
Maybe, maybe not. It's not as simple as you made it out to be. If you write a book with lots of stuff and you got inspiration from other books, and even put in phrases wholesale, but modified to use your own character names instead, I'm not convinced you would lose.
The court would look at the work as a whole, not single pieces of it.
They would also check if you are just copying things verbatim, or if you memorize a pattern and emit the same pattern - for example look at lawsuits about copying music, where they'll claim this part of the music is the same as that part.
It's really not as cut and dry as you make it out to be.
Broadly, there is the distinction between expressive and functional code. [1]
And then there are the specific tests that have been developed by the courts to separate the expressive and functional aspects of software. [2] [3]
In practice it is very expensive for a plaintiff to do such analysis. For the most part the damages related to copyright are not worth the time and money. Plaintiffs tend to go for trade secret related damages as they are not restricted by the above tests.
There are also arguments to be made of de minimis infringements that are not worth the time of the court.
Most importantly the plaintiff fundamentally has the burden of proof and cannot just say that copying must have taken place. They need concrete evidence.
[1] https://en.wikipedia.org/wiki/Idea–expression_distinction
[2] https://en.wikipedia.org/wiki/Structure,_sequence_and_organi...
[3] https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
comes to mind. I bet most js devs have written this verbatim.
Violate one or two copyrights, get sued or DMCAed out of existence. Violate billions, on the other hand, and you magically become immune to the rules everyone else has to follow.
But if you want to argue that copyright is counterproductive, I completely agree. That's an argument for reducing or eliminating it across the board, fairly, for everyone; it's not an argument for giving a free pass to AI training while still enforcing it on everyone else.
Just because their lobbies tend to push the boundary of copyright into the absurd doesn't mean these industries aren't worth saving. There should be actually respectful lawmakers who seek for a balance of public and commercial interests.
Citation needed. There are many ways to make money from producing content other than restricting how copies of it can be distributed. The owner should be able to choose copyright as a means of control, but that doesn't mean nobody would create any content at all without copyright as a means of control.
As it is now, especially in the creative fields (which I am most knowledgeable about), the current system has allowed for a incredible flourishing of creation, which you'd have to be pretty daft to deny.
Slapping 3 lines in LICENSE.TXT doesn’t override the Berne convention.
that's not the argument. The fact that there currently are restrictions on producing derivative works is the problem. You cannot produce a star wars story, without getting consent from disney. You cannot write a harry potter story, without consent from Rowling.
There's actually a huge and thriving community of people publishing derivative works, in a not-for-profit basis, on Archive of Our Own. (Among other places.)
Yes, and none of those people are making a living at creating things. That's why they are allowed by the copyright owners to do what they're doing--because it's not commercial. Try to actually sell a derivative work of something you don't own the copyright for and see how fast the big media companies come after you. You acknowledge that when you say there are "restrictions" (an understatement if I ever saw one) on profiting from other people's work (where "other people" here means the media companies, not the people who actually created the work).
It is true that without our current copyright regime, the "industries" that produce Star Wars, Disney, etc. products would not exist in their current form. But does that mean works like those would not have been created? Does it mean we would have less of them? I strongly doubt it. What it would mean is that more of the profits from those works would go to the actual creative people instead of middlemen.
Again, not true. One of the most famous examples is likely Naomi Novik, who is a bestselling author, in addition to a prolific producer of derivative works published on AO3. Many other commercially successful authors publish derivative works on this platform as well.
> It is true that without our current copyright regime, the "industries" that produce Star Wars, Disney, etc. products would not exist in their current form. But does that mean works like those would not have been created? Does it mean we would have less of them? I strongly doubt it. What it would mean is that more of the profits from those works would go to the actual creative people instead of middlemen.
Speculate all you want about an alternative system, but you really don't know what would have happened, or what would happen moving forward.
Sorry, I meant they're not making a living at creating derivative works of copyrighted content. They can't, for the reasons you give. Nor can other people make a living creating derivative works of their commercially published work. That is an obvious barrier to creation.
No, the current system has allowed for an incredible flourishing of middlemen who don't create anything themselves but coerce creative people into agreements that give the middlemen virtually all the profits.
This is a false dichotomy. It's not "free speech" to copy someone else's video game and then sell it for your own profit. By "copy", in the old days that was literally copying the distribution CDs and providing a cracked keycode (it was not even a question of trademarks being close or what not. It's literally people taking the stuff, duplicating it, and selling it for their own profit. Eastern European mafia were greatly financed by this and ran this type of operation at industrial scale).
> Laws also hardly prevent sharing of copyrighted content, they only make it illegal.
Yeah, that's the point. Without that, everything is bootlegged. Imagine video games - they get bootlegged. DVDs, all bootlegged. Clothing bootlegged. Whatever your business is - bootlegged. Zero copyright is not a utopia of free speech, it is people ripping everyone else off. Per lived experience, I'm just saying the other extreme is not a utopia.
Nearly two hundred years ago one man warned everyone this would happen. Nobody listened. These are the consequences.
"At present the holder of copyright has the public feeling on his side. Those who invade copyright are regarded as knaves who take the bread out of the mouths of deserving men. Everybody is well pleased to see them restrained by the law, and compelled to refund their ill-gotten gains. No tradesman of good repute will have anything to do with such disgraceful transactions. Pass this law: and that feeling is at an end. Men very different from the present race of piratical booksellers will soon infringe this intolerable monopoly. Great masses of capital will be constantly employed in the violation of the law. Every art will be employed to evade legal pursuit; and the whole nation will be in the plot. On which side indeed should the public sympathy be when the question is whether some book as popular as “Robinson Crusoe” or the “Pilgrim’s Progress” shall be in every cottage, or whether it shall be confined to the libraries of the rich for the advantage of the great-grandson of a bookseller who, a hundred years before, drove a hard bargain for the copyright with the author when in great distress? Remember too that, when once it ceases to be considered as wrong and discreditable to invade literary property, no person can say where the invasion will stop. The public seldom makes nice distinctions. The wholesome copyright which now exists will share in the disgrace and danger of the new copyright which you are about to create. And you will find that, in attempting to impose unreasonable restraints on the reprinting of the works of the dead, you have, to a great extent, annulled those restraints which now prevent men from pillaging and defrauding the living."
https://www.thepublicdomain.org/2014/07/24/macaulay-on-copyr...
Are you implying that these three pillars will be able to produce anywhere near the current amount of content we produce?
How in the world where digital copies are effectively free to copy and infinitum would a creator reap any benefits from that network effect?
A modern equivalent would be famous YouTubers who all they do all day is "watch" other people's hard earned videos. The super lazy ones will not direct people to the original, don't provide meaningful commentary, just consumes the video as 'content' to feed their own audience and provides no value to the original creator. The position to kill copyright entirely would amplify this "just bypass the original source" to lower value of the original creator to zero.
Do you think the vast "amount of content we produce" is actually propped up by copyright? Have you ever heard of someone who started their career on YouTube due to copyright? On the contrary, how often have you heard of people stopping their YouTube career due to copyright, or explicitly limiting the content they create? I have only heard of cases of the latter. In fact, the latter partially happened to me.
> How in the world where digital copies are effectively free to copy and infinitum would a creator reap any benefits from that network effect?
You are making an assumption that people should reap (monetary) benefits for creating things. What you are ignoring is that the world where digital copies are effectively free is also the world where original works are insanely cheap as well. In this world, people create regardless of monetary gain.
To make this point: how much money did you make from this comment that you posted? It's covered by copyright, so surely you would not have created it if not for your own benefit.
Yes, and better quality content too as it doesn't need to be compromised as much to allow for commercial exploitation in the current model.
But these are also not the only ways to fund content. Patronage in particular does not need to be restricted to singular rich patrons but can be extended to any group of people that decide to come together to make something exist. This does already happen to some extend (e.g. Kickstarter) but is actually hobbled by copyright where the norm is that the creator retains all rights while individual contributors to the funding are restricted in how they are allowed to share the creation they helped realize.
> How in the world where digital copies are effectively free to copy and infinitum would a creator reap any benefits from that network effect?
By having fans willing to pay him to create new content.
- Big Corps that buy IP
- Patent Trolls
- Companies that fuck over artists
Of course, money is a huge motivator, but so is self-expression.
The current manifestation of copyright is about rent-seeking, not promoting innovation and creativity. That it may also do so is entirely coincidental.
Funny how both the rhetoric and intentions are the same after three hundred years.
The Supreme Court's decision was a bunch of bullshit around "well, y'know, people live longer these days, and some creators are still alive who expected these to last their whole lives, and golly, coincidentally this really helps giant corporations."
Sounds like the same concept as commonly said of "murderer vs conqueror".
Could probably be applied to many other fields for disruption too. Not the murderer bit (!), more the "break one or two laws -> scaled up massively to a potential new paradigm".
Bottom line, if you're doing something considered relevant to the national interest then that buys you a lot of leeway.
As we're seeing in court, that's a very interesting question. It turns out that the answers are very counter-intuitive to many.
It is immutable.
What are you going to do about it? Confiscate everyone's home gamer PCs?
Even in the most extreme hypothetical where lawsuits shutdown OpenAI, that doesn't delete the stable diffusion models that I have on my external hard drives.
The tech is out there. It's too late.
I cannot think of a better example of how futile copyright enforcement has been than the example that you just brought up.
> The most recently dismissed claims were fairly important, with one pertaining to infringement under the Digital Millennium Copyright Act (DMCA), section 1202(b), which basically says you shouldn't remove without permission crucial "copyright management" information, such as in this context who wrote the code and the terms of use, as licenses tend to dictate.
> It was argued in the class-action suit that Copilot was stripping that info out when offering code snippets from people's projects, which in their view would break 1202(b).
> The judge disagreed, however, on the grounds that the code suggested by Copilot was not identical enough to the developers' own copyright-protected work, and thus section 1202(b) did not apply. Indeed, last year GitHub was said to have tuned its programming assistant to generate slight variations of ingested training code to prevent its output from being accused of being an exact copy of licensed software.
So (not a lawyer!) this reads like the point about GitHub tuning their model is not a generic defense against any and all claims of copyright infringement, but a response to a specific claim that this violates a provision of the DMCA.
I don't know whether this is a reasonable defense or not, but your intuitions or mine about whether there is a general copyright violation or what's fair are not necessarily relevant to how the judge construes that very specific bit of legal code.
Pretty much the exact opposite of all these AI companies :p
I suspect that a lot of copyright violations are enabled by cut-and-paste and screenshot-taking functionality, and maybe we need to be careful with autocomplete, too? It's the user's responsibility to avoid this. We should be careful using our tools. Do users take enough care in this case? Is it possible to take enough care while still using CoPilot?
I've switched from CoPilot to Cody, but I use them the same way, to write my code. There's no particular reason to use CoPilot's output verbatim and lots of good reasons not to. By the time I've adapted it to my code base and code style and refactored it to hell and back, it's an expression of how I want to solve a problem, and I'm pretty confident claiming ownership.
Is that confidence misplaced? Are other people more careless?
By the same token, the machine alone can't download pirated movies. Yet the sites hosting those movies are targeted as the infringers.
There's a point at which foisting this responsibility on the users is simply socializing losses. Ultimately Copilot is the one serving the code up - regardless of the user's request. If the user then goes on to republish that work as their own it becomes two mistakes. It'll be interesting to see if any lawyers are capable of articulating that well enough in any of these lawsuits.
> Is that confidence misplaced? Are other people more careless?
I would say yes, for two reasons. One is that using code of unknown provenance means you're opening yourself to unknown legal risks. The second is if you're rewriting it fully (so as not to run afoul of easily spotted copyright) that's not actually "clean room" and you're still open to problems. I'd also wonder what the point of using a code writing LLM is anyways if you're doing all the authorship yourself. It seems like doing double the work.
As long as we're not required to register copyright there's no reason to think the above will play out. International copyright agreements are not limited to verbatim copies only.
This has already been done[1] in music, though in their case they released them to the public domain. Admittedly I think that was more of a protest than anything.
[1]: https://www.vice.com/en/article/wxepzw/musicians-algorithmic...
First: every human is per se doing that already. We have – to handwave – a "reasonable person" bar to separate violations versus results of learning and new innovation.
Second: You can be a holder of copyright and your creations result in copyrightable artifacts. Anything generated by the program has been held as uncopyrightable.
Legal protections for source code are still pretty fuzzy, understandably so given how comparatively new the industry is. That doesn't stop lawyers from racking up huge fees though, it actually helps because they need so much more prep time to debate a case that is so unclear and/or lacking precedent.
Literally the bank account behind the action...
But what if the generative AI were used to create music instead of code would the court have ruled differently?
CONSIDER:
In 2015, a federal judge order Thicke & Pharrell to pay 50% of proceeds to the Marvin Gaye estate for being “too similar” to the song, “Gots to Give It Up”.
Comparison and commentary: https://youtu.be/7_UiQueteN4?si=SkClbyBMOcucigRm
Comparison of both songs: https://youtu.be/ziz9HW2ZmmY?si=3_VZzfoLT-NrozoK
Choosing function signatures is an art form but after that "copying" is hard to judge.
I'd argue there are infinite ways to implement any function, just almost all of them are extremely bad.
It looks like wilful obfuscation because the obfuscation is so simplistic. But as the obfuscation gets increasingly sophisticated, it becomes ever harder to distinguish wilful obfuscation from genuine originality.
for the purposes of copyright, originality is not required, just different expressions. It's ideas (aka, patent) that require originality.
The 'sufficiently complex obfuscation' is exactly what people's brains go through when they learn, and re-produced what they learnt in a different context.
I argue that AI-training can be considered to be doing the same.
(1) You leave your employer, don’t take any code with you, start your own company, reimplement your ex-employer’s product from scratch, but you do it in a very different way (different language, different design choices, different tech stack, different architecture)
(2) You leave your employer, take their code with you, start your own company, make some superficial changes to their code to obscure your theft but the copying is obvious to anyone who scratches the surface
(3) You leave your employer, take their code with you, start your own company, start very heavily manually refactoring their code, within a few months it looks completely different, very difficult to distinguish from (1) unless you have evidence of the process of its creation
(4) You leave your employer, take their code with you, start your own company, download some “infringement obfuscation AI agent” from the Internet and give it your employer’s codebase, within a few hours it has transformed it into something difficult to distinguish from (1) if you didn’t know the history
(1) is unlikely to be held to be infringing. (2) is rather obviously going to be held to be infringing. But what about (3)? IANAL, but I suspect if you admitted that is how you did it, a judge would be unlikely to be very sympathetic. Your best hope would be to insist you actually did (1) instead. And then the outcome of the case might come down to whether the judge/jury believes your claim you actually did (1), or the plaintiff/prosecution’s claim you did (3).
And (4) is basically just (3) with AI to make it a lot faster and quicker. Such an agent likely doesn’t exist yet, but it could happen.
Timing is obviously a factor. If you leave your employer and launch a clone of their app the next week, everyone is going to think either you stole their code, or you were moonlighting on writing it (in which case they may legally own it anyway). If it takes you 12 months, it becomes more believable you wrote it from scratch. But if someone uses AI to launder code theft, maybe they can build the “clone” in a few days or weeks, and then spend a few months relaxing and recharging before going public with it
If I find a dollar on the sidewalk and put it in my wallets, is that stealing? If I punch a man getting change at a hotdog stand and a dollar falls on the sidewalk and then I put that in my wallet, is that stealing?
It doesn't matter what the scenario is after you stole code from your former employer, all actions are poisoned after.
Imagine the ex-employee open sources it, and I’m an innocent third party using that code base, ignorant of its unlawful origins. Am I infringing their ex-employers copyright (even if unintentionally)? For (2), obviously “yes”. But what about (3) or (4)?
I think the argument is that the machine is not doing that, or at least there isn't evidence that it is doing that.
Specificly no evidence that github is doing both 1 and 2 at the same time. There might be cases where it makes trivial changes to code (point 2) but for code that does not meet the threshold of originality. Similarly there might be cases with copyrighted code where the idea of it is taken, but it is expressed in such a different way that it is not a straightforward derrivitave of the expression (keeping in mind you cannot copyright an idea, only its expression. Using a similar approach or algorithm is not copyright infringement)
And finally, someone has to demonstrate it is actually happening and not just in theory could happen. Generally courts dont punish people for future crimes they haven't comitted yet (sometimes you can get in trouble for being reckless even if nothing bad happens, but i dont think that applies to copyrighg infringement)
Why? This is no different than copy pasting and modifying a bit of code from some documentation/other project/tutorial/SO. Surely if that were a basis for copyright infringement most semi-large software projects would be infringing on copyright.
I don't think anyone here should be willing to open the can if worms that is copy pasting small snippets of code and modifying them.
The judge seems to argue that the non-identical copies are at issue here and that they only happen under contrived circumstances. My moral opinion is that this is irrelevant and that even the defendant is the wrong person. Even verbatim copies of code snippets shouldn't be copyright infringement and suing the company providing the AI is wrong to begin with, as the AI or its providercan not possibly be the one to infringe.
My analogy is that if Copilot doesn't provide 100% code from another repository it is OK to be used by other people trained with code available on GitHub.
I also think it's not just copyright. It's simply not right to create a product on top of the collective work of all open source developers monetize them on the absurt scale Microsoft operates and never ever credit the original creators.
That’s the entire reason “clean room reverse engineering” is done.
Using nothing but the binary itself, work out how things are done. Making sure that the reverse engineers don’t even have access to any material that could look like it came from the other organization in question. And that it is provable.
"Pierre Menard, author of redis"
I know from experience that parents are aggressively pushing their children into STEM to maximize their chances of being economically secure, but, I really feel that we need a generation of philosophers and humanists to sift through the issues that our technology is raising. What does it mean to know something? What does authorship mean? Is a translated work the same as the original? Borges, Steiner, and the rest have as much to contribute as Ellison, Zuckerberg, and Altman.
It sounds fair from how the article describes it
Also, even if this weren’t the case you can’t sue for damages to other people (they’d need to bring their own suit)
It would be more correct to say Quake III Arena was released to the public as free software under the GPLv2 license.
Copyright infringement could be emitting the code in a manner that exceeds fair use.
The license gives you permission to utilize the code in a certain way. If Copilot gives you GPLed code that you then put into your closed source project, you have infringed the license, not Copilot.
> If you don't meet the conditions, it's still copyright infringement like before.
Licensing and copyright are two separate things. Neither has anything to do with the other. You can be in compliance with copyright, but out of license compliance, you can be the reverse. But nothing about copyright infringement here is tied to licensing.
To be clear: I am a person who trashed his Reddit account when they said they were going to license that text for training (trashed in the sense of "ran a script that scrubbed each of my comments first with nonsense edits, then deleted them"). I am a photographer who has significant concerns with training other models on people's creative output. I have similar concerns about Copilot.
But confusing licensing and copyright here only muddies waters.
It'd be a long bow to draw to say that what is akin to a search result of a snippet of code is "redistributing a software package".
That said the implementation doesn't appear to be totally trivial and copilot apparently even copies the comments which are almost certainly copyrightable in themselves.
https://x.com/StefanKarpinski/status/1410971061181681674 https://github.com/id-Software/Quake-III-Arena/blob/dbe4ddb1...
However a twitter post on its own isn't evidence a court will accept. You would need the original poster to testify that what is seen in the post is actually what he got from copilot and not just a meme or joke that he made.
Also the plaintiffs in this case don't include id-Software and there is some evidence that id-Software actually stole the fast inverse sqrt code from 3dfx so they might not want to bring a claim here anyways.
When it was reported, I was able to reproduce it myself.
Absolutely there were a few outliers where a judge might want to look more closely. I'd be surprised if -under scrutiny- there wouldn't be any issues whatsoever that OpenAI overlooked.
However, it seemed to me that over half of the NYT complaints were examples of using the -then rather new- ChatGPT web browsing feature to browse their own website. In the case, they then claimed surprise when it did just what you'd expect a web browsing feature to do.
All the plaintiffs would need to do is provide evidence that copywritten code was produced verbatim. This includes showing the copyrighted code on GitHub, showing copilot reproducing the code (including how you manipulated copilot to do it), showing that they match, and showing that the setting to turn off reproduction of public code is set.
It makes no difference who owns the copyrighted code, it need only be shown that copilot is violating copyright. Microsoft can't say "uhh that doesn't count" or whatever simply because they own a company that owns a company that owns copyright on the code.
i agree from a philosophical pov, but this is clearly not the case in law.
https://en.wikipedia.org/wiki/Abstraction-Filtration-Compari...
Open source licenses allow sharing under certain conditions.
Rightly so, you have to show some sort of damage to sue someone, not just theoretical damages.
1. The copilot team rushed to slap a copyright filter on top to keep these verbatim examples from showing up, and now claims they never happen.
2. LLMs are prone to paraphrasing. Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it. The copyright filter is only a legal protection, not a practical protection against the issue of copyright infringement.
Everyone who knows how these systems work understand this. The copilot FAQ to this day claims that you should run copyright scanning tools on your codebase because your developers might "copy code from an online source or library".
Github has it's own research from 2021 showing that these tools do indeed copy their training data occasionally: https://github.blog/2021-06-30-github-copilot-research-recit...
They clearly know the problem is real. Their own research agreed, their FAQs and legal documents are carefully phrased to avoid admitting it. But rather than owning up to the problem, it's "Ner ner ner ner ner, you can't prove it to a boomer judge".
Isn't that akin to destruction of evidence?
In spirit? ... Probably?
Unlike most LLMs, Github copilot can trivially solve their copyright problem by just using only code they have the right to reproduce.
They have a giant corpus of code tagged with license, SELECT BY license MIT/Equivalent and you're done, problem solved because those licenses explicitly grant permission for this kind of reuse.
(It's still not very cash money to take open source work for commercial gain without paying the original authors, and there's a humorous question if MIT-copilot would need to come with a multi-gigabyte attribution file, but everyone widely agrees it's legal and permitted.)
The only reason you'd hack a filter on top rather than doing the above is if you'd want to hide the copyright problem. It's an objectively worse solution.
Absolutely not trivial, in fact completely impossible by computer alone. You can't determine if you have the right to reproduce a piece of code just by looking at the code and tags themselves. *Taps the color-of-your-bits sign.*
* I can fork a GPL project on Github and replace the license file with MIT. Okay to reproduce?
* If I license my project as MIT but it includes code I copied inappropriately and don't have the right to reproduce myself, can Github? (No) This one is why indemnity clauses exist on contracted works.
* I create a git repo for work and select the MIT license but I don't actually own the copyright on that code and so that license is worthless.
The people that think Copilot is infringng their copyright would be happy with that I would think? Unless they take a much stricter definition of fair use than current courts do.
Is taking away a drunk driver's keys (before they get in the car) destruction of the evidence of their drunk driving?
In the current case - its unclear if any crime took place at all, it seems clear that the primary intent was to prevent future crime not hide evidence of past ones. Most importantly the past version of the app is not destroyed (presumably). Github still has the version of the software without the copyright filter. If relavent and appropriate, the court could order them to produce the original version. It can't be destroying evidence if the evidence was not destroyed.
Well if the copyright filter is working they indeed aren't happening. Putting in safe gaurds to prevent something from happening doesn't mean you're guilty of it. Putting a railing on a balcony doesn't imply the balcony with railing is unsafe.
> LLMs are prone to paraphrasing. Just because you filter out verbatim copies doesn't mean there isn't still copyright infringement/plagiarism/whatever you want to call it
Copyright infringement and plagerism are different things. Stuff can be copyright infringement without being plagerized, and can be plagerized without being copyright infringement. The two concepts are similar but should not be conflated, especially in a legal context.
Courts decide based on laws, not on gut feeling about what is "fair".
> They clearly know the problem is real
They know the risk is real. That is not the same thing as saying that they actually comitted copyright infringement.
A risk of something happening is not the same as actually doing the thing.
> "Ner ner ner ner ner, you can't prove it to a boomer judge".
Its always a cop-out to assume that they lost the argument because the judge didn't understand. I suspect the judge understood just fine but the law and the evidence simply wasn't on their side.
Doesn't mean you weren't, at some point, guilty of it, either. It doesn't retcon things.
After all, you yourself probably cannot prove that you didn't commit the same offense at some point in time in the past. Like Russel's teapot, its almost always impossible to disprove something like that.
Actually, it does. The production of the output is what matters here.
People do clean room implementations because of paranoia, not because it's actually a necessary requirement.
The literal act of making modifications isn't infringement until you distribute those modifications -- and we're talking about a situation where you've changed the code enough that it isn't considered a derivative work anymore (apparently) so that's kosher.
> you already have license to access the code
This isn’t access, that occurs before the AI is trained. It’s access > make copy for training > AI does lossy compression > request unzips that compression making a new copy > process fuzzes the copy so it’s not so obvious > derivative work sent to users.
GitHub didn’t just copy open source code they copped everything without respect to license. As such attribution which may have allowed some copying isn’t generally relevant.
Really a public repo on GitHub doesn’t even mean the person uploading it owns the code, if they needed to verify ownership before training they couldn’t have started. Thus by necessity they must take the stance that copyright is irrelevant.
>You grant us and our legal successors the right to store, archive, parse, and display Your Content, and make incidental copies, as necessary to provide the Service, including improving the Service over time
>This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service, except that as part of the right to archive Your Content, GitHub may permit our partners to store and archive Your Content in public repositories in connection with the GitHub Arctic Code Vault and GitHub Archive Program.
https://docs.github.com/en/site-policy/github-terms/github-t...
I think the important questions are (1) whether "the Service" includes Copilot, and (2) whether GitHub is selling users' content with Copilot.
For (1), I'm unhappy to admit Copilot probably does fall under "the Service," which is nebulously defined as "applications, software, products, and services provided by GitHub." But I'll still say that users' could not agree to this use while GitHub was training The Copilot model but hadn't yet announced it. At that time, a reasonable user would've believed GitHub's services only covered repository hosting, user accounts, and the extra features attached to those (issue trackers, organizations, etc).
GitHub could defend themselves on point (2) by saying they aren't selling the code, instead selling a product that used the code as input. But does that differ much from selling an online service that relies on running user code? The code is input for their servers, and it doesn't need to be distributed as part of that questionable service. But it's a clear break from the TOS.
If you copy a whole book and do the same, there’s still lines-3 infringement left.
More than that: the fact that they claimed it wasn't possible before adding the filter, to filter out the thing that said wasn't possible. This doesn't help me trust anything else they might say or have already said.
My take on that was always: if it isn't possible, then why are MS not training the AIs on their internal code (like that for Office, in the case of MS with their copilot product) as well as public code? There must be good examples for it to learn from in there, unless of course they thing public code is massively better than their internal works.
Since you really need to work hard to make the AI spit out anything verbatim, and you have no knowledge of their internal code, how could you ever prove or deny it?
Because if they were, they would have said.
It would be an excellent answer to the concerns being discussed here: “we are so sure that there is nothing to worry about in this regard, that we are using our own code as well as the stuff we've schlepped from github and other public sources”.
The main issue, as I see it, is that they took copyrighted material and made new commercial products without compensating (let alone acquiring permission from) the rights holders, ie their suppliers. Specifically, they sneaked a fair use sticker on mass AI training, with neither precedent nor a ruling anywhere. Fair use originates in times before there were even computers. (Imo it’s as outrageous as applying a free-mushroom-picking-on-non-cultivated-land law to justify industrial scale farming on private land.) That’s what should be challenged.
I wonder, if MS and OpenAI win, does that mean it will be legal for anyone to take the leaked source code for a proprietary product, train an LLM on it, and then ask the LLM to emit a version of it that is different enough to avoid copyright infringement?
That would be quite the double-edged sword for proprietary software companies.
Someone is likely to design an LLM that is specifically trained to do exactly that.
Lots of money to be made...
That should be fun...
To put another way, the motivations to produce art in another artist's style can still land the artist/buyer in legal trouble regardless of fair use.
I have to imagine that it's likely quite popular to sell AI generated art that mimics or copies existing works.
Not that it’d look anything like the artist you are copying, but it’s a fun idea.
Scale matters, and the scale that computers/these AIs operate under are absurd compared to a person doing it manually.
The work of a person can be mitigated and a person can be held accountable for their actions.
Much of our society operates on the idea that we don’t need to codify and enforce every single good or bad thing due to these reasons; and having such an underpinning affords us greater personal freedom.
I believe that this kind of generative AI is bad because it approximates human behavior at an inhuman scale and cannot be held accountable in any way. This upends the entire social structure upon which humans have relied to keep each other in-check since the advent of the modern concept of "justice" beginning with the Code of Hammurabi.
In essence: Because you cannot punish, rehabilitate or extract recompense from a machine, it should not be allowed in any way to approximate a member of society.
This logic does not apply to machines that "automate" labor, because those machines do not approximate human communication - they do not pretend to be us.
Should those machines then be subject to your same philosophies? I'd suspect you'd say "that's different" somehow but it is only because you are alive at this moment and these machines have been normalized to you that you do not care about them. Were you to be born in a few centuries, you would likely feel the same way most do about the prior machines, and indeed, you'd be hard pressed to find anyone who think that future generation's AI (probably simply called technology then) is problematic as you do today. Recency bias is one hell of a drug.
Even then, their work as output can matter but that doesn't necessarily mean they (should) have a per se right to their work without other people also using it, especially in cases where their work is not used as outputs directly, which is what plagiarism is. If that were the case, no one could learn from a other's work, regardless of whether that one is a person or a computer.
I don't think there is a way to continue this particular branch of this argument without devolving into a debate on the value of human life like a couple of Macedonian philosophers - suffice to say, my point of view is that the work of others has intrinsic value tied to intent, and machines do not have intent.
If no output of humans has intrinsic value, then once machines can approximate humans sufficiently there is no reason for humans to exist - and that is an outcome that I, as a human, reject with all of my being.
The reason for humans existing is not because of the output they produce (indeed, that is dystopic), humans have worth inherently, regardless of what they output. This is also what nihilists have figured out, so maybe that is something you should look into if you seriously have such an opinion as expressed in your last paragraph.
[0] https://news.ycombinator.com/item?id=40919253&p=2#40920318
You aren't allowed to use photos featuring a non-consenting person to, for instance promote a product.
You are allowed to use photos including a non-consenting person.
There's a lot of complicated law, differing between different jurisdictions to cover this question, and to balance the needs of the public with commercial desires. It's not as simple as you make it sound, and there's no reason we should just default to bending over backwards for commercial interests.
Laws exist to serve society, not the other way around.
This is my whole point. There isn't a single, one-size-fits-all rule that a five year old can comprehend that describes any particular country's legal framework around the many, many different dimensions of tension between public and private interests on this incredibly broad question.
And none of the existing frameworks fit the new use cases well, and we should probably have an open political debate about what we want to do going forward.
Being an ass is generally not illegal. Particular behaviours might be, but no legal or social system intends to censure you for every possible one, and most people who are experts in law or ethics don't believe that they should.
If you identify particular problems with the particular paparazzi laws in your country, that's an interesting conversation, and maybe, if framed well, an interesting data point for this discussion, but is not in itself the 'last word' on it. Just because you can torture an analogy, doesn't mean the analogy has a lot of power.
Perplexity AI.
How does this describe Perplexity AI more than any other LLM?
Perplexity is in the business of using an LLM to paraphrase existing content, then serving that up as their own "work" in a way that directly harms the original content they took.
It's not even a question of "Is AI training copyright infringement", they're just doing copyright infringement with AI. And it's horribly common already.
https://www.theverge.com/2024/6/27/24187405/perplexity-ai-tw...
Somehow I feel if it was "Adobe vs dev that claims his code was spit by copilot" it would not end the same.
> Specifically, the judge cited the study's observation that Copilot reportedly "rarely emits memorized code in benign situations, and most memorization occurs only when the model has been prompted with long code excerpts that are very similar to the training data."
That almost sounds like it'd be fine to train an "art transformation model" which takes an image and transforms it, which for all the frames of a specific Disney movie just so happen to output the very next frame...
There is a reason a famous AI model architecture is called transformer, it is pretty much optimised to be good at transforming artistic and intellectual works.
It would probably be legal to do this, as long as no one could reasonably show that you intentionally trained the LLM on said leaked source code with the intent to reproduce the product.
Of course, civil suits could be another matter entirely. If you pick a product to rip off that's owned by a multi-billion dollar company, all that can save you is the ethical limits of their legal team's consciences.
Far less people care about Bing or Cortana.
They cloned the bios by observing how it behaved and writing code that behaved the same way. Nobody even looked at the bios code.
To do something similar with AI, you really need to train one AI on the source code and then have it explain that code to a second AI that never saw the original code.
Hilarity will ensue :)
But if it was made public and then if an unrelated third party were to re-write the code in such a way that it was non-infringing, then it would be non-infringing. That’s just a tautology.
The LLM has nothing to do with it, and isn't required here.
> A few devs versus the powerful forces of Redmond – who did you think was going to win?
I hate that kind of obnoxious "journalism". Sometimes the little guy is actually wrong. To clarify, I'm not commenting on the specifics of this case, I just hate how fake our online discourse has been by appealing to "big guy evil" before even bringing up the specifics of the case.
I think it merely implies MS has more resources to throw at the legal case.
I'm not claiming Microsoft doesn't have tons of resources, I'm claiming that the plaintiffs attorneys should be sufficiently funded that the difference in outcomes is negligible.
aka the plaintiffs were wrong and had no idea what they were talking about
There is actually zero evidence that the judge issued his ruling based on Microsoft's superior legal team, so why even put that sentence in there anyway?
He is, sometimes. Also sometimes, the moon passes exactly between the sun and Earth, a new star appears in the sky, the magnetic field of our planet reverses, a proton decays (jury is still out on that one, actually). Etc.
Tools like Copilot are plagiarism machines. We know the data they're being trained on, and a conclusion of "that's plagiarism" is not - or anyway should not be - controversial. I'm not terribly against the notion of a plagiarism machine but I am against the owners of such machines reaping profits from them to the exclusion of the people who provide the source material. This is theft.
More importantly, getting back to big guys and little guys: big guys gang up on little guys all the time. It's usually how they get to be big. They tend to be the ones who realize that working together against the rest of us is to their benefit. So, in the interest of pushing back on that a little, and recognizing that I am after all a fellow "little guy" (figuratively speaking anyway), I tend to support the "little guy" unless I have overwhelming evidence confirming that they are, in fact, both wrong and that supporting them anyway would be against my best interest. Neither is the case, here.
At any rate, the subtitle here references a pretty ubiquitous and, I'm happy to report, increasingly well-known and understood facet of our economic and social institutions, which is that they absolutely positively do not work for us or further our interests in any sense.
And obnoxious individuals gum up enterprises. It's lazy to the point of dismissal to conclude based on bigness.
EDIT: And by "win" I mean not who the judge will side with, but who will end up chugging along fine financially and who will end up broke.
I can certainly agree with that sentence, but that is definitely not how the Register was referring to "win" (they clearly just meant the judicial outcome), so it's obnoxious to imply the legal ruling went Microsoft's way solely due to their greater resources.
My maintainable code gets published, my nightmares get banished to private repos so no one else thinks it's a good idea to replicate.
If I were Microsoft, I’d really be concerned that I’m going to kill my golden goose by causing a large-scale exodus from GitHub or open source development more generally. Another idea I’ve considered is publishing boatloads of useless or incorrect code to poison their training data.
As I see it, people should be able to restrict how people use something that they gave them. If some people prefer that their code is not used to train LLMs, there should be a way to enforce that.
I think this is a rather radical approach. You're undermining the OSS movement because you dislike Microsoft (I do too). I think adding a clause or dual licensing your work is more effective at stopping big-tech funded AI crawlers than just not adhering to open source.
You can host your code on sourcehut or Codeberg (Forgejo), you don't NEED to host it on a Microsoft owned platform.
Not everyone is multi-generationally rich or absurdly frugal. Most people like having good jobs.
> The anonymous programmers have repeatedly insisted Copilot could, and would, generate code identical to what they had written themselves, which is a key pillar of their lawsuit since there is an identicality requirement for their DMCA claim. However, Judge Tigar earlier ruled the plaintiffs hadn't actually demonstrated instances of this happening, which prompted a dismissal of the claim with a chance to amend it.
So, the problem is really one of the lack of evidence, which seems... like a pretty basic mistake from the plaintiffs?
They could've taken a screencap video back when Copilot still produced code more verbatim, and used that as evidence, I assume.
Found this: https://github.com/non-ai-licenses/non-ai-licenses
Legally sound or not, these should at least prevent your code from being included in Copilot's training data, hopefully without affecting any other use case. I'm going to use one of these next time I start a new project.
Has microsoft said this or something?
Turns out I was wrong. They don't care.
https://web.archive.org/web/20210708165143/https://twitter.c...
That doesn't mean anyone has to follow it.
If it's legal to train on other people's stuff, without their permission, this would still apply to your code even if your code includes a license that said "I double extra declare that you can't train AI on this!!".
Also, remixes almost always do contain verbatim lyrics and/or samples from the original song. LLM output isn't supposed to contain verbatim copies, but I've been told that sometimes it does. (I don't know much about LLMs and I don't think Copilot is useful. I want my 2010-era Intellisense back, when it was extremely fast and predictable.)
Colloquially, I generally expect a remix to be comprised of the original instrumentals/beat (potentially edited, but virtually nothing actually new added), potentially new lyrics, and to still be recognizable as the original.
The "still be recognizable as the original" part is a huge problem for fair use, and why I don't think remixes generally qualify. If it doesn't sound like the original then it's not a remix, but if it does sound like the original it can't be fair use.
I think the underlying issue is the resulting work, not the process that went into creating it. I think (but am in no way sure) that copying parts of songs would be fine if you did something to them so they aren't recognizable as the original.
As an example, if I take a song by the Beatles and repeatedly compress it until it's entirely compression artifacts, I would bet that I could publish that. I don't think it would matter that I started with a copyrighted work, what matters is that my finished product bears no resemblance to any other copyrighted work.
That would mean it's just a normal "is this work too similar to existing works?" standard applied to humans as well.
There is still an ancillary question of whether it's okay to train on copyrighted music, but that's really a different question than whether the works it creates infringe.
If I made an “advanced music engine” which rips Taylor swift files and duplicates them, I would be sued to oblivion. Why does calling it an AI suddenly fix that?
They should have to train them on information they legally own.
"You train them by comparing the output to the original." To the best of my knowledge this isn't correct; can you expand or cite a reference?
"You train them by comparing the output to the original." ->
You train neural networks by producing output for known input, comparing the output with a cost-function to the expected output, and updating your system towards minimizing the cost, repeatedly, until it stops improving or you tire of waiting. Cost functions must have a minimal value when the output matches exactly the expected to work mathematically. Engineering-wise you can possibly fudge things and they probably do so ... now.
I don't agree with your critiques. It isn't an oversimplification, published code literally works as stated.
Ah I see what they meant by that statement. It is true that supervised learning operates on labelled input/output pairs, and that neural networks generally use gradient descent/back propogation. (Disclaimer: it's been a few years since I've done any of this myself so don't quite remember it that well, and the field has changed a lot). Note since the parameter space of the neural network is usually _significantly_ smaller than the training data set, a network will not tend to minimise that cost function near 0 for an individual sample since doing so will worsen the overall result. There is inherent "fudging", although near identical output can potentially happen. The statement here is more reasonable and similar to the training process than the first.
It still doesn't.
This would mean excluding non-github members, and excluding members that opt out.
Wait, I'm forced to use Teams at work but Microsoft employees are on Slack?!
How did they reach this conclusion? How can you prove that it never copies a code snippet verbatim, versus just showing that it does for one specific code snippet? The latter is a lot easier to show, but I don't know what is it exactly that the prosecution claimed. I guess the size of the copy also matters in copyright violations?
Legal proof is I think different (not a lawyer). They're more pragmatic. If, observing a lot of cases where it does not verbatim copy, and, if an expert provides a reasonable argument as to why it is unlikely to verbatim copy, that is enough legal proof for a judge to conclude that the output is not identical enough to the developers copyrighted code.
The function of the Patent System is to incentivize search for solutions by temporarily securing exclusive right to market novel devices and processes for the discoverer.
Everything requires attention to be seen, once somethign becomes "obvious" is fully determined where you're looking and the scope you're zoomed in on.
E.g. "matter is solid" until you zoom in and realize matter is mostly made up of space.
And bringing things more expediently is the actual opinion here, unsupported, where arguably it actually slows down not only progress but the value of that progress not being as widely distributed as it otherwise would be.
If your code is readable, the public can learn from it.
Copyright doesn't extend to function.
People have the right to learn non-copyrightable elements from your code.
The claim is that AI learns copyrightable elements.
I agree it's certainly possible for AI to produce infringing output.
Nevertheless, people don't have the right to enforce a limitation on training.
If I "process data" by doing a word count of a book, and then I publish the number of words in that book (not the words themself! Just a word count!) I haven't created a derivative work.
Processing data isn't automatically infringement.
You think that a model that's capable of being prodded into producing an infringing output in addition to all the other non-infringing outputs it could produce is no different than a compression algorithm?
i would package the entire code as a series of comments, [ideally this would be snipped by the pliagarists] leaving a snippet of example code that no one of sound mind would allow to execute, being proffered by copilot.
That's a reach, these days...
I'm seeing some really ... interesting ... behavior, being exhibited by folks that, at first blush, I think are kids, just out of bootcamp, but, on further inspection, turn out to be middle-aged professionals.
I really think Teh Internets Tubes have been rather corrosive to collective mental health.
Smart people still exist. They just aren't online.
That said, I expect the ease of such will continue to decline as we approach a largely dead Internet, primarily consisting of bots talking to bots trying to sell each other herbal brain force supplements or whatever
Socrates: I heard, then, that at Naucratis, in Egypt, was one of the ancient gods of that country, the one whose sacred bird is called the ibis, and the name of the god himself was Theuth. He it was who invented numbers and arithmetic and geometry and astronomy, also draughts and dice, and, most important of all, letters.
Now the king of all Egypt at that time was the god Thamus, who lived in the great city of the upper region, which the Greeks call the Egyptian Thebes, and they call the god himself Ammon. To him came Theuth to show his inventions, saying that they ought to be imparted to the other Egyptians. But Thamus asked what use there was in each, and as Theuth enumerated their uses, expressed praise or blame, according as he approved or disapproved.
"The story goes that Thamus said many things to Theuth in praise or blame of the various arts, which it would take too long to repeat; but when they came to the letters, "This invention, O king," said Theuth, "will make the Egyptians wiser and will improve their memories; for it is an elixir of memory and wisdom that I have discovered." But Thamus replied, "Most ingenious Theuth, one man has the ability to beget arts, but the ability to judge of their usefulness or harmfulness to their users belongs to another; and now you, who are the father of letters, have been led by your affection to ascribe to them a power the opposite of that which they really possess.
"For this invention will produce forgetfulness in the minds of those who learn to use it, because they will not practice their memory. Their trust in writing, produced by external characters which are no part of themselves, will discourage the use of their own memory within them. You have invented an elixir not of memory, but of reminding; and you offer your pupils the appearance of wisdom, not true wisdom, for they will read many things without instruction and will therefore seem to know many things, when they are for the most part ignorant and hard to get along with, since they are not wise, but only appear wise."
They nailed us, what, four thousand years ago?
Pre-writing 'texts' (such as the Iliad) were memorized by poets, which is reflected in their forms which made more use of memory-friendly forms like rhyming, consistent meter, and close repetition.
Writing allowed greater complexity and more complex/information dense literary forms.
I feel that intelligent, critical LLM usage is just writing with less laboriousnes, which opens up the writer's ability to explore ideas more widely rather than spend their time on the technical aspects of knowledge production.
Worth noting that people were smoking plain old opium back in those times; I'd be reluctant to apply their reasoning to fentanyl.
Until the open tech community is chicken enough to not boycott their no open source stuff such as github and linked in a proof nothing will happen.
You don't get to go after GitHub because you have no contractual relationship with them. At best, you can get an injunction forcing them to take it down, though getting them to un-train copilot may not be feasible. At best you'd get a small cash offer, since you're unlikely to be able to justify any damages in a suit.
... the copyright owner may elect, at any time before final judgment is rendered, to recover, instead of actual damages and profits, an award of statutory damages for all infringements ... in a sum of not less than $750 or more than $30,000. ... in a case where the copyright owner sustains the burden of proving, and the court finds, that infringement was committed willfully, the court in its discretion may increase the award of statutory damages to a sum of not more than $150,000.
<https://www.law.cornell.edu/uscode/text/17/504>
The issue isn't contract. It's copyright infringement.
I'm curious if Github's ToS make uploading GPL software you don't own a copyright violation.
> You don't get to go after GitHub because you have no contractual relationship with them
What makes you say that? If someone eg uploads my copyrighted work to YouTube, I file a DMCA notice with YouTube to stop distributing my work. If YT ignores the notice then I can pursue them with a lawsuit.
How is this situation different?
Amazon has the rights to publish a book, and you have the right to receive a copy of the book, but neither of those gives you the right to re-publish the book under your own name.