Does free software benefit from ML models being derived works of training data?
mjg59.dreamwidth.org
mjg59.dreamwidth.org
As far as I understand it, Stallman wants users to have the right to hack on software that they use and to be protected from bad software because of that ability to hack, like right to repair protects consumers on the hardware side.
Well, if GitHub Copilot can wash away copyright, companies will quickly realize that opening their source just allows their competitors to wipe away whatever edge they get from their software by keeping changes secret. They will use that reason to close more of their software to preserve their trade secrets.
In other words, if copyright wouldn't apply to software anymore, then companies would protect their software by trade secrets instead, leading to less open source, which would lead to less user freedom to hack, especially since such trade secret controls might be used against users themselves.
So I would argue that if Stallman wants copyright to disappear, he wants the wrong thing because it conflicts with his goal of user freedom, and Copilot is a problem. If he doesn't, then he will absolutely care about Copilot because Copilot will still be a problem since licenses use copyright to protect users.
I can't stress that amount. Our copyright on our software protects our users, and as such, we should guard it jealously. Not for our sake, but for theirs.
Edit: I would like to add something about licenses and contract law.
While FOSS licenses can also be contracts (https://writing.kemitchell.com/2019/03/09/Deprecation-Notice...), if copyright does not apply, then users (and companies) do not have to accept the license in order to use the software because they already have access to it since the source is open.
However, with closed source software, the user does not get access until after they accept the EULA. This means that full contract law applies to the user, regardless of copyright on software.
That is why closed source software wins if copyright does not apply: because it makes an end run around copyright anyway.
Just checked: the Windows EULA [0] does not allow you to copy the software. And since it is a contract, that clause does not need copyright to be in force. So the person copying it to you can be targeted by Microsoft, and if you copied Windows without that user's permission, you are probably going to be charged with a violation of the CFAA. [1]
These companies have been practicing law for longer than you. Don't underestimate them.
Also, in my opinion, you are not really arguing in good faith. That is probably because my arguments really tear your article apart, and you are trying to find some way to discredit what I am saying. You can try, but I guarantee you that companies have a lot of experience protecting their closed source software. If copyright does not apply anymore, I also guarantee you that they will find, or probably already have found, a way to protect their software without it.
Edit: The Windows agreement specifically says that using the software is an implicit agreement to the EULA, and Adobe's [2] says the same thing. Thus, even if you copy from someone else's computer, you are under the EULA.
[0]: https://www.microsoft.com/en-us/Useterms/OEM/Windows/10/UseT...
[1]: https://en.wikipedia.org/wiki/Computer_Fraud_and_Abuse_Act
[2]: https://www.adobe.com/products/eula/tools/captivate.html
If I'm given permission to use a computer, and I then use that access to duplicate binaries that are not subject to copyright law and which the level of access I have been greanted gives me access to, which provision of the CFAA do you think I'm violating?
> Edit: The Windows agreement specifically says that using the software is an implicit agreement to the EULA, and Adobe's [2] says the same thing. Thus, even if you copy from someone else's computer, you are under the EULA.
That's something you can enforce if use of the software is controlled by law, and you can refuse me permission to use it if I don't agree to the EULA. If the code isn't copyrightable then you have no ability to do that - I'm entitled to use it without agreeing to the EULA, and asserting that I'm implicitly agreeing to it carries as much weight as me asserting that you agree to pay me $1000 by reading this comment.
The CFAA covers use of a computer that extends beyond the authorized use.
> That's something you can enforce if use of the software is controlled by law, and you can refuse me permission to use it if I don't agree to the EULA. If the code isn't copyrightable then you have no ability to do that - I'm entitled to use it without agreeing to the EULA, and asserting that I'm implicitly agreeing to it carries as much weight as me asserting that you agree to pay me $1000 by reading this comment.
Good luck when a company comes after you for not agreeing to the EULA.
You and I might agree that you are entitled to use it despite the EULA, but that is not what we are discussing; we are discussing how companies will protect their software if copyright does not apply.
Saying that me reading your comment implies that I agree to pay you $1000 is wrong because of two major things:
1. You posted in a public forum.
2. I did not need to copy your comment to read it.
Windows has not posted their source code in a public forum, and the use of their software on a new machine requires copying it.
Actually, this talk of $1000 gave me an idea, so let's make a bet: if you can train a Copilot-like model with a decompiled copy of Windows and not be sued by Microsoft, I'll give you the $1000 you claim I owe you. But if they sue you, then obviously, they think there is something protecting their code, whether EULA or copyright.
Good luck with that.
Edit: Also, implied contracts are a thing. [1]
[1]: https://www.investopedia.com/terms/i/implied_contract.asp
Right. I have valid credentials for a system which grant me read access to a bunch of binaries. Are you saying that I need explicit authorisation to copy those?
> 2. I did not need to copy your comment to read it.
Agreed! I'm not going to be able to cause you to implicitly accept a EULA by doing something you already have permission to do, just as Microsoft wouldn't be able to cause me to implicitly accept a EULA by copying a binary that isn't subject to copyright law.
> But if they sue you, then obviously, they think there is something protecting their code, whether EULA or copyright.
You seem to be conflating two different situations. At present, since Windows is considered to be under copyright, Microsoft are in a position to enforce a EULA that would forbid me from doing that, even if the model I generated would produce output that wasn't considered to be a derivative work of Windows. If software couldn't be copyrighted at all then their ability to enforce that would be significantly weakened.
Yeah, you do.
> Agreed! I'm not going to be able to cause you to implicitly accept a EULA by doing something you already have permission to do, just as Microsoft wouldn't be able to cause me to implicitly accept a EULA by copying a binary that isn't subject to copyright law.
Then take the bet. I'm not even asking you to give me $1000 if you lose.
> You seem to be conflating two different situations. At present, since Windows is considered to be under copyright, Microsoft are in a position to enforce a EULA that would forbid me from doing that, even if the model I generated would produce output that wasn't considered to be a derivative work of Windows. If software couldn't be copyrighted at all then their ability to enforce that would be significantly weakened.
Ah, but you forgot one thing: Microsoft and GitHub are claiming that training a machine learning model is under fair use, which is an exemption to copyright. So if they sue you for training a model with their code, then they either do not actually believe that training an ML model is fair use, or they think that they can enforce the EULA without copyright.
Either way, take the bet. Prove me wrong.
Every time you access a file share, you request explicit authorisation to copy any of the files in that share? I think you'd have trouble finding any court who accepted that this was the intended outcome.
> Then take the bet. I'm not even asking you to give me $1000 if you lose.
I'd gladly take that bet in a world where software wasn't subject to copyright. Since this isn't that world, Microsoft's in a position to compel me to agree to the EULA if I want to get at the binaries.
> Microsoft and GitHub are claiming that training a machine learning model is under fair use, which is an exemption to copyright.
A EULA can require you to give up rights that you would otherwise have - in many cases it's legal to engage in reverse engineering, but in order to gain the right to use the software you may have agreed not to do that. So even if it would be fair use to analyse the binaries, if I can't obtain those binaries without either violating copyright law or agreeing to a EULA, I'm not in a position to do so without putting myself at risk. If copyright law didn't apply to the binaries in the first place then I wouldn't need to agree to the EULA in order to obtain them, and there'd be no problem.
If you don't own the files, yes, you do need authorization. I think courts would not have trouble with that concept.
> A EULA can require you to give up rights that you would otherwise have - in many cases it's legal to engage in reverse engineering, but in order to gain the right to use the software you may have agreed not to do that. So even if it would be fair use to analyse the binaries, if I can't obtain those binaries without either violating copyright law or agreeing to a EULA, I'm not in a position to do so without putting myself at risk. If copyright law didn't apply to the binaries in the first place then I wouldn't need to agree to the EULA in order to obtain them, and there'd be no problem.
You are dodging and deflecting.
As you said yourself, if copyright doesn't apply, you can obtain copies without worrying about the EULA or copyright.
Well, fair use is an exemption to copyright. If you do something under fair use, it does not matter if you agree to a EULA or not; if your use is under fair use, you can do it. Period.
So I gave you a way around copyright by using GitHub's own argument that it is fair use. If you won't take it, that means you don't believe what you are saying because you are not willing to put your actions where your mouth is. Yet I am willing to front $1000 if you do.
So clearly, you are not arguing in good faith, especially since your argument is about the consequences if copyright is done away with. In this one case, where GitHub claims that copyright is done away with, you won't take action, which means all that you are saying is just empty air.
Because you are not arguing in good faith, I think I am done with this conversation. And I will take the bet; I will challenge Microsoft to issue their legal opinion on this.
No. A contract can involve me agreeing not to do something I would otherwise be permitted to do (eg, an NDA is a contract in which I agree that I won't engage in certain forms of constitutionally protected speech). If I agree to a EULA that says I won't do things that would be permitted by fair use, that's probably enforceable.
> In this one case, where GitHub claims that copyright is done away with
They don't claim that! They claim that what they're doing is permitted by copyright law, which is a very different thing.
> If you don't own the files, yes, you do need authorization. I think courts would not have trouble with that concept
And how many people have been successfully prosecuted (and that prosecution upheld) under this interpretation?
You do not get it. I am assuming you did as you originally mentioned and got a hold of a copy of Windows without agreeing to the EULA. If that's the case, then you don't need to worry about the EULA, according to your argument at the beginning of this discussion. Then you can train the model under fair use.
Unless, of course, you actually don't believe your own arguments where you claimed that you can get a copy of Windows without agreeing to the EULA, in which case, your entire first argument is moot, and you admit that EULA's do put implicit contracts on people that copy the code, even without copyright, which means that I am right.
Once again, you are dodging. You keep trying to come up with "gotchas," but you're really digging your own grave here.
> They don't claim that! They claim that what they're doing is permitted by copyright law, which is a very different thing.
They are claiming an exemption to copyright law that is spelled out in the copyright law. Thus, they are claiming the same thing as no copyright in this particular case.
> And how many people have been successfully prosecuted (and that prosecution upheld) under this interpretation?
Ever heard of Aaron Swartz? He was charged under the CFAA and with grand larceny, among other things.
He was charged with unauthorised access to a network (he wasn't affiliated with MIT), and with violating JSTOR's explicit terms of use regarding what types of access were authorised. This is not analogous to someone having authorised access to a computer without any other terms of service. And, uh, it's a case that's pretty notable for not having been successfully prosecuted.
If you want to make shit up, go ahead, but this clearly isn't a productive discussion.
And it was only not prosecuted because he committed suicide. If you're going to reject true facts, then this discussion isn't productive.
(Arguing IP on the Internet? what is this, the 90's)
And what follows doesn't need to be completely unrestricted. You might ban for example all clauses that restrict copying this piece of software in any way or using it as is.
> ...nothing other than this License grants you permission to propagate or modify any covered work. These actions infringe copyright if you do not accept this License. Therefore, by modifying or propagating a covered work, you indicate your acceptance of this License to do so.
From too much freedom I presume? Or what did you mean exactly?
> if copyright wouldn't apply to software anymore, then companies would protect their software by trade secrets instead
As if they didn't do that already? What part of Photoshop is not secret? Or ar least wasn't at some point until people reverse engineered it? Copyright is the last resort. The first and most solid line of defense for commercial software is alwas secrecy, of source code, of file formats, of binaries in case of server side software, of anything they can keep secret.
> leading to less open source
Why would that lead to open source since companies keeping their software secret, closed source and copyright protected was exactly what lead to open source in the first place?
From things like what Audacity did: adding spyware or other bloat that serves the company rather than the user.
Also from being taken advantage of, as right to repair has shown. If we have right to repair, it's harder for companies to engage in monopolistic practices. Open source does the same thing for software.
> Why would that lead to less open source since companies keeping their software secret, closed source and copyright protected was exactly what lead to open source in the first place?
Because right now, if a company uses open source code in their code, then they will sometimes contribute back to the original software. Without copyright, they won't because their change might reveal something about their file format or something like that.
It's because of copyright that companies have played somewhat nice when using FOSS. Without it, there is no incentive to.
So we either 1) declare this to not be copyright infringement, or 2) basically outright outlaw training of interesting ML models.
I personally believe that (1) is the way to go, and I find the whole outrage about Copilot to be essentially akin to collectively shooting ourselves in the foot in the long run.
I think we shouldn't forget why copyright exists in the first place. To quote the copyright clause of the US constitution (emphasis mine):
> "*To promote the Progress of Science and useful Arts*, by securing for limited Times to Authors and Inventors the exclusive Right to their respective Writings and Discoveries."
If anything we already went too far with how draconian modern copyright enforcement is, up to the point that it often actively hampers our progress. Please don't make it even worse!
Would you dramatically expand what you can legally borrow without infringing or just allow wholesale automated washing from one copyright to another?
What do you do when random corp starts shutting down open source over ownership of code it swiped from project claiming infringement?
In other words, to what extent does the color of the bits matter? [1] And does the output of a ML model have a new "color"?
You're acting like we don't already have something similar. We do. It's called fair use. (: I'd probably just expand on that. If the portion of the work generated by a model is insubstantial, and/or doesn't harm the interests of the original rights holder then it should be okay.
If your model can generate a single excerpt of Harry Potter does that constitute harm to J.K. Rowling? Can you use it to reasonably generate a substantial portion of the book with it? Is it a reasonable substitute for buying the book? No? Then it's okay.
Some countries in fact do already have such exceptions, e.g. Japan amended its copyright law in 2018 to explicitly allow for this. To quote the relevant law:
> "It is permissible to exploit work in any way and to the extent considered necessary ...where such exploitation is not for enjoying or causing another person to enjoy the ideas or emotions expressed in such work..."
So if you want to train an ML model on a bunch of commercial novels, without the permission of their owners, and then exploit that commercially - you can do it, and it's completely legal.
Personally, I'd like not have my IP used in training ML models. If they want to incorporate my stuff into their training data, they can do it on my terms, which would be either paying to license my work, or something along the lines of a new kind of copyleft license where the model has to be public and copyleft too. By including my work anyways, I didn't get the terms I would have needed.
Scale makes it weird, because it's the a-drop-of-water-in-an-ocean issue, but harm is still real harm when it's split into a billion separate pieces of infinitesimal harm.
Saying that it’s infringement doesn’t mean we’re outlawing interesting ML. These definitely aren’t the only two options. A few years ago, this kind of infringement wasn’t even possible, and now it’s just beginning. This stuff is in its infancy.
Reliability and controllability is an issue our field has to confront as applications move from research to application. This is just another instance of that. We’re all eager to deploy this stuff as soon as it can do something cool, but that’s not enough, as all the discussion around copilot demonstrates.
> I think we shouldn't forget why copyright exists in the first place… To promote the Progress of Science and useful Arts
Right. It’s to encourage creators to create things. I wrote about this elsewhere [0] that the issue of generative ML models is that, as far as incentives go, you can pay a creator for use of their things, or you can throw money at AWS and get use of something like that creator’s things. The fair use equation is ‘$$$ + my stuff = your stuff,’ except that that money doesn’t go to the creators, it goes to nvidia and aws.
I think we’ll need new copyright laws for this, to make a middle ground. I personally want to be able to exclude my works from being training data unless I license then openly enough. Just like with open code licenses where you opt in. That option doesn’t seem to exist right now.
Wouldn’t it be great if a model benefiting from all this public data had to be public itself, or something along those lines? Wouldn’t that promote the arts and sciences even more?
- We could extend copyright statute to explicitly allow training ML models only under particular conditions (for example, some kind of "non-commercial use" clause).
- We could enact mechanical licensing for inputs to ML models, as exists for musical cover performances. Creators of the works used as inputs would get royalties.
- We could require asking for permission to use a work as input, but standardized generous terms once permission is granted.
- We could flip that and just require notice that you're about to use a work as input, but allow the creator a certain window to object.
- We could declare that inputs are a free-for-all, automatically non-infringing, but that you (the creator of the model) are liable for extra damages if the outputs are proven to be able to infringe.
Many, many different ways this could shake out. I'm not particularly optimistic that any of (what I see as) the better choices will come to pass, but they exist.
https://github.blog/2015-03-09-open-source-license-usage-on-...
CHAPTER ONE
THE BOY WHO LIVED
and it’ll do it super duper reliably with CHAPTER ONE
THE BOY WHO LIVED
Mr.I have been effectively using GPT-3 a lot recently because I added chapters with OpenAI APIs examples to my Common Lisp and Clojure books (you can get free copies at [1] by setting the price to Free).
At Capital One, I used LSTM models to generate AWS JSON logs for testing purposes. The original logs contained sensitive data that was not in the generated data. JSON structure and names of keys was learned and reproduced verbatim, but not the values. GPT-3 has massively more parameters than my models. BTW, I also generated spreadsheet data using GANs: similar idea, but different tech.
If you have a piece of code that would be infringing if you wrote it in Emacs. How can it not be infringing if it was written by VSCode and Copilot? I just don't see how any court would hand down such a judgment.
What I do think that Github could get away with, is not being liable for copyright infringement in my code if I used Copilot and it gave me some code that infringed. But that would just move that liability to me. Good for Github, not so good for me.
If running the data through a ML model, removes the copyright. Then we could always train models with specific input to remove copyright on that data, and we follow through on that. We could easily remove copyright on anything, and that would, if the courts upheld that. Be the death kneel for copyright. Can't really see that happening. But maybe that is just my limited imagination.
Maybe I can't read, but it's just not there. This ruling giving precedent to train generative ML systems with any source material seems to be nothing more than a shared fiction entertained by the ML industry.
This.
Ianal but I feel pretty sure that the content of the work is what is in question, not how it was created with regards to copyright. I think wikipedia more or less says this:
https://en.wikipedia.org/wiki/Copyright_infringement#Limitat...
- You can keep things secret without copyright being involved at all. Something like the Coca-Cola recipe isn't protected by copyright, but by secrecy. No problems to apply this to proprietary software.
- SaaS and PaaS make distribution terms irrelevant and secrecy very easy.
Open source for the user lost a number of years ago, the reason companies embrace it is because it commoditizes components, i.e. lowers cost, because it's literally free stuff.
IANAL, but it can be more fine grained than that. For example, courts may very well care if you obtained a copy of the source code legally before you trained the AI on it, or whether the source code copy comes from a hack or a leak.
There is a difference in a court declaring that a rental agreement's clause that no parties may be hosted by the renter is void, and someone breaking into someone elses home and hosting a party there.
Proprietary source code that you can obtain legally does exist, but it's a much smaller category than proprietary code in total.
In the worst case, this might render copyleft unenforceable while doing nothing to most proprietary code.
It's something like:
Having been unable to defeat free/open software through the courts and copyright, lets try it this way -- let's see if we can, without too many people noticing, gain control of free/open source through its biggest repository with "Copilot." We'll let just enough people who are not us use it in ways that we see fit, but also probably later on will fight against "scraping our data" to retain control (like we tried to do with LinkedIn.)
Nevertheless, there was a link a couple weeks ago for a NN that replicated GTA. I wonder if that is more along the lines of what the author wanted to argue? I.e. if I can have a NN learn and replicate windows, would MS be ok with it?
> there will not be derivative works of closed source at least from a copilot use case point of view.
which is inaccurate, since many public repositories on Github aren't open source and ended up in Copilot's training corpus anyway.
Because I'm pretty sure it doesn't. Section D4:
> This license does not grant GitHub the right to sell Your Content. It also does not grant GitHub the right to otherwise distribute or use Your Content outside of our provision of the Service...
This isn't really true when software is distributed in binary form - or, increasingly, run on servers and not distributed at all. The outcome Stallman wanted is that end users would have the ability to fix bugs, and add new features, in software they were using. If software is uncopyrightable then that removes one barrier, but not necessarily the most severe one; if you're relying on cloud software whose actual implementation kept as a trade secret (for example) then you still can't fix the problems that are disrupting your workflow.
> If Github's interpretation of copyright law holds, we can train a model on proprietary code and extract concepts without having to worry about being tainted. The proprietary code itself won't enter the commons, but the ideas it embodies will. No more worries about whether you're literally copying the code that implements an algorithm you want to duplicate - simply start typing and let the model remove the risk for you.
I think the closest analogy is to source code leaks like the famous Windows 2000 one. Would having such leaks be non-copyrightable help free software? I doubt it - even a deliberately released "code dump" is still difficult to learn much from, and even free-software projects have to put a lot of time and effort into documenting their code structure, having a contributor onboarding process and so on. The value of the GPL isn't that it forces people to make their raw code available - "hostile" (L)GPL code, like early WebKit or xlvns, has never made for successful projects - the value is that it creates an environment where people are actually cooperating on software development in a positive way. I think that's more likely to happen in a world where you can only use Copilot-like tools if your software is free than in a world where no such rule exists.