Since I have two free complimentary months I decided to sign up even though I'm not super thrilled with it (see previous comments). I was given two options:
1. allow code from public repositories 2. allow copilot to learn from my code
I disabled both of these options. Presumably I am now using an AI model which learns and suggests based on the context of my project.
I just gave it a test run. I have a function with this code:
if (!card.IsFaceUp && !card.IsBlocked)
{
FlipTableauCard(card);
card.SetIsBlocked(false);
break;
}
I then added this comment afterwards: // if the card is face up, flip it
And this is what copilot produced: if (card.IsFaceUp)
{
FlipTableauCard(card);
card.SetIsBlocked(false);
}
I'm pretty positive that is code generated based on my comment and the surrounding code.The "allow code from public repositories" doesn't do what you think it does. All it does is add an extra filtering step to avoid producing code found in it's training set verbatim. The model you are using was still trained on those repositories, it's not limited to your project.
But my comment still stands. You can turn off the verbatim copying feature that people keep talking about and the "AI model" will generate code based on your own codebase.
When I'm using it with Unity or a JS project that has NPM modules, does it use those as context to fill in some code as well? No clue.
Was it trained on open source code and is that ethically and legally shady? Yes.
Is it copying verbatim at this point? No.
Does it help me be a better programmer and will I pay for it? No and only if I forget to cancel my trial subscription.
US Copyright law states that fair use and derivative work are not infringing - and said law supersedes licensing.
It says no such thing:
I wrote some code. Released it under the GPL, and my only expectation is that if you use my code in your product you make source available to users (and GPL does require you tell the user how to get that source code). That on small requirement, the one thing that I'm asking you to do if you want to use my code, is not being respected by Copilot. It recommends my code, and obfuscates it, and does not tell the user where it was synthesized from, nor provides a way to get to the original source. From a certain point of view, Codepilot could be seen as a willful infringement machine. It will be interesting to see how this gets sorted out.
[0] https://twitter.com/mitsuhiko/status/1410886329924194309
Ironic that those who generally purport to champion FOSS fail to understand that Free Software was all about defeating copyright. The GPL was meant to turn copyright against itself.
For this reason, I've worked at places that forbid employees even reading open-source code. If we were having difficulty with an open-soruce component to the point where we needed to look at the code, we'd hire a contractor, explain the problem, and then they'd explain a solution, and all the communications would go through a company lawyer.
If you learn to write novels by reading other authors, is that a crime? No.
If you reproduce their work, sometimes word by word, yes.
I'm less worried about MS getting sued for this and approaching 100% expecting that users are opening themselves up to legal exposure. I can't see any legal department saying go ahead with using copilot code, but by all means ask.
it depends on what this "bunch" means. It's not clear cut at what granular level does the copyrighted parts become so small, and sources so many, that the new works is considered transformative.
I think the interesting legal question here will be, are *you* (the user of the service) infringing, or is *copilot*?
I suspect Copilot's legal team has already worked license terms such that they're passing the buck onto you.
it would seem reasonable that copilot should not be liable for anything that a user instructs it to do.
Copilot is a tool, much like the "copy" command.
If you choose to use what it's suggesting, then the fault is completely yours.
maybe my co-pilot reproduces code verbatim from your GPL'd project because you and a dozen other developers all copied the same solution from stack overflow.
Some open source developers are not allowed by their employers to read source code with a different license for fear of infringement.
If you were to write all those ideas and idioms down and pass them to someone who had not seen the Linux source code, who then used it to reimplement similar functionality, neither of you would probably be guilty of copyright infringement. (there are still patents of course). https://en.m.wikipedia.org/wiki/Clean_room_design
Copyright doesn’t protect programming idioms and concepts. It protects against verbatim copying, more or less.
It all comes down to how you characterize what Copilot does. We will just have to wait for new caselaw or even legislation that accounts for autonomous systems in defining legal wrongdoing.
This response doesn't relate to the example provided by CameronNemo, as it's a different scenario.
At any rate, there is no clean room because copilot has "seen" the literal source. It is not comparable to clean room implementation in any fashion.
the software running the model does not have access to source code
as the parent said, it depends how you characterize it, which is why this will be decided by whoever can afford the best lawyers.
Copilot has reproduced entire functions from existing codebases, which invalidates the idea it doesn't have access to source code.
But if that were the case, one could get away with copying music by merely compressing it - I am not copying the data, I have a totally different set of data that happens to get decoded into a similar performance.
That's roughly analogous to a machine learning model, isn't it? compressing an enormous dataset into a "model" that is capable of being decoded in myriad ways depending on context.
At best, if copilot told you explicitly (by doing the very hard work of identifying the likely sources of the code output ) you could make some (more) informed decisions as to if it's worth the risk to include it.
That's not a fundamental statement about all machine learning systems. GPT-2 did a lot more direct regurgitation than GPT-3. GPT-3 tends to be much more transformative, but does still sometimes spit out code / text verbatim.
Copilot and codex spit out close enough to my own code that it's clearly creating a derivative work, at least by my read.
This is untrodden legal ground, but I think that a lot of this comes down to issues of reasonableness. The reason I used the AGPL license was to create a commons. If copilot played within some reasonably friendly way around that commons, I might not feel bad about it.
However:
1) Copilot wants me to pay to use something derived from my own code, where I stuck a license there designed precisely NOT to be in that position.
2) Copilot provides a competitive advantage to proprietary projects who are more likely to be able to afford it, over open-source / community ones. The reason I used an AGPL license was because I thought we needed this type of code to be open and transparent. I work in a domain where transparency is essential (I don't disclose domain, but you can think of transparency in government, education, voting, medical, police, etc.)
3) I have no way to have a conversation with anyone at github / Microsoft. They took my stuff, and they won't talk to me about how they use it. It's automated systems all the way down.
4) The whole Open AI nonprofit -> for-profit transformation is just sleazeball. Given all the talk about ethical use of AI, something like this really leaves a sour taste in my mouth. I don't mind DeepMind, FAIR, etc., since they're honest about their goals. Open AI feels like a Silicon Valley get-rich-quick scheme with a lot of nice marketing copy and legally-questionable tactics.
Jury, judges, and developers are swayed by common sense. People like me can be swayed to testify one way or another based on whether we feel cheated. What Microsoft / github / Open AI did here wasn't very reasonable, friendly, or sensible.
TL;DR: I support the concept of co-pilot in essence. The specifics here feel illegal and sleazy.
I was about to admonish you for phrasing it this way when we all uploaded code willingly to github, giving up certain rights according to the ToS, but then I remembered microsoft straight up bought all of github, so "took my stuff" is pretty accurate. I would be interested to see a diff of the ToS since the purchase.
A lot of code on github (albeit not mine) is uploaded without the original party's agreement. Richard Stallman doesn't use github, but a lot of his GPL-licensed code has been incorporated into projects hosted there. If the terms-of-service allowed github to violate GPL licenses, I think most projects would need to migrate to gitlab. It'd be neigh-impossible for project authors to know that no GPL code in their project came from someone who did not have a side-license to github.
Even if that argument fell apart somehow, their terms-of-service state (https://docs.github.com/en/site-policy/github-terms/github-t...):
This license does not grant GitHub the right to sell Your Content. It also
does not grant GitHub the right to otherwise distribute or use Your Content
outside of our provision of the Service, except that as part of the right to
archive Your Content, GitHub may permit our partners to store and archive Your
Content in public repositories in connection with the GitHub Arctic Code Vault
and GitHub Archive Program.
github is now selling My Content. To add insult to injury, they're trying to sell it back to me!/joke
Has anyone tried taking source from leaked copies of old MS code and tried to get copilot to reproduce it?