Also I work in Rust these days, and use CLion. Not sure if it would be much help in that environment.
Since Microsoft uses material for copilot outside of licensing on the basis that it is Fair Use, that would probably have no effect in practice on whether or not the material is used in training something like Copilot. For that to matter, you’d first have to win a lawsuit on the basis that training something like Copilot requires permission of the copyright owner of the training material, to invalidate the premise of Microsoft’s action.
i think that any formal violation of an open source license would be very bad in terms of public relations - even if they can claim fair use.
I think that's the part that people who don't think it's worth the money aren't getting. This kind of system is godsend for the likes of Infosys, TCS etc. So the immediate threat is to the jobs there - but the side effect is that it'll make it all even cheaper, so we'll see more "outsourcing to the cloud", so to speak. Often to the obvious detriment of quality, but that doesn't seem to matter in this market.
Anything generated by copilot, which if it is a derivative work, is not something that copilot can hold the copyright on.
From the auditor's perspective it doesn't matter if you copied it out of stack overflow, from some GitHub search, or copilot. You, the human, didn't check the license / plagiarism detecter. It is you, the human, claiming copyright on the work you are creating which may incorporate material from other sources.
Copilot isn't claiming fair use.
You could argue that the model that copilot runs from is a derivative work (and this is going to be interesting when it gets to the courts, because, frankly no one will come out the 'winner' on this when trying to explain it to a judge) - but that's not the code that a human is claiming to be their creative work and is ultimately the license violation.
Personally, I (not a lawyer), believe that copilot is on ok ground - but anyone using it needs to do their due diligence in verifying that the code that they've incorporated is licensed appropriately - just as if they've copied something from Stack Overflow - who knows where that copied was copied from.
I have less concerns with identifiable code from copilot than humans not caring about the licenses of their source material in creating human generated content.
The moment a fragment of code from a GPL'd or AGPL'd project shows up almost verbatim in someone's closed source or non-copylefted, etc. project, and someone proves it, sparks are going to fly.
And it's probably already happening, just people haven't discovered it yet.
How many years did the Oracle/Google lawsuit go on for? And in the end about a handful of a lines of code only tangentially related to the issue at hand?
Part of the lesson from that must be: employers should be telling their workers to say TF away from Copilot or things like it. And be careful in general when browsing source. License literacy is critical.
I don't touch it because I need to feed my kids. I don't need my career exploded. Overcautious? Maybe. I'll let someone else find out. I make a living in and around open source software.
I'm hoarding popcorn and can't wait for that, honestly.
> License literacy is critical.
It's beyond critical, but most people I have talked says that they see the right to copy and use any code they see online. They don't care.
Building this open source corpus was not easy, and we need to defend it too. This is a culture.
ISTR that Amazon's version of copilot has licensing info output too.
This fact solves many "we need this lib but it's too open" problems. The author can make a special license just for you if they want.
It is the nature of the beast, however I make my stance clear and stand by my principles. It is equally important for me.
It's not even clear how often training machine learning algorithms on code results in copyright violations. CoPilot does have a setting to detect and disallow direct copying, but how well does it work?
This legal uncertainty is enough that I wouldn't advise using it, but maybe people who use it will be fine?
Europeans will probably preempt that with legislation.
The phrase “derived work” is, IIUC, a phrase from copyright law. And you’d have a hard time convincing me that Copilot-generated code is not a derived work from its training data.
> And given that they’re open source code bases I expect licenses to explicitly disallow things, and consider anything not explicitly disallowed as permitted.
That is very much not how copyright and licences work. Copyright law gives the copyright holder the exclusive right to make copies of the work, making derived works, (and to do some other related things, like making a public performance of it, etc.), so to do any of those things, you need explicit permission, i.e. a license from the copyright holder to do it. A license is not a list of things you are forbidden to do; on the contrary, it is a list of things you are permitted to do, which you would not otherwise be legally allowed to do according to copyright law.
This is a novel scenario. It seems unclear how the courts will interpret it? Never mind what we think, will they decide it's a derivative work, or is it a transformative use?
Suppose we create a new AI image generator, and use as training input every image ever made of a Disney character (official images by Disney, that is, no fan art), including every frame of every Disney movie. Could we just use the output images of that AI however we wanted to? (Not withstanding trademarks.)
https://en.m.wikipedia.org/wiki/Copyright_protection_for_fic...
I just don't understand the OSS community sometimes. "Software should be open and free (libre) for me to study and modify" includes what Github did for copilot. If you don't want your software to be free (in either sense), don't host it on an open source platform, especially one that makes it available gratis to the public.
There's possibly a valid argument that any private repo code that was used for copilot doesn't fit the proper definition of "open" (or gratis). But I haven't actually read the Github license around this, so I don't know.
Copyleft, free software, GPL style licenses do not have their source open purely for the purpose of studying and modifying. Their licenses also require that derivative works also be free and that such modifications be distributed.
Copilot does not comply with this. And so violates the spirit of those licenses, and probably also the letter of the law.
And it’s at best questionable whether copilot is derivative of the code it is trained on.
And by that logic, gpt3 and dall-e are also copyright violations. Except this has never been proven to be the case.
The completions are often more “similar to stuff I’ve typed before” rather than “generate this working function”, and it’s not an order of magnitude improvement over just regular intellij completion.
…but I find it’s generally better.
Up to you if you consider it worth the cost of paying for it, and the macro completion is so-so, but “its rust so copilot can’t do it” isn’t really true.
I only used it for a language I'm very familiar with. I'd be a lot more hesitant using it for a language I'm less familiar with because I won't be able to spot the problems so easily.
I wrote the grammar in a comment at the top of the file, wrote and imported the AST enum, and wrote the first production. After that, I just prompted Copilot and it worked its way down the grammar, producing the parser functions one at a time. The CLion integration was able to consider the imported data structures as part of the prompt, so it even stored everything in the correct AST nodes.
For something like that, it's easy to verify that it did it correctly (through visual inspection and testing), and it allowed me to write the entire parser in about 2-3 seconds per rule.
Here's an example I just now tried in the codebase I had open. I literally typed in the following Go in a function:
mean := geom.Vec3D{}
for
Copilot then autocompleted this, me pressing tab to accept each line, to: mean := geom.Vec3D{}
for _, detection := range filtered {
mean.X += detection.X
mean.Y += detection.Y
mean.Z += detection.Z
}
mean.X /= float64(len(filtered))
mean.Y /= float64(len(filtered))
mean.Z /= float64(len(filtered))
It's not a perfect platform by any means, but it's also pretty useful for this sort of autocompletion.I'd rather see Copilot used to expand existing libraries. Pulling potential additions to your own library off of Copilot would be an interesting twist on the situation. A hacktoberfest-alike based off this would be weird.
I think you should give it a try and see what you think about it. I was hesitant about it at first but was very surprised at how much time it could save me from having to look up things on Google. I've found it especially useful when I'm switching to a language I may be less familiar with. I can understand the basic logic of what I want to do, but would have to spend time looking up how to do it in this specific language. Orrr... I just have copilot help me out and generate such a solution.
It's obviously not going to be a tool that is applicable to everyone. Just like how many manual labour oriented contractors have a bunch of tools, each of them may have their own set of tools that slightly differs from the other person. That is okay, and there would be no need to try and bring someone else down for their choice to use a certain tool.
My concern is that I'd get locked into a specific IDE.
Besides, the generated example seems to be missing code to gracefully handle the case where len(filtered) is zero. Maybe there's a precondition that prevents that from happening or maybe a division by zero is exactly what you'd want, but at face value it looks like the bot did a rush job.
Saturation is not. This is what really bugs me: If I'm going to drag in a billion GPUs of external computation (or a dependency, which is basically the same thing but with human brains), I want it to provide the hard algorithm I can't write, not the easy one I can. I am not limited by typing speed.
I also wonder what a "detection" is.
(Also, in--say--Ruby and JavaScript 1.0/0.0 is Infinity and not NaN.)
Both good examples of coding where you should be thinking instead, though.
For what it's worth, Copilot (correctly) inferred a loop variable called "detection", I imagine based on similar usage earlier in the function. And there is already a conditional in place to prevent invalid operations; if I remove it I see a new suggestion:
if len(filtered) > 0 {
This tool is far from perfect, but it very much sounds like you folks haven't used it. If that's the case, I would encourage you to research it like all tooling and draw some informed conclusions about it's applicability instead of making assumptions.It automates a great deal of boilerplate crap that you have to do especially in web frameworks such as angular.
Until then, copilot is a giant liability that ensures I can't use it for code that my company ends up owning, nor can I contribute code I write with it to literally any open source project because in a very real sense: I didn't write it. I just assembled it from parts unknown, and those parts may end up being lawsuits.
Obviously IANAL so this is largely conjecture, but until we _actually_ see how this would play out in court, I'm leaning towards this being less of a legal issue than folks here act like. For personal projects, I'd say the likelihood of some other engineer reading your code, noticing a similarity or duplication, and dragging you to court for it is near 0.
- "you were hired to write code for us, not to use an autocomplete service that makes us liable for both copyright and patent lawsuits, I hope you like getting fired."
- "as per this project's license, we can only take code on board that you contributed under our license, but an audit shows that your PR/MRs contain tons of GPL/MIT/Whatever licensed code instead. We're going to have to back all of that out, and we're going to revoke your contributor status"
- etc.
If you don't know where the code in your autocomplete comes from (and copilot can autocomplete large swathes of code) then literally anything that comes out of "you didn't write this code" may apply. From fraud (depending on what contract you signed) to trademark infringement, to license violations, to even just simply misrepresenting your skills to an employer. As with all things, it's a sliding scale, but just because the majority of incidents will be on the bening part of the spectrum doesn't mean the litigating part doesn't exist, and that's what your legal department plans for.
Work for a big company? Good bet you're not allowed to use copilot. And depending on the company, not even "for your personal projects" because you might accidentally read someone else's license encumbered code that you would not have come up with yourself and may now open your employer up to "you stole our ideas instead of properly crediting/paying for licenses".
Copilot is a legal nightmare.