> or assure the developer, that the inserted/suggested code is not a verbatim copy of existing code
No, it does not do that.
> How can developers be sure, that they are not violating licenses by using Copilot
There are no clear answers.
That will destroy the propaganda that the Law is there to protect content creators better than anything the people against copyrights can come-up with.
""" We built a filter to help detect and suppress the rare instances where a GitHub Copilot suggestion contains code that matches public code on GitHub. You have the choice to turn that filter on or off during setup. With the filter on, GitHub Copilot checks code suggestions with its surrounding code for matches or near matches (ignoring whitespace) against public code on GitHub of about 150 characters. If there is a match, the suggestion will not be shown to you. We plan on continuing to evolve this approach and welcome feedback and comment. """
That's a bold statement considering how easy it was for testers to quickly find examples of this in initial testing.
> against public code on GitHub
... and how some of those examples found were from code not hosted on Github.
Ultimately though, what matters here is not whether this is true but whether it's plausible enough for legal departments in companies buy it.
The example I can remember was Carmack's* quick square root - but I'd probably call that "folk code" given it was passed down/altered before being misattributed to the Quake dev, and appears in hundreds of Github repos (many with permissive licenses like WTFPL, so a well-intentioned human may do the same).
> That corresponds to one recitation event every 10 user weeks
> This investigation demonstrates that GitHub Copilot can quote a body of code verbatim, yet it rarely does so, and when it does, it mostly quotes code that everybody quotes, typically at the beginning of a file, as if to break the ice
A year old post now, YMMV.
So to avoid a violation a developer needs to perform a mind-wipe?
If I am writing a novel and I copy a section verbatim from another novel, I am infringing on the other novelist's copyright, regardless of whether I wrote it from memory or not.
And this makes sense. For a trivial operation, there might be only one way to write the code. That's not copyright infringement, just like you're not infringing on an author's copyright by occasionally writing a sentence that was similar to theirs. For a nontrivial operation, you can easily write your own code without copying someone else's work.
Remember also that you can use others' ideas. Copyright only cares about the code itself. If there's a clever trick that you've seen someone use, you're free to use the same clever trick as long as 1) they didn't patent it and 2) you're not actually copying their code
If you draw Micky Mouse from memory, Disney still owns the copyright.
An example for wine/proton/reactos developers from a moderator on the forum about the leaked windows xp code:
"You look at the code? You worked for MS? No dev for us! It's that easy."
https://reactos.org/forum/viewtopic.php?t=20189
There are many instances of large lawsuits where just seeing the old code made you in eligible to even touch the new code
Maybe I'm using it wrong but I've hardly seen it pump out a mass volume of code.
Copyright violations are a genuine concern from the outputted code, GitHub themselves have admitted it may emit raw training data rarely.
> We will both continue to work on decreasing rates of recitation, as well as making its detection more precise.
That’s not a knock on Copilot, I think it’s a great product and I happily subscribed today after using it the last few months!