The court of law would probably have you unveil and tear apart your process and find that you were trying to plagiarize in a roundabout way.
The court of law would probably have you unveil and tear apart your process and find that you were trying to plagiarize in a roundabout way.
This is the whole idea. Copilot is spitting out considerable chunks of code that is licensed under GPL and it will be up to GitHub to prove that Copilot is not trying to plagiarize this code in a roundabout way.
In the very least, Copilot should have separate data stores for different groups of licenses: public domain, attribution-only, copyleft, etc.. That would already make it much more usable than the current "here's some code, it came from I don't know where, don't ask me" that literally looks like black market deals except they are GitHub-branded.
One regurgitation in 10 weeks per user. Not considerable. Could be just skipped with a simple search by Github.
The only reason the law doesn't work this way for Microsoft Copilot is because the copyright holders are individuals who do not have the capital or expertise to file suit.
If Microsoft instead released a video editor addon that was trained on Disney movies and which would sometimes insert scenes of _any_ Disney movie you can bet your ass we wouldn't be having the same discussion.
Well, yes, that's the point.
So using machine learning with one movie is illegal, but using it with a million isn't?
Note: I’m not really trying to comment specifically about the code/movie examples - just the general notion that the more input there is (from different sources), the likelihood that the use will be considered “Fair Use” increases.