We already have the same fuzzy line for writing. Am I forbidden from ever reading other author's books because I might accidentally "generate exact copies" of some of the sentences? Clearly not, that is how people learn a language. Does that mean I am allowed to copy the whole book? Also clearly not.
Where do you draw the line? Somewhere.
Forbidding the training of AI on public code would definitely be a step too far though.
Edit: I'd also like if they provide a tool for checking if your code matches copyrighted code too close so you can confirm if you are violating anything or not when you use copilot.
There's absolutely nothing different whether the creator is ML or a human.
Generally, if you train an ML network to generate an almost exact copy of a thousand lines, it's obviously not fair use. If it's five simple lines, it obviously is fair use. If it's somewhere in between, there are a lot of different factors that need to be weighed in a fair use decision, which you can easily look up.
My simplistic view is that the following is legally equivalent:
input -> ai network -> output
input -> huffman coding -> output
So, whilst:
* compressing and decompressing a copyright work is permissible;
* output and weights are deterministic transformations of the inputs;
thus:
* not eligible for copyright (lacking creativity); and
* are derivative works of the inputs;
copyrighted input -> compiler -> copyrighted output
So portions of a derivative work are covered by the original copyright, and other portions may be under a distinct copyright as a derivative work, and several copyrights may apply to a work as a whole.
In the case of a Huffman transform, the transformed work does not meet the "creativity" requirements to be eligible for copyright, over that of the original works.
That may be true but I fail to see how any process that produces the same content that was input into it somehow strips the license. If the generated code is novel, then there is no copyright and it is just the output of the tool. If the code is a copy, but non-creative (example a trivial function) then it isn't covered by copyright in the source anyways, so the output is not protected by copyright either. However if the output is a copy and creative I don't think it matters how complicated your copying process was. What matters is that the code was copied and you need to obey copyright.
Again, I don't think that novel code generated from being trained on copyrighted code is the problem. I think it is just the verbatim (or minimally transformed) copying that is the issue.