> It's pretty obvious that without the open source code corpus, such tool would not have been possible, hence the (IMHO) justified derivative work question.
Note though that a work not being possible without your work is not sufficient to make that work a derivative work of your work. It just suggests that you need to take a closer look at the relationship between your work and the other work.
For example Windows applications, even ones that make intimate use of the behavior of Windows and would take significant rewrites to port elsewhere or to run under current Windows compatible operating systems like ReactOS or under things like Wine, are not automatically derivative works of Windows.
To be a derivative work the work has to include copyrighted elements from your work in a way that is not covered by fair use. That's why clean room reverse engineering works--by making sure the coders do not have access to the work being reverse engineered they cannot copy any copyrighted elements from it and so cannot produce a derivative work.
I suspect that under current copyright law it is possible to do something like Copilot without the output violating copyright but it may need to be more sophisticated than the current Copilot.
From the few examples of Copilot output I've seen it seems to output stuff that would probably either be covered by fair use or that doesn't have enough creativity to be copyrighted. But from what people have said it occasionally spits out longer things that seem likely to be copyrighted and not covered by fair use.
What may be necessary for systems like this is to couple them with a second AI that can recognize when the first AI is making a suggestion that goes beyond fair use and stops it. I don't know if it is currently possible to make such an AI. Where would you get a good set of training data?
The above was about the output of Copilot. Another question is whether Copilot itself is legal. When you train an AI on some data is there a copy of that data in the AI? If there is then Copilot may be an infringement of the copying right.
In the US copyright law defines copies in 17 USC 101, where they are defined as
> [...] material objects, other than phonorecords, in which a work is fixed by any method now known or later developed, and from which the work can be perceived, reproduced, or otherwise communicated, either directly or with the aid of a machine or device. The term “copies” includes the material object, other than a phonorecord, in which the work is first fixed.
Is a collection of neural net weights something from which you can perceive, reproduce, or otherwise communicate the individual works the net was trained on? Or is it more like some kind of hash of the work?
My guess is that both the output of AIs and the AIs themselves are sufficiently beyond what anyone was contemplating the last time there was a major update of copyright to deal with new technology that to fit AI in we are probably going to need a major update to the law.