> If it is, as you claim, permissible to train the model (and allow users to generate code based on that model) on any code whatsoever and not be bound by any licensing terms, why did you choose to only train Copilot's model on FOSS? For example, why are your Microsoft Windows and Office codebases not in your training set?
I'm not sure I buy this argument. In the first point, the authors state that Copilot was trained on public data. In the very next point, they slightly tweak it by saying training was done on "any" code which loses the distinction between public and private code. Obviously Windows and Office are not public code.
I also interpreted "public data" to mean they trained on codebases that explicitly specified, say, MIT licenses or other permissible licenses. That seems like fair use to me. Those licenses don't explicitly restrict training AI models on their codebases do they? It's ironic if these licenses started banning AI training now though. That would effectively mean Copilot would be sole trained AI model.
I'm happy to be proven wrong though. In general I have a distrust of Copilot. I fear it would make individuals worse programmers in the end at the cost of productivity