Are you saying that you believe that untested but technically; models trained on GPL sources need to distribute the resulting LLMs under GPL?
Are you saying that you believe that untested but technically; models trained on GPL sources need to distribute the resulting LLMs under GPL?
If an AI outputs copyrighted code, that is a copyright violation. And if it does and a human uses it, then you are welcome to sue the human or LLM provider for that. But you don't get to sue people for perceived "latent" thought crimes.
That being said, I don't think that your analogy is valid in this case.
> GitHub and Windows and IDEs need to be open source because they can output FOSS code
They can output FOSS code, but they themselves are not derived from FOSS code.
It can be argued that the weights of a model is derived from training data, because they contain something from the training data (hard to say what exactly: knowledge, ideas, patterns?)
It can also be argued that output is derived from weights.
If we accept both of those claims, then GPL training data -> GPL weighs -> every output is GPL
> If an AI outputs copyrighted code
Again, the issue is not what exactly does AI output, but where it comes from.
(I say 'relatively easy'. Not that it would be trivial.)