Maybe training on copyrighted data should be allowed if the size of the training set is huge, as each individual example is justa drop in the ocean compared to the full training set.
If you train a model 20B parameters on 20T tokens, even with 1000 tokens per example, the model extracts about 1 byte of information per example. What is the value of 1 byte of copyright infringement?