What I can grasp is "output of copyrighted data is violation of copyright", and the fix seems to be a straight forward dumb software contentID filter on the output.
What I can grasp is "output of copyrighted data is violation of copyright", and the fix seems to be a straight forward dumb software contentID filter on the output.
One of the major dividing lines between legal and illegal is whether it's a purely mechanical process, or done by human hands. The training of LLMs and diffusion models is a purely mechanical process, more like photocopying or photographing than an artist gaining inspiration"
That's part of where I think a lot of confusion is - LLM training doesn't do any of that. There is no coherent corpus of data in these models.
The argument is that the 400k books being shared here are themselves copyrighted works that the poster probably does not have the right to be distributing, and that is not very nice of them.
"It's just a dataset, bro!" being used to bulldoze people's rights to their own work might have negative consequences.