Should they pay too?
If this was enforced would there be search engines?
I think it's 0. Maybe a handful if pirate libraries of readable ebooks and papers didn't already exist, but even a handful of copyright violators wouldn't be a serious commercial threat to the publishers.
So far, governments have been unable to do anything about pirate libraries. Complaining about AI datasets, poorly formatted for human consumption, seems misplaced.
Criticizing this as "stealing" or "piracy" is a vote for a future where only big tech, or the major publishers themselves, can train models on good datasets. Nobody else will have the money or market power to license that much material.
This doesn't even affect the major AI companies that are doing training, does it? Don't you think Google, OpenAI, Anthropic all have their own datasets by now? If they wanted to use bulk content from pirate ebook and academic paper libraries, they would've already mirrored most of those sites a long time ago.
What I can grasp is "output of copyrighted data is violation of copyright", and the fix seems to be a straight forward dumb software contentID filter on the output.
One of the major dividing lines between legal and illegal is whether it's a purely mechanical process, or done by human hands. The training of LLMs and diffusion models is a purely mechanical process, more like photocopying or photographing than an artist gaining inspiration"
That's part of where I think a lot of confusion is - LLM training doesn't do any of that. There is no coherent corpus of data in these models.
The argument is that the 400k books being shared here are themselves copyrighted works that the poster probably does not have the right to be distributing, and that is not very nice of them.
"It's just a dataset, bro!" being used to bulldoze people's rights to their own work might have negative consequences.
I am sure there are wild times ahead. At some time it will become outlawed in the way regular piracy is. Not because of book authors or publishers, but either because of Hollywood or AI safety being used by bigtech to limit competition.
What's the rate of forgetting for the ingested material for human brain?