Do you think that DeepMind, OpenAI, etc. don't have this dataset copied 10 times over? Think again!
Do you think that DeepMind, OpenAI, etc. don't have this dataset copied 10 times over? Think again!
All the current LLMs are trained on data scraped from random websites, with no regard for the website's copyright, aren't they?
Presumably because these big businesses have a theory training an LLM is 'fair use'.
Why wouldn't they treat books the same way they treat web pages?
That being said, it feels like there's also a shade of perspective from the old quote:
"In its majestic equality, the law forbids rich and poor alike to sleep under bridges, beg in the streets and steal loaves of bread." - assuming everything in public is fair game, then everyone is welcome to build a multi-petabyte database of text and use millions of dollars worth of GPUs to train an AI on it.
But it also feels like standing at the edge of the sea complaining about the tide coming in...I'm not sure there's really much that can / will be done about it.
But saying "we steal to stay relevant, because we don't have as much funds as our competitors" is not the answer.