Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
Doing the scan within their own organization does not because it's within a single legal entity and there is no copyright issue involved.
You could overcome it through various means, including helping to set up a clearinghouse for bulk rights, implementing a data clean room approach that models could train on in situ, lobbying Congress for copyright law changes especially around orphan works, using Section 108 of the Copyright Act to set up a specific preservation vehicle like the HathiTrust, and other options.
I find it ridiculous that so many in this thread are acting as though AI companies are simply powerless to do anything but buy up and destroy these rare books.
That’s the whole point: Companies are doing this because it’s easier and cheaper for them, and those doing it now are first-movers who don’t care about how catastrophic the end result will be for the rest of society as they’ll have their training data moat.
It should be obvious, but the things that are most profitable are not always the things that are best for society as a whole. That’s why we have regulation.