Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
Half the reason the books get trashed in this process is because the first sale doctrine keeps copyright from strangling all the freedom in this narrow area.
And if the books are not in the public domain, then they should not be allowed to train their AI models with the material without some kind of license or agreement with the owner of the copyright.
That'd be an easy fix with new law -- if you are an AI company with book data, you have a burden to make openly available (or require your suppliers to) all public domain book scans.
But they won't share them because they don't want their competitors to have the data.
Emphasis mine. Just provide a digital copy, no questions asked. They will distributed to various archives globally. I'll pay for the drives and shipping. I understand and can appreciate the potential liability, and am willing to launder it to preserve the subject collection(s) and dataset(s).