Training an LLM on copyrighted works is not illegal.
This whole debate has been tried in court already. Calling it IP theft only stands on individual moral grounds, but the law allows for derivative works.
> HathiTrust Digital Library is a large-scale collaborative repository of digital content from research libraries, administered by the University of Michigan. Its holdings include content digitized via Google Books and the Internet Archive digitization initiatives, as well as content digitized locally by libraries.
https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._HathiTr... and https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,...
> Authors Guild, Inc. v. HathiTrust (2014) was a following case related to HathiTrust, a project by the libraries of the Big Ten Academic Alliance and the University of California systems that combined their digital library collections with those of Google's Book Search. The HathiTrust case differed in two primary factors which were raised by the plaintiffs: that for viewers with disabilities, they could view the scanned text through a screen reader to make it easier to read, and offering to print out the scans as replacement copies for members of the universities if they could verify their original copies were lost or damaged. Both uses were deemed also to be fair use by the Second Circuit.
> The subject of the copyright of orphan works – works that may still be under copyright but with no identifiable rights holder – was a significant point of debate after both this and HathiTrust. Normally, libraries have been hesitant to loan digital copies of orphaned works as libraries may be liable for copyright violations should the copyright owner step forward to claim ownership.
The bill on orphan works that didn't pass was https://en.wikipedia.org/wiki/Shawn_Bentley_Orphan_Works_Act...
https://www.hathitrust.org/the-collection/search-access/copy...
And there are exceptions for copyrighted works allowing them to lend them out.
> Protected by copyright law, but made available: Protected by copyright law but made available on a strictly limited basis in accordance with the statutory limitations including, but not limited to, Section 107 provisions for fair use, Section 108 provisions for libraries and archives, and the rights provided to registered users with disabilities. In the absence of an applicable exception, no further reproduction or distribution is permitted by any means without the permission of the copyright holder. Lawful uses of works are provided only under the following conditions ...
Scanning still continues. https://www.hathitrust.org/member-libraries/contribute-conte... - though it's not at the same rate as it was during google books project.
People obviously feel bad about companies doing this. People reading these stories don't care what's legal, they care what's ethical. Heck, re-publishing long lost material would make AI companies heroes instead of bad guys.
I don't think you have any idea how expensive it is to acquire the copyright for a single book with the intent of making it freely available online. That's equivalent to asking the rights holders to perpetually forgo all possible earnings from the material, and they expect to be compensated accordingly. Even paying lawyers to begin assembling what's needed to make this happen would be five figures per book to get started.