https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
https://en.wikipedia.org/wiki/Google_Books
https://arstechnica.com/tech-policy/2015/10/appeals-court-ru...
https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli...
I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.
Big companies will read up the books and make their AI recite them from memory, but Archive.org was sued for renting one book on an exclusive basis (unless one would return, another wouldn't be able to rent)
No, this is what they were doing before, but they explicitly started lending out "unlimited" copies, which is why they got sued.
> The crux of IA's first factor argument is that an organization has the right under fair use to make whatever copies of its print books are necessary to facilitate digital lending of that book, so long as only one patron at a time can borrow the book for each copy that has been bought and paid for. See Oral Arg. Tr. 31:10-15. But there is no such right, which risks eviscerating the rights of authors and publishers to profit from the creation and dissemination of derivatives of their protected works. See 17 U.S.C. §§ 106(1), (2). IA's wholesale copying and unauthorized lending of digital copies of the Publishers' print books does not transform the use of the books, and IA profits from exploiting the copyrighted material without paying the customary price. The first fair use factor strongly favors the Publishers.
> In this case, there is a "thriving ebook licensing market for libraries" in which the Publishers earn a fee whenever a library obtains one of their licensed ebooks from an aggregator like OverDrive. Pls.' 56.1 ¶¶ 577-578. This market generates at least tens of millions of dollars a year for the Publishers. Id. ¶¶ 170, 172. And IA supplants the Publishers' place in this market. IA offers users complete ebook editions of the Works in Suit without IA's having paid the Publishers a fee to license those ebooks, and it gives libraries an alternative to buying ebook licenses from the Publishers. Indeed, IA pitches the Open Libraries project to libraries in part as a way to help libraries avoid paying for licenses. See Pls.' 56.1 ¶ 383 (presentation IA gave to libraries asserting that pairing with IA means that "You Don't Have to Buy It Again!"); id. ¶ 382 (different presentation promising that the Open Libraries project "ensures that a library will not have to buy the same content over and over, simply because of a change in format"). IA thus "brings to the marketplace a competing substitute" for library ebook editions of the Works in Suit, "usurp[ing] a market that properly belongs to the copyright-holder."
No, it was not, even supported by the quotes you pulled. Libraries right now, with publisher blessing, offer all manner of controlled digital lending. The suit was because IA did it buy undercutting the publishers copy rights to that legal market. Had IA simply done what every other library has done to provide controlled digital lending, there would be no suit.
In contrast to this, the e-book lending practiced by most libraries with publisher blessing involves the library purchasing special library-specific e-book licenses from the publisher. These licenses contain various contractual restrictions, such as the library having to re-purchase the e-book after a certain amount of time or after a certain number of borrows.
So making it seem as if all "controlled digital lending" is not allowed under current law, even under the description you give, is not that simple.
> “At bottom, [the Internet Archive’s] fair use defense rests on the notion that lawfully acquiring a copyrighted print book entitles the recipient to make an unauthorized copy and distribute it in place of the print book, so long as it does not simultaneously lend the print book,” Judge John G. Koeltl of the U.S. District Court in Manhattan wrote. “But no case or legal principle supports that notion. Every authority points the other direction.” [0]
[0]: https://www.insidehighered.com/news/tech-innovation/teaching...
I'd be worried about how much they'd be charging for access once they had the monopoly on so many rare books.
The copyright owner must make new copies of the work available; the price must be no greater than the original price (not inflation adjusted). And if they fail to do so, anyone may produce copies and escrow the original price (less the cost of production) for collection by the copyright holder.
That means that orphan works are effectively in the public domain. Calculus professors can ask students to get the cheaper 2nd edition, not the latest 22nd edition. And a company like Google could make scanned works available in their entirety for a small amount of money for each work. And the copyright holder still gets their end, without having to arrange a printing or hold stock.
What happens to limited editions of print runs? Can I not make a run of 200 prints anymore because the 201st will be something that someone could request?
Does a musician have to license any song they made to anyone who asks? Can they refuse to license a song to some organization they disagree with and not have it fall into the orphan works category?
---
Amending copyright to the way you describe requires a renegotiation of the TRIPS agreement ( https://en.wikipedia.org/wiki/TRIPS_Agreement ) with all the nations of the WTO (or withdrawing from the WTO).
I suppose the rule could make reference to a rival good, i.e. the 22nd and 23rd editions of the calculus textbook. It should not be reasonable for the publisher to make the 22nd edition only available for $1,000, and the 23rd edition for $250. But that definition would invite many lawsuits and chilling litigation in general. A clear definition based on the historic price is much simpler.
You can still number your limited runs, and your limited runs still have increased numismatic value over some other reproduction.
I don't envision licensing of performance rights in this system, just recordings or reproductions.
Is this a reduction of the rights granted by copyright? Yes, intentionally so. It is stripping copyright holders of the right not to copy, against the interests of society in granting that copyright in the first place.
It keeps about 50% of the books submitted. The other half are donated
AIUI registration only allows for additional penalties during infringement claims.
It's the logic of cutting off one's nose to spite one's face, which prefers that no one benefit rather than Google see any benefit.
It's a very serious issue, very well known to the people with the connections and power to affect it.
> [ it wouldn't ] make a ton of people vote for you to get re-elected
The tons of people are moved by the media, people are oblivious to the tricks of that trade, for the same reasons, obviously. In other words, this issue isn't something that happens to slip below the radar, it's kept stealthy by well organized engineering and considerable expense.
> and it won't create a ton of new jobs
Nothing ever creates tons of new jobs, the "tons" are reserved for promises and other useless noise.
> I therefore feel the same way about Google Books that how I felt when I learned that What.cd went down: that I don't gain or lose anything anyway because I never had access to begin with, and that by not making it 100% publicly accessible you're asking for the data to one day disappear forever.
Can you still search the restricted parts? If so there's still value to it: it helps you identify the book so you do an inter-library loan to get at the full content. Sure, it's not frictionless, but I wouldn't be all or nothing about it.
I don't think they were overcome. As far as I remember Google couldn't make the books available so they abandoned the project. They possess the scans (if they didn't delete them) but they won't be made public.
It's great they did this, but the Google that is today cannot be trusted with data of public value anymore.
This doesn’t incentivize them to be good stewards of this data and making anything in the public domain available. It incentivizes less access to the source material, having to blindly trust their tools, and is effectively automating plagiarism.
Book scans, secreted away, are worthless to the public.
Can you do anything else with the book? What are its costs for storage in a way that retains the value of the book? If the assets of the warehouse are sold to another company (see also https://paizo.com/blog/paizo-restructuring-a-difficult-updat... ), what are your obligations for the format shifted copy that you retain?
These questions imply that there's a liability that exists when retaining the original that has little value to the company. And they (the books) aren't assets that can be resold.
It's easier (and cheaper), has no ongoing costs for physical storage, and answers those questions without creating legal entanglements for the future company.
https://btaa.org/library/programs-and-services/book-search/f...
> Will scanning harm the books?
> No. Google developed innovative technology to scan the content without harming the books. Any book deemed too fragile will not be scanned by Google, but may be treated by expert library staff. Once scanned, all print volumes are returned to the library collections.
That was an inherently different goal (borrow the books from the library, scan them, and return them) than the Bartz v. Antrophic ruling.
https://cases.justia.com/federal/district-courts/california/...
> Storage and searchability are not creative properties of the copyrighted work itself but physical properties of the frame around the work or informational properties about the work. See Texaco, 802 F. Supp. at 14 (physical), aff’d, 60 F.3d at 919; Google, 804 F.3d at 225 (informational); Sony Corp. of Am. v. Universal City Studios, Inc. (“Sony Betamax”), 464 U.S. 417, 447 (1984) (rightful interests). In Texaco, the court reasoned that if a purchased scientific journal article had been copied “onto microfilm to conserve space, this might [have been] a persuasive transformative use.” 802 F. Supp. at 14 (Judge Pierre Leval), aff’d, 60 F.3d at 919 (reducing “bulk[ ]” “might suffice to tilt the first fair use factor in favor of Texaco if these purposes were dominant“). In Google Books, the court reasoned that a print-to-digital change to expose information about the work was transformative. Google, 804 F.3d at 225 (Judge Pierre Leval). And, in Sony Betamax, the Supreme Court held that making a recording of a television show in order to instead watch it at a later time was copying but did not usurp any rightful interest of the copyright owner. 464 U.S. at 447, 455. Important to the Supreme Court’s reasoning was the expectation that most such copiers would not distribute the permanent copies of the work. Finally, in A&M Records, Inc. v. Napster, Inc., our court of appeals recognized the reasoning just explained, and therefore rejected by contrast a digitization effort that was touted as space-shifting but in fact resulted in the multiplication of copies shared with outsiders through a file-sharing service. 239 F.3d 1004, 1019 (9th Cir. 2001), aff’g in this part 114 F. Supp. 2d 896, 912–13, 915–16 (N.D. Cal. 2000) (Judge Marilyn Hall Patel) (citing Sony Betamax and Texaco).
> Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).
---
The AI training isn't borrowing from libraries and returning from libraries. Instead, it is buying a book, format shifting, and retaining that format shifted version from its own use. The company can't do anything else with the book once they've format shifted it. They can't donate it and they can't resell it. In that light, destructively scanning the book is the best option. There is no value in trying to non-destructively scan it because otherwise all it would do is sit in a warehouse and cost money to pay for storage of something they can't sell.
Perhaps that's an important distinction here, I was really just trying to explain what I thought the commenter was trying to say. Calm down with the copy pasta walls.
But I still think that AI companies could make a very similar argument to the one Google made in your copy-pasta. The laziness I refereed to earlier is the fact that they haven't even tried. They could donate the books afterwards which would go a long way towards helping that argument in court.
First, in today's environment people will state AI hallucinations as fact or "google it yourself" as the reference. If someone doesn't know where to find that information they're left with the "someone on the net said XYZ".
Secondly, I'm not always certain that if I do provide a link to a large document that people will find the relevant section in there. And second and a halfly, if someone comes back to it in a year or two or five that the link will still be live. I've had situations in the past where I've provided a link and then the domain changes hands and the new owners of the site put up a retroactive robots.txt and make it inaccessible on the wayback machine.
---
The question for an AI company that can't return, sell, or donate the material after it has been scanned - are they then to pay to store it in a warehouse indefinitely?
After they've format shifted the content and retained the format shifted content for continued use in training, it becomes at least a very gray area to donate the original.
And even if the books were scanned like Google Books did, and they could donate them to libraries - the libraries don't want those books. These aren't rare books in that people would expect you to wear cotton gloves while handling them... they're books that have mostly disappeared from availability.
https://old.reddit.com/r/books/comments/1vugion/the_federal_... provides an example of what is being seen in a used book store:
> I work for a large used bookstore with an online component. We're getting slammed with orders for books like the proceedings of an obscure 1992 Dutch geology conference or $500 festschrifts about D-module applications we would have previously sold to some university library. We've never once had an order for anything anybody would actually want, and most of this shit has sat on our shelves for years, if not decades.
Neutron Radiography: Proceedings of the First World Conference San Diego, California, U.S.A. December 7–10, 1981 is technically a rare book. https://www.amazon.com/Neutron-Radiography-Proceedings-Confe...
If you had a copy of it, I would challenge you to find a library that would accept it as a donation.
In the event that you wanted to read a copy of it, there is a copy of it in the Library of Congress. https://search.catalog.loc.gov/instances/b220c9bd-63ad-5a8b-...
and that you're arguing that a library could accept a donation.
From Bartz v. Anthropic which was mostly about the pirated copies, but included some information about the destructive scanning, format shifting, and use as training as fair use.
> Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others).
> ...
> As a result, Anthropic’s format-change from print library copies to digital library copies was transformative under fair use factor one. Anthropic was entitled to retain a copy of these works in a print format. It retained them instead in a digital format, easing storage and searchability. And, the further copies made therefrom for purposes of training LLMs were themselves transformative for that further reason, as above.
Note especially the second sentence in the first paragraph. "The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company."
If the physical copy was shown, shared, or sold outside of the company the fair use defense would likely be much weaker than it currently is.
After the company has format shifted the original and retained the digital version, it cannot share or donate the original back to some other organization.
A library might take the books. It would be a legal liability for an AI company to give those books away and weaken the fair use claim that it stands on from Bartz v. Anthropic. The destruction of the original enhances the fair use claim.
From the ruling:
> For the print library copies that Anthropic purchased and then converted into digital library copies, Anthropic already enjoyed entitlement to keep the copies in its library. The purpose of the copying was to keep them in its library but with more favorable storage and searchability properties. Copying the entire work was exactly what this purpose required. There was no surplus copying. The source copy was destroyed.
> The third fair use factor favors fair use for the purchased library copies converted from print to digital.
There are no surplus copies when the original is destroyed.
They never should have been trusted in the first place.
Info on where to send books not yet in their collection: https://help.archive.org/help/how-do-i-make-a-physical-donat...
Mobile apps to determine if they need a book: https://help.archive.org/help/donate-books-app-for-ios-and-a...
Web app: https://archive.org/want/?mode=donation_book
For example, I donated a copy of Systems Bible (out of print, hard to find imho) and paid for it to jump the digitization queue (https://archive.org/details/systemsbiblebegi0000gall/). The original book will remain stored as a physical backup. It's not fully publicly available of course due to copyright (it will eventually be made public by the Internet Archive once its copyright expires ~2084 and it enters the public domain), which is where shadow libraries|archives like Anna's Archive and Z-Library fill the gap.
If you have rare books you would like digitized, archived, and distributed, I am very interested in providing assistance.
While the hard copy is a bit pricy for my shelf of curious books, it's also available on kindle. https://www.amazon.com/SYSTEMANTICS-SYSTEMS-BIBLE-John-Gall-...
This is was what ended up spinning up AWS as a business.
I dislike the idea of destroying rare books--but how rare? A digital copy has a lot more benefit.
At the time it was obvious and innovative but over time it was clear Google was establishing a technical precedent to corrode what libraries have the power to do. I'm hoping we can continue to do the good work but it's exhausting.