Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered.
My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
If the book was just rotting away in some forgotten bookstore, it would more likely be unceremoniously disposed off in the future without anyone scanning it first.
And refusing to do this exercise just means that you behave as-if you put a really silly number on the value of human life, and probably not consistent between different parts of the project.
So I don't quite agree with these taboos in the absolute.
(I'm still against capital punishment on practical grounds.)
For books it's similar: if you taboo book destruction for the AI training folks, that doesn't rescue books from their ordinary pre-AI life cycle of getting destroyed all the time in the course of running a publisher or a library or a second-hand book store.
In fact, the AI craze is what's giving rare books _value_ and incentivises people to dig them up and preserve them. Or at least preserve them long enough to be scanned.
The scanning might destroy the physical copy of that book, but they save the contents. That's the whole point of scanning after all.
Roads that cause random deaths or are the reason for accidents (whatever by their physical design, control systems, surface coating, or allowed speed) are generally illegal in most places and require strict exemptions for geographically challenging places.
https://en.wiktionary.org/wiki/wastebook
Oh right. But anyway, nobody knows what needs preservation, it's a basic problem of life, somebody usually mentions the BBC throwing out boring old Doctor Who tapes to save archive space because nobody liked it any more at that point in time. Some things should probably be thrown out now and then, I suppose.
So at least with the AI companies they are scanning them and preserving them digitally. Not just in the trained weights, but also as raw training data for future runs.
P.S. I'm not sure why you need to make fun of your own ignorance? Just look up the word you don't know and don't mention it?
Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash.
That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
In the UK at least, people usually take them to a second hand / charity shop, who sort through them and send the valuable ones to auction (typically early editions, 100+ years old) and then either sell them themselves (for recent books that are easy to get rid of) or sell them to specialised second-hand bookshops.
Most of the specialised second-hand bookshops rarely throw books away, usually if nobody buys them after a couple of years they end up in the extreme discount piles (20p, 50p etc) and probably only trashed if they still don't sell from there.
I do get the _idea_ of preserving books, but... people don't care. I just threw out well over a thousand books from my grandparents house this spring.
There were ~6-10 "valuable" books there. Two because I personally knew someone who wanted old war-time books and a few 100+ year old bibles. And maybe two dozen books worth saving, mostly because they were from big-name authors or had stuff that nobody would print anymore (I have detailed instructions how to make laughing gas and how to build an underground chemical lab - hobby books in the 50s were ... interesting :D )
I literally couldn't give away the rest. And I tried. It was all just "interesting, but..." - no way to justify using the shelf space for books that, realistically, nobody will actually ever read again.
There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further.
In short, it's a mess.
At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
No, you have a translation.
>Even if, like the Odyssey, there are hundreds of wildly varying translations?
Precisely why translations are not considered equivalent to the original text.
>What about an abridged copy? What about the Sparknotes version?
An abridged copy is not a copy of the unabridged version.
>If you have a copy of Pride And Prejudice And Zombies, do you have a copy of Pride And Prejudice?
No.
I'm honestly surprised these were the questions you chose to ask, when you could have asked what if you have 90% of the pages, or what if most of the pages are missing pieces because the book was shot with a shotgun, or what if the book was scanned and OCRed and all the "rn"s were replaced with "m"s and all the lower case Ls with ones. Hell, is a scan of the book close enough to having the book, or is it far enough that one can no longer be said to have the book anymore?
If I have a translation of a book, I think I have more of that book than if I had nothing at all. It's fuzzy.
You seem to be arguing that a translation, etc is not literally having the book, which is something that has always been my stance as well.
We just have different opinions about what it means to have a book. I think that having a copy of "Pride and Prejudice and Zombies" is more like having a copy of "Pride and Prejudice", than having no book at all is like having a copy of that book. You disagree and that's fine.
The reverse is not true. Just because someone asks a rhetorical question, does not mean they are using the Socratic method.
There's no technical reason why an LLM couldn't reproduce verbatim some of the training material. It's sort of a lossy statistical compression engine. Enough of the info will survive to the output in the original form. With the amount of data and the commercial nature it's hard to argue fair-use. But nobody tested this in court. I'm not even sure the US wants to ever test this. Why even attempt something that has a non-0 chance to sabotage your most promising industry/bubble in ages?
These generally are not books people care about. The information contained therein was doomed.
Now the information has been digitally preserved and a digestion of the information will be made publicly available.
But you'd think that the Library of Congress and such would actually prevent stuff from vanishing just by collecting it themselves.
I'd rather a digital copy exist in someone's hands than a rotting physical copy.
It almost seems like you're suggesting that having Claude generate a paraphrased book is as good as having the original book but i don't think that could be your intention?
One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.
Most of these owners aren't doing anything particular to preserve them, they're stacked up with 10,000 other TV Guides in a hoarder's moldy basement. Anyone interested in keeping October 1994's TV Guide pristine has had over 30 years to procure and protect their copy.
These particular buyers are converting their copy to a digital one that's getting some kind of use, which is better than the fate of 99.9% of the other copies.
Everybody imagines libraries hoping against hope to get their hands on all these rare books so they can shelve them and preserve them for generations to come. But if you donate a box of old books to a public library, there's a good chance the clerical staff handling that donation is going to roll its eyes and cart the books out to a dumpster. Large public library systems have stopped taking book donations, for this reason. "Bring them to a thrift store" is what they'll tell you.
The argument wouldn't hold even if you had been on the library's case, because it misconstrues the normal lifecycle of a book. But we don't even get there unless you've been consistent about this, rather than instrumentalizing it at just this precise moment to jump on a trend story about hating AI.
I’d personally choose the latter, especially given that the 1994 tv guide is not going to meaningfully improve the utility of the language models.
Direct access to pre-digital history is drying up rapidly, why accelerate that for incremental benchmark gains in a domain that isn’t even relevant to the most useful forms of a nascent technology?
A physical book in a museum archive is useless for 99% of the worlds population even if they really wanted that specific book and were able to find it, as they'd have to arrange for access, then travel (at incredible expense) to access it.
Maybe they could ask the museum to digitize it, but that's still going to be days of delay and tens of dollars of cost to access parts of that book, if the museum even offers that service. If we go slightly beyond your "melted into" statement, the chances of the book becoming useful to the public are much higher in the AI company's digital archive, which might turn into something like what Google Books could have been, given the right incentives and copyright law changes.
And of course that presumes that the TV guide is going to stay in the museum rather than been thrown out as part of curation (or realistically, long before it makes it into a museum). Neither museums nor archives hoard everything, throwing stuff out is - as far as I know - one of the key jobs of an archivist. And a 1994 TV guide, while useful to understand what it was like to be alive in 1994, likely doesn't contain much unique information. You don't need that specific guide.
If there are 52 weekly editions, of 10 different guides, you would likely get most of what you want from any one of them. And for the parts that you wouldn't - there's a good chance that you'll have a much easier time getting the essence of this knowledge from the anonymized data pool that all the content was melted into, rather than chasing 10 different museums to find the original magazines.
https://genome.ch.bbc.co.uk/about
Historic TV guides are also the sort of strange ephemera that people collect. They ought to be digitized like newspapers and other magazines, but this was always the purview of libraries anyway.
I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
The legal ruling from Judge William Alsup declared that if AI companies purchased the books legally and then copied them to their servers, it was fair use as a "transformative" operation, but the originals had to be destroyed in that case, because then there was only one copy still in existence (the one on Anthropic's servers):
From https://www.theguardian.com/commentisfree/2026/aug/05/anthro...
> Under US copyright law, the “fair use” doctrine allows you to make “transformative” use of copyrighted works without the owner’s permission. Anthropic took printed books and scanned them, “transforming” or remediating them into a new, electronic format. They then disposed of the original printed copy: the “destructive” part of destructive scanning. Along the way, Anthropic’s vendors had already sliced the spines and edges of the books, to scan them more easily before destroying them. “One replaced the other,” as Judge William Alsup wrote, noting: “There is no evidence that the new, digital copy was shown, shared, or sold outside the company.”
This site, https://copyrightalliance.org/education/copyright-law-explai..., states "It is important to note that this exception for backup copies only applies to computer programs and not to other copyrighted works, such as digital movies, music, or photographs or ebooks." But it seems unbelievable to me that they would have spent millions copying all these books and not have backups.
They just don't want to pay what the copyright holders want to charge
I’m not aiming this at you directly by: ISBNs or STFU
Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for.
People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy.
It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
BBC good enough for you?
https://www.bbc.com/news/articles/cp3rprx2wl4o
"A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's.
"But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century.
"It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them. This isn’t complicated. Amazon/etc aren’t breaking into museums and libraries, they are buying books on the open market.
If these books are so rare and important, then why has no one cared until now to actually preserve them?
You're repeating the seller's contextual point, as if it's a counterargument. Do I need to explain to you that what makes the lone surviving 18th century edition important is that there aren't 75 copies of it in museums?
> And lastly, if these books are so important, then don’t sell them, hold onto them, digitize them without destroying them.
It's good to see you agree any digitization of this category of book should be non-destructive.
> If these books are so rare and important, then why has no one cared until now to actually preserve them?
(Lastly for realz this time, eh?) Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.
> It's good to see you agree any digitization of this category of book should be non-destructive.
I don't. Digital or physical, it's the same, there is no difference in my mind. I have many paper books but they are art, not functional, I also have the ebooks which is what I actually read. The _ideas_ are what's important, not a dusty, decaying shell in which the ideas are contained. I wouldn't shed a tear over every library digitizing their books and destroying the physical versions, nothing is lost. More importantly, once you've bought something it's yours, yours to read, yours to display, yours to destroy. On hacker news, of all places, the people fighting _against_ first sale doctrine is appalling.
> Why has no-one cared to actually preserve the actually preserved book being sold by the bookseller... Bit of a strange question, that.
Our definitions probably differ here but preserving is not storing a book, preserving is ensuring that even if this copy is destroyed the ideas inside live on. I think that people that hoard (actually) rare books without a thought or care to making sure the text inside is preserved for future generations out of some desire to simply own something rare are the actually monsters here. And let's dispose with the notion that booksellers are "preserving" books, they are holding inventory, inventory they were happy to sell to Amazon/etc. If AI companies were raiding museums at gunpoint we'd be having a different discussion. They are buying books for sale, if they keep them in a library at corporate or scan and destroy them it makes no difference.
Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.
On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
Why are AI companies forced to shred books?
But how about the importing into the AI tool?
Does this "transformation" somehow override the authors complaints of AI taking their book and not compensating the authors for its mass usage in the AI tool.
Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them?
You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
Source? You state that in a tone that implies you have verifiable knowlege of this. All the information I found says that the exact number, titles and authors are under NDA.
1. This is not the 404 media story you quoted as source
2. Maiberg's answer is most likely a book that has no emotional value. Translation: they do not know
After scanning and destroying 10 million books, AI companies will have destroyed 10 such books.
(I am obviously assuming a Poisson distribution here where I pulled the parameter p = 1/1000000 out of thin air; point is p is non-zero; and at this scale the undesirable event is bound to happen a bunch of times).
I assume no such things of AI companies.
There is also a fundamental difference in intention. AI companies are seeking out old books and will destroy them. Public libraries are trying their best but simply don‘t have the resources to find rare books in every collection.
And finally there is a fundamental difference in scale. Public libraries are not buying millions of copies to destroy, completely eliminating the chance they will ever be discovered.
---
PS. I don‘t like the anti-academic tone of your post. Librarians know what they are doing, their expertise is valuable, and when they act according to their specialized knowledge it does have a positive effect on the world.
I'm not "anti-academic". I'm just very involved in my local community and I read the annual reports from our library system. We accept book donations, like a lot of suburban library systems do (from our extremely book-y community), and we are open about the fact that most of those books get trashed. We do not have the personnel on staff you claim libraries generally do.
I would say that between the two of us, I'm the one arguing more respectfully about libraries. I see them as real institutions with real pressures and constraints that do an important public service, and you see them as an instrumentality in your argument against AI.
I pushed back on that comparison. These behaviors are in fact not comparable. What the AI companies are doing is bad actually, and what libraries are doing is, while unfortunate, acceptable, given the limited resources they have.
What I don't understand is your response to this. An AI company digitizes a book before destroying it: "bad actually". A public library takes a cartload of books to a dumpster without so much as opening the front cover of any of the books: just fine.
AI companies buy boxes and boxes of books with unknown content staff anybody who can operate a scanner and have no idea which books they are destroying.
If a public library comes across a book they didn’t know they had, it is very likely that somebody will notice and know how to continue, to find out if this book is worth saving etc. AI companies will treat this book exactly like any other and destroy it to feed the plagiarist machine.
The very act of cataloging donation books takes a huge amount of resources that the collection managers don’t have.
Many public libraries don’t have a collection management department at all! They outsource that function entirely to vendors and those vendors mostly source new books and spend resources preparing them for the hard use of a library (changing bindings, uploading and cleaning metadata, tagging, etc).
Decommissioning is done by volunteers who just look at the condition of books and chuck the bad looking ones in a bin for recycling.
My wife worked in this industry on the vendor side, your model of public libraries is closer to a small subset of certain big city libraries narrowed to their rare and research departments. The median book bought by a library is a Danielle Steele romance novel packaged for library consumption and sent to a small town library system that will be lucky to have a single professional librarian for the whole system.
> Most public libraries do not have professional archivists
Most reasonably sized library systems do in fact. Even small ones have trained librarians with an some degree in library science who took at least an introductory class in archiving as a part of their degree, and very likely has some idea how to prevent rare books from being destroyed.
* It's happening on a much smaller scale
* They're actually digitizing the books before they destroy them
This obviously isn't about the books. It's about people want reasons not to like AI companies.
As does the intention matter. The AI companies are doing this because they want to profit off of it. They don’t have to do this, and the fact that they do is part of what makes them bad for humanity.
I honestly don‘t understand what kind of a mission you are on here, and why you feel the need to defend AI companies by whatabouting libraries into this conversation. But whatever you are doing here, I think you have succeeded. This thread was supposed to be about the unambiguously bad behavior of AI companies, but now it is about libraries.
But OK. The numbers don’t add up. The 404 article talk about millions of copies, you mentioned hundreds of thousand. If we take these (extremely vague and unworkable) numbers at face value (which I am willing to do with the 404 article but not you because I trust journalists more than strangers on the internet), the scale of AI companies is one order of magnitude larger then libraries.
I still think this is a bad comparison regardless of the numbers. Seeking out books to destroy is worse behavior then not accepting donation. There is no equivalence in behavior here to compare.
So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.
This is just one example, but it has become unfortunately common across all social media platforms.
It's an insane kind of hubris to do what they did, thinking it would change the world, and that no one would notice what they did. So I don't even really agree with your characterization, as they were convinced what they were doing was a risk worth taking, on the assumption it would be successful.
Unfortunately 50 years after the death of the author (or 50 years after publication for corporate owned works) has been locked in as a minimum term through international treaties, so it'll be somewhat hard to lower it beyond that, but many countries (including the US) enforce much longer terms, so that would be a first lever that could be applied quickly.
Maybe countries could could also establish an exception for out of print books offered to the public for free, or a general "library exemption" for public archives after a certain number of years?
I'm sure one of the AI companies would be willing to host a LibGen style library as a PR measure if legally allowed (with sign up required for rate limiting and as an extra benefit for the company to get daily active users).
The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.
What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap.
Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
Not "need to ingest"
Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred
So in the end it's really a Congress problem as usual
Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand.
The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory.
Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?
Who is going to claim a copyright violation?
Most developed countries have a 'legal deposit' system with a national archive/national library that requires publishers to send a copy of their works to them. They've existed in some form for centuries in some countries.
Example for the UK - British Library guidance: https://www.bl.uk/services/legal-deposit
They are not forced to shread them by copyright.
Nothing forces them to shred books, they do it because it's slightly cheaper that way.
https://en.wikipedia.org/wiki/Project_Panama
So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.
They can contact the copyright holder and ask/buy a license to make multiple copies.
Also, who’s forcing AI companies to “ingest” books in such a destructive way?
Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
Anna's Archive, for one, would be more than happy to host at no cost to the author.
The law is currently forcing these companies to destroy the books after scanning them.
No, the law is stopping them from digitally sharing their scans. They are perfectly capable of reselling or donating or storing the books they buy. (Wasn’t Amazon originally a book seller?)