The Vanishing Page: AI Firms Scan Then Destroy Rare Book Editions
dallasexpress.com
dallasexpress.com
What's blowing my mind lately though is the people complaining about copyright today are some of the same people I saw 30 years ago screaming "Information wants to be free" and running around with the DeCSS code on their T-Shirts. But suddenly, now that they hate AI, copyright is important and the information should be locked away. I can't wait for this all to settle down.
Define excessive. Who makes that determination and how?
Among other factors: time value of money and future-value discounting mean that economically ever-expanding copyright terms offer virtually no present-value benefit.
Most works see virtually all of their economic value in the first few years of publication. There are some (rare) long-tailed exceptions. The original US copyright term of 14 years, with extensions, actually fits the economics pretty well.
What extreme copyright duration does do is:
1. Create vast copyright-holding monopolies, often of works for which authors were poorly compensated if at all. (For the latter: academic publishing.)
2. Create a vast category of "orphan works" which aren't or cannot be published, in the latter case because rights simply cannot be clearly established, or competing claims (survivors, estates, publishers) exist.
"Seventeen Famous Economists Weigh In On Copyright : the Role of Theory , Empirics , and Network Effects" <https://jolt.law.harvard.edu/articles/pdf/v18/18HarvJLTech43...> (2005)
"The true impact of shorter and longer copyright durations: from authors’ earnings to cultural creativity and diversity" Jimmyn Parc & Patrick Messerlin (2020) <https://www.tandfonline.com/doi/full/10.1080/10286632.2020.1...>
Orphan Works: <https://en.wikipedia.org/wiki/Orphan_work>
For example, nobody cared about data centre water or electricity use when it was serving us up cat pictures. Collectively the past 20 years of doing that is still significantly more use than AI data centres so far. We suddenly care about it because its anti-AI ammunition, not because we care.
Things being done as part of a massive boom are usually harder to accommodate then things done as part of longer term growth and data centers are no exception.
I don't think I've ever seen people query with such fervour; the energy usage of our existing internet structure. or supply chains to provide us all with electronic devices, despite them being inherently problematic in similar ways.
I was in a British supermarket today and there were apples on the shelves that were transported there from New Zealand. Again, no protests. Do we really care about energy usage or do we simply find convenient ammunition to attack ideas that we are already apprehensive about? We live in a world that is overheating as a consequence of the energy bounty of fossil fuels and yet people are still driving cars, SUVs are more popular than ever. There is no global carbon tax in sight while every nation on the planet slinks from their responsibility for all the co2 in the atmosphere, blaming each other and charging towards our scientist's worst case outcomes. So who really cares about energy consumption on its own?
They validated the concept of copyright by employing it themselves to create the gpl, to make information free and arm it's ongoing freedom.
More troubling to me is the focus on pre-2022 data. This indicates that the companies are seriously worried about model collapse, where the future models become worse becuase they are being fed by generated data instead of real human data.
This is actually my worse case scenario for AI: the models become good enough to disrupt entire industries in the next 10 years, but then stay frozen at that level. Model drift then kicks in because the world will keep changing but the models do not, and in a few decades we actually regress instead of progress becuase they won't be enough skilled humans to drive advances.
But changing how non-AI people write, that's an interesting angle. Because where do we go from here? In 100 years will we still be overvaluing pre-AI sources? That doesn't make sense. Of course a lot can (and will) change in 100 years.
But i'm reminded of pre-atomic steel, which is steel made before the first atomic bombs were detonated and thus have really low background radiation. This is necessary for making MRIs and such. People will go and find it from shipwrecks and such (ironically, many of which are from WW2). It's also a finite resource. What happens when we run out?
Now pre-2022 texts aren't consumed (other than destructively scanning books of course) but it is also finite. We can't make more of it.
Ancient Rome stopped creating aqueducts because they had all the ones they needed. They failed to pass that knowledge on to the next generation and so they just forgot how to create aqueducts.
Once AI starts making the majority of content, people will simply forget how to make content. Before we know it we're all fat slobs in floating chairs like in Wall-E.
Best case scenario I think is similar to what we see in Ian M. Banks Culture books where the machines basically take care of us out of the goodness of their hearts and we just kinda fuck off into obscurity.
The next generations are also losing computer skills, being brought up with devices that handhold you and have no opportunity for customisation and are essentially thin clients for mega corporations.
Each time it’s some different outlet - but when you dig in, the piece is just verbal framing around the original story written by 404 media:
What’s even more stunning is that the original article doesn’t provide evidence that “rare” books are being destroyed. That doesn’t even appear in the original article title.
Will I get sued?
Edit(seriously): >Transformativeness is a characteristic of such derivative works that makes them transcend, or place in a new light, the underlying works on which they are based. In computer- and Internet-related works, the transformative characteristic of the later work is often that it provides the public with a benefit not previously available to it, which would otherwise remain unavailable.(Wikipedia)
I am sure Napster or the like argued that what they were doing was transformative.
Not sure how AI companies are arguing about PUBLIC benefit if you have to pay...
Don’t hate the player, hate the game.
For physical books, they believe they’re not allowed to, hence the destruction.
It’s abhorrent, I agree… maybe if your world domination plan involves destroying rare books, you should change the plan rather than saying “well our hands our tied, the law says we have to”
If I were running the AI company and my lawyers came back with “teeeechnically we can still scan them if we burn them after”, I would reply with “I guess we’re not scanning them”, but I lack the sociopathic instincts required to be a tech CEO I guess.
You should be ashamed of yourself.
"Rare and out-of-print" is fabricating a lot of aura here. It's technically correct (the best kind of correct) but I've yet seen evidence that these are culturally significant copies being destroyed for scanning.
From https://nltimes.nl/2026/06/25/rare-book-dealers-fear-tech-fi...:
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
The oldest is from 1999! If that's the best they can actually enumerate to bolster this outrage farming cycle you could just wonder how irrelevant the rest are.
Really, please, kill this news cycle. There's a lot of issues deserving proper attention right now and this one is a straight-up nothingburger.
Roughly nobody would have had access to that copy. Roughly everyone can now benefit from its content.
You mean, roughly everyone who uses that particular model provider’s products. For truly rare books, this has an anticompetitive flavor, since it ensures others can’t train models from the same knowledge.
Oh no, a lump of cellulose is gone. Will no one think of the fibers?
But it ... won't be? The companies have no motivation to do so. The copyright on lots of these books are surely already expired, so if they wanted to they could be putting these up now. I'm sure internally this is viewed as a corpus of knowledge they have that their competitors do not, so they will not release them unless something forces them to.
> The attachment contained 3,000 English-language titles organized by ISBN number, including books such as Distinct Element Modelling in Geomechanics by K.R. Saxena (1999); Barrett's Traditional Fairy Tales (2021), an academic study of Irish folklore; and Laser Shock Peening of Advanced Ceramics by Pratik Shukla (2018).
Even if the authors died immediately after publication there is no way the copyright has expired.
a) When a company has made a digitized copy of it
b) When a few previous hard copies are kept somewhere in the world
Idk what Anthropic plans to do with this, but note that my original comment was not about Anthropic at all. I simply think it's a good tradeoff to destroy a few copies of rare books if that helps digitizing them. I am not against simply legally forcing companies that do this to release those copies in due time. In the meantime I am also happy to have them as part of the LLM corpus.
> illegally?
No, copying rare books is very likely not illegal, because, presumably, they are quite old.
You mean every Anthropic?
This story has a lot of details left out and is a magnet for illinformed commentators: the reason the originals are destroyed is because it is legal to do so for the purposes of format conversion, but not legal to retain the originals under copyright law.
Confuses "(not all that) old books" with "rare books".
Ragebait, nothing more.
"This is not a novel to be tossed aside lightly. It should be thrown with great force."
Jokes aside, I don't know what I would do in this situation either. Maybe find a way to mark "rare" or "out of print" books and sell/auction the "pages"? Or make the scanned book available via some mechanism?