Anna's Archive Owes $340 Million, Lost Several Domains, but It's Still Online
torrentfreak.com
torrentfreak.com
I deeply admire the people who are obsessed with their passions and strive to build things that will lay the foundations for others.
Coming out would be a bold move for them.
(not legal advice!)
One wonders if they're still doing it.
Plus this is not legal in the EU (and Canada, and ... let's just say the entire rest of the world, and accept that I'll be wrong for one or two smaller countries). Doesn't that matter? Or is only Mistral disallowed from training on copyrighted materials? Je veux ma chaton fat, goddamit!
And where are you getting the idea that Mistral doesn't train on copyrighted data? There's not a lot of code written by people who've been dead for more than 70 years, but somehow Mistral has been able to release coding models anyway.
Hell, I know that for one "lab" (kindof AI lab) since 2020 or so has determined wikipedia quality is dropping fast. It was already dropping slowly before that, but now it's getting bad.
Anna's Archive didn't exist in 2019.
It is almost like we need a "for use for training" agreement across the board. This would not fix the current issues (at least without substantial work), but going forward would allow for creators (or publishers/rights holders) such as this to designate a work as crawl-able for AI. A robots.txt just for Claude.
I always wonder why y'all feel the need for these impressive mental gymnastics. You can use the models /and/ think they are trained unethically. Living through the ambiguity without abandoning your ideals completely is a valuable skill these days.
Instead, the ongoing lawsuits focus on the idea that AI training involves making additional copies, for which they would need a copyright license instead of just one legal copy.
Also, as one more example, I find it hard to believe that their models could generate 'Studio Ghibli' style images without training on the movies. There is no licensing deal between them.
I think the real issues here are two-fold:
Firstly, Copyright is very ill equipped to handle these cases. Just because the model is tuned not to output the exact training data does not mean that compressing mostly-copyrighted datasets into a proprietary model is ethical, fair or /should/ be allowed, simply because they might destroy entire livelihoods. If you take those copyrighted works away you are left with, in OpenAIs own words, a cute little experiment.
Secondly, there is absolutely no transparency. Datasets are easily deleted and its impossible to tell what the models have been trained on, especially after fine tuning. Moreover, only the biggest most successfull works would be easily identifiable without the fine tuned model. Once again, sticking it to the little man.
Citation?
* We should not sell powerful chips or chipmaking equipment to China
* We should crack down on industrial-scale distillation operations
That's not quite a ban of local ML inference, but it basically says that he doesn't want companies in China to create their own state of the art models.
1. Citation
2. Source
3. Reference
Long ago in a career based on original research, I/we ONLY used "reference."
This seems like a "heads I win, tails you lose" type of argument. If Anthropic was pro-piracy I can imagine everyone getting mad that they're flouting law and want to "steal from artists" or whatever.
>and continue doing it
Source? AFAIK they were caught and stopped. That's why there was the recent story about how they were destroying old books to scan them.
https://cdn.arstechnica.net/wp-content/uploads/2025/06/Bartz...
Sounds pretty pro piracy to me.
I thought they could've bought just a single copy of each book and use the content to train their models. In that case, it falls into the fair use doctrine and they wouldn't need to pay the fine. And that will be way less expensive than the $1.5B price tag.
But when it was a proof of concept, they were using pirated data.
Just like Spotify did.
Furthermore, every pirate wants to be an admiral. None of the big tech companies are actually in favor of any amount of copyright reform. They never have been. There is a huge gulf between "personally benefitting from copyright theft" and "actually wants to legalize the theft". Anthropic still believes they deserve to be paid for their models, they just have this delusion in their head that doing a bunch of computation on stolen data is equivalent to actual human creativity.
[0] OpenAI, Anthropic and Facebook have been shown in court to be using shadow libraries, I don't know about Google.
INB4: "Here is one time one of those organizations broke the law". Don't go there, absolute lowest level of conversation.
Although I partially applaud what Anna's Archive is doing, I like their model the least (they want to CHARGE to download while doing Copyright infringement?).
In any case, I strongly believe that Anna's Archive is the wrong approach, as it has a single point of failure. We have been doing massive P2P sharing for more than 26 years; we have the algorithms for fully distributed file sharing and databases. Why are we still depending on HTTP/DNS based interfaces that are easily taken down by people wanting to limit knowledge?
I'm glad and thankful that the people behind Anna's Archive dedicate their time maintaining the huge base of human knowledge (Encyclopedia Galactica Asimov would say), but we (the people) should make it really distributed, really infallible and accessible (no, downloading 10TB torrent files doesn't make sense, except for archiving purposes).
We should have something like Popcorn Time but for knowledge.
This is bollocks. AA gives users a means of paying to enjoy faster speeds as a means of contributing to costs, but the downloads are free to anyone who doesn't want to pay, and very often quick enough.
PlayStation recently sent out emails regarding this, effectively reminding people that they did not, in fact, own what they purchased[0]. It's very confusing - someone here would have a much higher chance of knowing that it's not a direct purchase or right to ownership, though I'll bet most of the general public do not.
[0]: https://windowsreport.com/sony-emails-playstation-users-to-r...
I will say, however, that they've gone up against some very scary people with very large teams of lawyers, and for that I salute them.
Those are arguably the hardest protocols to block on the open Internet without causing major issues for all other sites, forcing those trying to take them down to play "whack-a-mole". If they were to create a new "AATP" for distributing data, it would make it trivial to block on every ISPs firewalls.
They railroaded Aaron Swartz (which eventually lead to his suicide) for much, much less.
One side wants to freely share knowledge with all of humanity, the other wants to restrict it to make a buck. I know who I support.
The only reason we even have dead-tree libraries at all is because 100 years ago, that was what John Rockefeller and Andrew Carnegie put forth to whitewash their horrible capitalist behaviors across the USA. And because it was done by those generations' billionaires, public libraries because acceptable.
If public libraries were created in the last 20 years, they would have been banned and felony copyright charges would have been levied. In fact, thats exactly what happened WHEN people tried to create free digital libraries.
I wonder why nobody created a Netflix-like subscription for digital books
The economics aren't there.
Additionally, modern Netflix is a streaming company, and even DVD Netflix was sending the materials directly to your house, which is a differentiator versus the library, whose materials need to be picked up. There's a convenience factor as an incentive to pay. Library streaming exists, but it's awful - very limited library and very limited watches - versus Netflix where once you sub you can watch as much as you want.
There was already one that existed in the web 2.0 era. Founded in 2012, raised a $3M seed from Founders Fund and a $14M A. Completely failed though and was acquihired by Google to have the founders lead Google Play Books.
Amazon now has Kindle Unlimited but I think it's mostly romance slop and self-published books. Seems like publishing rightsholders were just too inflexible to let the business model take off.
The reason I don't subscribe is it takes me a lot longer to finish a book/audiobook than a movie. I just get them from the library/Libby - even if it's a 4 month wait, there are plenty of available books to read while I wait.
In general I don't have many qualms with modern piracy, it just seems very hypocritical and I'm confused.
I'd imagine the backlash to AI companies from us white collar workers to be less severe if they have to publish their weights. In fact if you look closer you will see HN is actually pretty content with Chinese open models. It's the American AI corps with closed models that attract criticism.
One directly threatens the livelihoods of the commenters here, the other doesn't ;-)
But, in general, Stewart Brand’s sentiment — expressed during a panel discussion at the first Hackers’ Conference in 1984, when he said that “information wants to be free” — is broadly shared here.
His full quote was: “On the one hand information wants to be expensive, because it’s so valuable… On the other hand, information wants to be free, because the cost of getting it out is getting lower and lower all the time.” Steve Wozniak famously responded: “Information should be free, but your time should not.”
Even among those on HN who believe that copyright serves a societal purpose, many still want copyright terms to be drastically shortened (back).
So I’m skeptical that the hostility toward AI companies is primarily based on their training on, and profiting from, the world’s knowledge without licensing it. I suspect it is rather based on their doing so without making that knowledge accessible in return. In the recent news of their destroying books, it even means reducing the chance that that knowledge will ever become open.
Yes, opening the world’s knowledge can lead to good and bad uses — just like free software can be used for great as well as for nefarious purposes. We’re generally too techno-optimistic to judge technology solely by its worst possible outcomes. But for such a benevolent view to be justified, the potential benefits need to be also clear.
Might as well be a trillion dollars as it's never getting paid.
Fred and Wilhelmina are in bed together. Fred tosses and turns, unable to sleep. "Fred, what's troubling you?"
"You know Bjarne from the bank? Well, a balloon payment is due tomorrow on my loan, and I don't have the cash flow to pay it."
Wilhelmina thinks for a bit, then reaches for her phone. "Betty? Yes, this is Wilhelmina. Sorry to call so late. Would you tell Bjarne that Fred can't make the loan payment tomorrow? Yes, that's all. Good night."
Fred states at Wilhelmina, aghast. "What did you do THAT for?" She smiles. "Now it's Bjarne's problem. Let him toss and turn, you can go to sleep."
I remember when RIAA was shaking down 15 year olds for $7000 for a single Metallica download from Napster. Or how Aaron Schwartz was executed by proxy by JSTOR and the feds, for what should have been free to access for all.
But hey, Anthropic, OpenAI, X, and others can pirate to their hearts content with for-profit piracy, but "we" (royal) are OK with that. We just cant have the poors have access to the sum of human knowledge.
Oh yeah? Well, my dad works at Anna's Archive.
Why didn't Metallica's fans rebel? That would have stopped it quickly and set a precedent and example for the rest.
What an incredible strategy - and it worked! And in something that's entertainment and fandom-driven; it's not like the fans need a medication, job, or their auto warranty. How pathetic. The day rock'n'roll died.
> Take, for example, the case of the Tammy Lafky, a 41-year-old sugar mill worker and single mother in Minnesota. Because her teenage daughter downloaded some music last year—an activity both mother and daughter believed to be legal— Lafky now faces over $500,000 in penalties. The RIAA has offered to settle for $4000, but even that sum is well beyond Lafky’s means—she earns just $21,000 per year and receives no child support.
It gets worse:
> Among those sued was Brianna Lahara, a twelve-year-old girl living with her single mother in public housing in New York City. In order to settle the case, Brianna was forced to apologize publicly and pay $2,000.
How could one justify tearing cash from the hands of poor children? The irony is that they may have been downloading music because they couldn't afford it in the first place, yet wanted to keep up with their peers.
In response to the 340M fine, .... "We have set up a LLM and are deleting all our books ? "
Eventually US will have its own great firewall like China.
Politics is about powergrab and what we’ve seen is more and more power grab.
It’s likely the billionaires are able to buy the govt goons to pass the laws that allow them to be the gatekeepers.
Anthropic, OpenAI and the model builders massively benefited from the archived information. Distill entirety of archived human knowledge.
Piracy is morally justified at this point.
Thankfully I don't live in the US, and likely won't set foot there ever again.
Maybe if they use encryption, which would make them all pedophiles according to those governments. Hope you have your VPNs ready if visiting countries without a free internet, like North Korea, China, UK, or EU.
The term "piracy" is used by record companies to demonize sharing and cooperation by equating them to kidnaping, murder and theft.[1]
[1] https://stallman.org/articles/end-war-on-sharing.htmlbetter not miss!
They're not more or less virtuous than any others, they're in it for the money and you're a fool if you think otherwise.
I'm not against piracy, it's great, but to claim they are moral saints is a joke.
Anyway, bring in the downvotes as I know will happen.
If you just Google 'free media heck yeah' you'll see a great deal of them!
Is your angle merely that shadow library hosts should be operating their services entirely for free (to the point of refusing payment) or is there something we're missing here?
OTOH, if you think that AA is robbing creators and publishers of their hard-earned proceeds, and that loss is greater than the loss to humanity from the destruction of AA, then you wouldn't think AA was a moral good.
Of course, it's all gray in the middle.
You cannot claim that AI labs are bad for illegally downloading content to build their AI without paying creators, and that AA is good for illegally offering the same content.
I think both are immoral, but I definitely pirate all the content I can (except video games but only because it's unsafe, not to remunerate creators).
It is not grey at all, this is some bullshit that people tell themselves to feel better. Whether it's using the output of AI models or downloading a book on Anna's Archive, you are ultimately robbing the creator of profits.
Nobody on HackerNews would argue otherwise if their employer stole their code from their mind and didn't pay for it.
Of course on a related note, most work that's interesting to me was created by people who are dead anyway, so they don't mind.
Anyway, are you sure most people that support AA don't also support e.g. Deepseek and Qwen's efforts?
You're correct in the above, or at least I'm willing to accept the premise for the purposes of this argument.
However, it's not that simple when it comes to the morality of the actions since there are secondary effects.
You can't just make a blanket statement like "stealing is immoral", for example, when there are circumstances where stealing is absolutely the moral imperative (e.g. the food is held by barons who charge unaffordable prices and the populace is starving to death).
Or, to make an example closer to the topic, what about the stealing of scientific papers? There's a strong arguments to be made that millions of people have been helped through that theft. And yes, the gatekeepers of that knowledge have been harmed.
And it's like that with AA and AI (and their code) as well. People consider the secondary effects and whether or not the greatest good is served.
And, of course, "greater good" is gray.
If you want information to be free? Great. But most of those authors wouldn't be putting the work in to making that information/literature in the first place if they know they aren't going to get paid.
So I would say it is morally acceptable to download stuff older than that. Torrents offer brand new things, so I wouldn't say they are morally good. I don't think they're morally worse than those locking up old media though.
My understanding is that the number of authors that can make a full-time job of writing is a rounding error compared to the population of authors. I believe that people should be paid for their work, but I don't think the current copyright regime is actually very effective at paying authors for their work. So I value copyright enforcement proportionately less based on that observation.
Meanwhile, on the flip side of the coin, copyright rulings are causing companies like Atheropic to destroy books as they scan them, creating a rising sense of panic around knowledge scarcity. This panic is a direct result of the scarcity of the copyright intended to create in the first place. It's entirely artificial.
I've been saying this for 25 years, but copyright is essentially broken.
I don't really agree
Information itself should be free, authors deserve the right to monetise other ways (merch, physical media, exhibitions/shows, etc.) but society should pay artists and creators to do their thing[0]
Not everything needs to be a business, and I think art is one of those things
[0] I don't know much about it but maybe Ireland's Basic Income for the Arts scheme is a model. I think some other European countries do similar-ish things, too
it doesn't seem like it would be high revenue at least
Anna's Archive is very limited in performance if you don't pay.
It also helps that AA has the best UI and search of any shadow library I've seen. Some really established bittorrent sites win out on curation but otherwise AA is top tier, not something I'm accustomed to seeing from a (free, no-signup!) clearnet/direct-download site
As far as GP's question I think you can be in it for the money (though I personally have my doubts that Anna is) and also be a force for moral good as well. Whether the operator's morals perfectly align with all of our ethics is an open question but the room at large generally tends to agree.
Also Torrentfreak's reporting has been extremely solid for an extremely long time so the link itself fits here too
Maybe if youre trying to pull an AI lab and download their full corpus, although I suspect thats not too onerous either. I have never had any issues getting stuff from annas archive for free
It's not a competition, and being a moral saint isn't a good KPI; being a net positive for society is.
Maybe you can own (some) things, you can't own information