You can help Anna's Archive by seeding torrents
annas-archive.org
annas-archive.org
Note: This is 30M+ books, 100M+ papers. Depending on your philosophy and jurisdiction, you might be stealing a few billion $. Or not.
> By seeding these torrents, you help preserve humanity’s knowledge and culture. These torrents represent the vast majority of human knowledge that can be mirrored in bulk.
One specific host I'll mention is vsys.host (UA) (which anna's archive uses too) but they're not going to be the cheapest option.
Edit: I knew Sweden is a piracy hub but I was under the wrong impression regarding how Sweden legally sees piracy. Yes this is where TPB was raided, but also this is where TPB launched, and there seems to be a certain ideology in Sweden that propels sites like TPB. Either way, please stay safe online.
Sci-Hub is down (as in, you can't search/view/download papers from sci-hub.se), AFAIK.
Anna's Archive aggregates into a unified database material from Libgen.rs, Sci-Hub, Libgen.li, Z-Library, Internet Archive Controlled Digital Lending, and DuXiu 读秀. (See: https://annas-archive.org/datasets )
<https://torrentfreak.com/domain-registry-takes-sci-hubs-se-d...>
Z-Library was significantly attacked recently. There was a huge takedown in 2022, and I'm finding reports as of March 2024 as well:
2024: "FBI Carries Out Fresh Round of Z-Library Domain Name Seizures" <https://torrentfreak.com/fbi-carries-out-fresh-round-of-z-li...>
2022: "Z-Library operators arrested, charged with criminal copyright infringement" <https://www.theregister.com/2022/11/18/zlibrary_operators_ar...>
The Z-Library arrests were of Russian nationals, but were arrested in Argentina.
The first Torrentfreak article mentions two other actions in Spring and November 2023.
Sci-Hub has played cat-and-mouse with domains for years, and AFAIU has still withheld posting new scientific articles given pending litigation/prosecution in India.
Another Torrentfreak article from Nov 2023 gives an update on numerous issues with text liberation sites, including Anna's Archive, Sci-Hub, and Z-Library. At the time, the Anna's Archive twitter account had been "wiped out", so much for that platform's "free speech" stance.
<https://torrentfreak.com/copyright-piracy-news-brief-1-extra...>
I am having a hard time reading that claim as anything other than a bad-faith justification. Torrent nodes are not at all a good way to "preserve humanity's knowledge and culture".
EDIT: I'm no longer convinced I'm correct.
> These torrents are not meant for downloading individual books. They are meant for long-term preservation. With these torrents you can set up a full mirror of Anna’s Archive, using our source code and metadata (which can be generated or downloaded as ElasticSearch and MariaDB databases).
I'm not an expert in any of this, but it doesn't strike me as a bad-faith justification of anything.
I was sad to see that happen, but it’s important to be objective and plan future actions accordingly.
(And sure, there’s always the chance that some random person on GitHub just so happens to be named Anna and is an archival enthusiast, but a jury of one’s peers may find that it passes the reasonable doubt threshold.)
My legal troubles with books3 weighed on me pretty heavily, and I wasn’t even the target. Yet. I can only imagine what it feels like to be waiting for an indictment.
There ought to be some sort of protection for preserving books in bulk. No one is going to read two million books. But of course, one could also argue that having a readily available archive is harming the economic profitability of the works, on the basis that content licensing for AI is now a multimillion industry. It’s weird, because it feels like important work, rather than criminal — someone should put into words exactly what the distinction is.
Preservation of works is very obviously not why Anna's Archive is asking for torrent seeders. Seeding is for distribution and availability, not preservation. It would be more honest to say "preservation of ML training fair use access".
Can you elaborate on that? They write:
"These torrents are not meant for downloading individual books. They are meant for long-term preservation. With these torrents you can set up a full mirror of Anna’s Archive, using our source code and metadata (which can be generated or downloaded as ElasticSearch and MariaDB databases)."
Even such half-aborted systems such as Hathi Trust permit downloads only by the page, even from out-of-copyright works, which is absolutely infuriating. Full access to the Hathi archives (that is, in-copyright works, which cannot otherwise be viewed at all) is restricted to college libraries only, not public libraries generally.
The Internet Archive, LibGen, Z-Library, Sci-Hub, Anna's Archive, and other similar efforts are really the only viable means of access to much of the world's published information, whether inside, outside, or straddling extant copyright law. Which of course was written and lobbied for by existing copyright holders.
I'd much rather we burnt down that law than our true libraries.
As to Anna's Archive and the question of preservation vs. access: AFAIU the full archive is already available and stored in multiple fully-independent copies. As such it will all but assuredly survive even extreme legal attacks, let alone other threats. Access to those works is what torrent seeding provides, and as a means of making the archive available and useful is key to its function and service, but (probably) not its survival.
>It’s weird, because it feels like important work, rather than criminal — someone should put into words exactly what the distinction is.
Important work can be criminal
Yeah, illegal/criminal merely means that it's against the law and will be prosecuted if it can be pinned to you.
As an outlandishly severe example: helping Jews in Nazi Germany was criminal.
Edit: I'm asking as I find no mention of this elsewhere. Notably not on Torrentfreak, which tends to keep on top of such things, on Anna's subreddit, blog, or her Telegram channel
> On information and belief, in addition to her extensive online presence, she has a GitHub (a software code hosting platform) account called, "anarchivist," and she developed a repository for a python module for interacting with OCLC's WorldCat® Affiliate web services.
I think the civil suit will probably be dismissed due to lack of evidence, but it’s likely the police have probable cause to start a surveillance warrant. At that point, either she’ll stop all activity (which it doesn’t seem like she’s doing) or it’s inevitable they’ll catch her in the act.
The slim hope is that maybe it’s not her, or that police won’t pour a lot of resources into surveilling her. Maybe she’ll get lucky. But given the aggressive ways law enforcement has gone after people for e.g. scraping AT&T’s website, I wouldn’t bet on it.
Here's the 2-month-old discussion on the Torrentfreak article:
<https://news.ycombinator.com/item?id=40143549>
And that seems to be the only HN submission with > 5 points that matches this description since 1-Apr-2024:
<https://hn.algolia.com/?dateEnd=1718327852&dateRange=custom&...>
A similar search for "OCLC" turns up nothing.
IIRC, libgen used IPFS for preservation efforts.
Anna's Archive (seemingly the successor) appears to have migrated to BitTorrent.
I wonder what motivated the move?
Edit: asking as someone who works daily on building p2p software. We've abandoned mainline BitSwap (IPFS) in our work for similar reasons as the rest of the rust-libp2p community, but haven't found a particularly good "successor" protocol for a generalized use case yet. We are currently using our own ad-hoc hand-rolled chunking/transfer protocol as needed.
> Starting with v0.3.0, Iroh is a ground-up reimagination of the InterPlanetary File System (IPFS) focused on performance.
Also see, A New Direction for Iroh[1].
Many bittorrent clients let you click a button to continue seeding the data over time.
Libgen still uses torrents primarily for preservation. It also hosts on IPFS but that is more for access, and there are very few IPFS seeders.
We tried IPFS for a bit but found it not stable and usable enough for preservation purposes. We're closely watching IPFS development and hope that it will get there, since it would be wonderful to merge the preservation and access use cases in one system.
I've found the BitTorrent protocol tries to be more suited to accessing popular data on-demand (i.e. streaming a popular file) vs. archival.
IPFS' BitSwap protocol strikes me as trying to be optimized for longer-term preservation (higher latency time to first byte in exchange for more resilient pinning/discovery/propagation of rare data).
It's cool you're observing the opposite. I've had a growing suspicion that both protocols haven't quite realized the benefits they were hoping to get from the trade-offs they made in their transfer/discovery protocols.
Would love to compare notes at some point if you'd be open to it.
We've been playing around with both BitTorrent and IPFS. Some of the datasets we are working towards supporting are approaching the scale you work at (100TB archives).
Ultimately both BitTorrent and IPFS have fallen short for me when trying to seed 100TB datasets.
I've got a hunch that we're going to need to roll a new protocol to tackle these larger datasets that merges some of HTTP's, BitTorrent's, and IPFS' approaches to sharing content.
I have personal R&D list for pushing a file sharing protocol past the 100TB limit (not in any particular order):
* Better chunking using a mix of:
* Rolling hashes
* File boundary splitting
* (should enable deduplication of identical files across nonhomogeneous archives, and allow for adding content to an archive with without losing the existing seeders)
* (inspired by prior art in container storage: https://github.com/hinshun/ipcs)
* "online" deterministic archive formats w/ detached metadata * Ability to share a directory as an archive, or partial slice of an archive, without having to generate the archive on disk. (Announce a "tarball" like archive on the DHT without having to generate it by being able to generate the "chunks" on demand from the directory)
* Detatch the manifest containing the archive's contents from the archive, so you can download/parse the manifest without downloading the full dataset. (You can use this to find the chunks specific files are in. So you can download a single file from a 1TB archive, and the client can seed that file back to the network as part of the archive.)
* Chunking of manifest files for large datasets, since the manifest itself might grow to many GBs in size (manifest resolution inspired by IPLD's data structure)
* Normalize file metadata in the archive header so timestamps etc. don't muck up your CIDs
* Deterministic ordering of files in an archive
* Chunking/Transfers/Announcing/Discovering * Supporting increasing the chunk sizes for large files past 1MB. A 100TB dataset w/ 1MB chunks requires ~209M CIDs just for the chunks, that's a lot of load on the DHT and a lot of work on the seeding node to keep the data available.
* Support interruptible/resumable/recoverable downloads from peers using something similar HTTP RANGE header semantics
* Merge BitTorrent's DHT query approach w/ IPFS' DHT query approach, asking connected peers for CIDs and tit-for-tat reciprocity while simultaneously hedging your bet by kicking off the slower DHT traversal to find more peers
* Connectivity * Bringing mobile devices and browser tabs into the fold as first class peers that can both download and seed content
* (i.e. WebRTC: https://github.com/libp2p/rust-libp2p/tree/master/examples/browser-webrtc)
* (proof-of-concept NAT hole punching appliance for end-users: https://github.com/retrohacker/turn-it-up)
Thank you for everything you doI want to give 5GB (don't have much storage), so I put "Max TB:0.005" and "Type:URLs".
It give me this url:
https://annas-archive.org/dyn/generate_torrents?max_tb=0.005...
Who has this two torrent files:
https://annas-archive.org/dyn/small_file/torrents/external/l...
https://annas-archive.org/dyn/small_file/torrents/external/l...
I put that torrents on Transmission, one is 5GB and the other 4MB, the 5GB is not downloading/seeding, the 4MB was downloaded and is seeding:
Any help?
We're soon releasing another few hundred TB of completely unique materials (mostly books, also lots of magazines), so help with preserving all of this is sorely needed.
Much, much thanks to everyone who has contributed in one way or another already!
EDIT: Thank you to everyone for your recommendations. I shall find an anon seedbox. I don't mind if it's nuked. I imagine I'll just pay per month and losing one month's spend won't hurt.
Though it does strike me now that any such service will be primarily used by criminals.
Just think of them as "anti-state actors"
Getting a small VPS with crypto and tunneling from your home seems like a better pathway to me?
https://kycnot.me/?t=service (not just hostings)
https://bitcoin-vps.com/ (an extensive list for ones that accept BTC, most accept other coins too)
There are hundreds of VPS hosters that accept crypto, but the important part is that a lot of them are not happy about abuse reports either way, so you'll probably have to use a VPN (like Mullvad) on the VPS itself to not get it suspended :)
Alternatively, there's /r/seedboxes over on reddit where most vendors accept bitcoin and get you a complete setup.
They're not a VPS in the traditional sense, but they give you a slice of a server, with a torrent client pre-installed.
Perhaps Great Depression 2.0 will set our priorities straight.
I'm a happy VPN customer, they've been excellent.
Given how much data Cloudflare (and other similar giants providing this service) has, this pretty much identifies the person to them.
Any plans to change that to a more privacy-protecting solution?
To permanently seed some of these torrents I'd have to keep it on all the time. Would have to keep my desktop running 24/7 too?
Besides, gluetun+chihaya+qbit containers do the job without breaking a sweat, and without ever having to remember that you run a VPN - as it’d only be tunneling the containers of your choice. gluetun is the best image ever made!
This is extremely suboptimal. "I only hide my activity when it's worth hiding" paints a giant target on your back. VPNs as a matter of course protect you from all manner of anti-consumer tactics, and if you don't obfuscate all of your traffic, it tips off surveilling parties to only focus on the subset of traffic that routes through a VPN.
I'm definitely already on a health insurance bad persons list. My employer forwarded myself and other employees terminated over "that thing" to the FBI!
Thanks for the clarification of your stance on VPNs and privacy, I appreciate the insight into why one should care.
As for the VPN side you should be able to configure your VPN to always tunnel your torrent app but not always tunnel your entire computer. The best/easiest way to do this varies by the specific VPN application and your OS.
You could build a seed box out of a old ARM board running the Transmission daemon and a USB key mounted read only to avoid wear; power draw would be just a few watts and total cost could be less than 50 bucks. The desktop would be needed only when adding torrents or changing configuration from its web interface, although Transmission also has remote control apps running on phones and tablets. If the router permits it, QoS rules can be set up on the router so that the seed box can use all bandwidth, although at lower priority than other machines on the LAN, so that it will never clog the network, which comes handy for example with online gaming.
E.g. Consider someone with a gigabit connection that their VPN can't keep up with. It would make sense to seed the legal items without the VPN, and the other items with the VPN.
edit: this is maybe the link he's after? looks like you need to be logged in to see it though.
The second reason is that libraries tend to operate under government control, and governments has done things to enable libraries and work around copyright law. An old one that my country (used to?) have was to require publishers to send copies to the national library. In return, national authors got a symbolic sum (very tiny) each time a copy was taken out. Being forced to send a copy to the government isn't technically against copyright law, since no unlawful copying is being made, but the result has a very similar feeling as unlawful copying.
That being said while the IPFS protocol is decent the implementations kind of suck. Bittorrent is well established with many high quality implementations.
Really, BitTorrent could do this by making all torrent files a small fixed size and then having "torrent files of a directory of torrent files" where the torrent client knows to queue the sub-torrents as they're discovered+downloaded in the parent torrent. But that's not how any part of the ecosystem works. IPFS is a "do over" that allowed them to fix this.
BitTorrent v2 would in theory be able to seed individual files even if they come from a different torrent. But clients have no reasonable way to look for other versions of a torrent that contain a file they already have.
The main Bittorrent clients already support creating and seeding v2 torrent. But there's just no infrastructure for seeding at the individual file level.
If you don't distribute further, it should be pretty ok to download and play with this?
I would if I could though, so excited to see anyone answer the call!