ArchiveTeam Warrior
warrior.archiveteam.org
warrior.archiveteam.org
For example: I don't think that Reddit ever felt threatened by the fact that a bunch of people are pulling their data and creating a backup that can not be easily accessible. But I think that Reddit would be very much afraid of a distributed network of nodes running partial copies of Reddit and making it available for local-first clients and/or Lemmy mirrors.
(Note that Archive Team is separate from Internet Archive)
edit: they already did.
I have a somewhat of a conspiracy theory that deep down they don't implement a search feature on all their content on purpose. Essentially if WB made easy to discover stuff, you end up having to deal more and more with all those shenanigans of people requesting information to be removed. By making the information there, but somewhat unfindable or at least very hard to find, they essentially preserve the information, without having to deal with such problem (I know this happens even nowadays, but if it was easier to find information, it would happen even more).
On related note, Internet Archive backup would probably cost between 20M and 60M USD. Many EU countries would have incentive to do this as public culture preservation projects.
The archive is roughly 70 PB. Decentralized storage projects have achieved 7PB already. Thus attempting a decentralized backup would ALSO work.
There are multiple archive projects around the world, which often isn't understood when the Internet Archive (the biggest and original) is discussed.
Various countries including some EU already have those as "public culture preservation projects", targeting their own nation's web presence.
In that context, and with scarce funding already, there is not really an incentive to back up a load of irrelevant (in the sense it's not their country's) archive material.
Now, notice that the budget of many of these projects are X billions:
https://www.ne-mo.org/cooperation-funding/funding-opportunit...
Putting a 100 million EUR into an European IA backup would be more cost effective than any of these projects.
Alternatively or additionally:
https://en.m.wikipedia.org/wiki/Wikipedia:Fundraising_statis...
Wikipedia could actually also probably – being dependent on IA – invest 50M into the project. In fact, this would probably do more what the donations were meant to do than anything else they could do with the ("excess") funds.
Truth to be spoken, NSA probably has an IA backup. But it still sort of drives me insane to know that political change or natural catastrophes could lead to loss of public access to the IA. No one seems to care about IA enough except IA itself.
Current digital preservation projects are likely a tiny fraction of that 100 million on national levels and will include additional activities like those you ascribe to the IA, but carefully attuned to each nation's priorities while collaborating internationally with each other including organisations in the USA like the IA.
Importantly, they will also be operating strictly according to national and intra-national legislation (which IA has gotten into severe trouble within the USA).
In the context of a complicated international environment with many different local political, cultural, commercial and other factors, it's difficult to see how your proposal to replace local projects with an IA backup would be either more cost effective or legally practical.
The Internet Archive is inspirational and does a terrific job, but canning the many disparate entities that do national equivalents on much more limited budgets in favour of moving those funds to the IA (or any other global corporation or organisation) would risk invoking the classic problems of centralisation with associated detrimental effects on local requirements.
I am not sure if you opened the link I gave. It seems the budget in EU is particularly high.
I don't believe the optimal solution would be to move anything to IA: in fact, a separate legal entity would be much better option due to decreased legal risks. This only needs to be updated perhaps once or twice a decade, or even less.
[1] https://www.pewresearch.org/data-labs/2024/05/17/when-online...
Let me explain where I am coming from. I'm working on an (open source, self-hosted) service to help people migrate from Reddit to Lemmy, called Fediverser [0]. It offers the following:
- A crowdsourced map of "reddit-to-lemmy" alternatives. - Use the list of subreddits and some preferences to find a Lemmy instance that is suitable for you.
- Lets people sign up to a "fediversed" Lemmy instance directly via Reddit OAuth. Simplifies the registration process, can let an admin skip the verification process (e.g, reject redditors whose account are less than 3 years old)
- Using the crowdsourced data, automatically subscribe the user to the communities that correspond to their favorite subreddits.
- If the admin of the Lemmy instance so chooses, it can also set up mirror bots, which will create "shadow accounts" for each reddit author. This shadow account can then be "taken over" by the real redditor if/when they sign up to the instance.
I believe that these features together would lead to a credible threat to Reddit's dominance. My remaining "problem" to solve is, simply put, that I need more people running this, because it's just too much data for a single node. I set up an instance that was mirroring ~100 reddits (posts and comments). In three months, my database was already recording ~3 million "shadow" users and ~10 million posts + submissions.
For this to work, I either need to have more instance admins willing to run the Fediverser software, or I need to move the "shadow" users and the mirrored content straight to the client and only bring to the Lemmy server the content from users who actually migrated.
1) A lot of Reddit's usefulness comes from the bots. If they shut down the API entirely, they would lose a lot of value which would accelerate their demise.
2) A more cynical person would say that without bots, Reddit would lose 30-40% of its "traffic", and that they can not afford to do that. This is why the current API is still quite generous. It's enough for most bots, but just too expensive for third-party clients.
3) Even if Reddit shut down its API tomorrow, the majority of interesting content has already been copied/archived.
4) Scraping old.reddit is quite easy, and getting rid of it altogether is not something that they are willing to do.
All in all, I'd say that if we ever get to a point where Reddit is cutting down the API, it will be a time to celebrate.
Really? Unless there's an ironclad public statement to the contrary, the vibe I get from old.reddit is that it won't be long until the axe falls, especially since they are fine with "new" reddit features that malfunction in old.reddit. And I'd imagine their attitude towards unhappy users once they kill old.reddit will be the same as when they killed the API (which was "go fuck yourselves", to put it plainly).
My feeling is that Reddit will continue to push and nudge people to new UI, but will not fully retire old. I think that if they ever do it it will be the final straw and mass migration will be inevitable.
They certainly don't want to have other Big Tech companies exploring "their" data, but I'd argue they would only after someone who had a clear commercial interest. They are not going to chase a few dozen people on /r/DataHoarder running an Archive Warrior.
Reddit's become low quality now anyway. All the good old high quality information that people search Reddit for is in the dumps - they're not generating much more of it.
I can dream.
ArchiveTeam Warrior: archiving as much of imgur as possible - https://news.ycombinator.com/item?id=35983510 - May 2023 (2 comments)
Help preserve the internet with Archiveteam's warrior - https://news.ycombinator.com/item?id=30524842 - March 2022 (51 comments)
ArchiveTeam Warrior backing up Reddit - https://news.ycombinator.com/item?id=29584622 - Dec 2021 (71 comments)
https://github.com/ArchiveTeam/warrior-dockerfile/tree/maste...
https://github.com/ArchiveTeam/warrior-dockerfile/blob/maste...
I wish I had spare time to try to figure out how to get it set up inside QEMU on a Raspberry Pi... Seems like publishing such a compose file would unlock it for even more people.
There's a specific "wget-lua.raspberry" file in the repo, so Raspberry Pi seems to be supported almost natively.
You can get Docker to emulate x86, though.
It also comes with a retrieval function, where you can say "I want to get X content from this network" and it will find it for you.
It's a fairly thin wrapper over torrents, but it something that doesn't currently exist. Unfortunately, it wasn't met with much interest when I contacted some archivists.
looks like it's currently broken :(
They need somewhere to download things before pushing them to archive.org
For HTTP this seems impossible. For HTTPS, who does the TLS termination? It could be safe if the warrior was just at TCP proxy and Archive was the TLS client, but https://wiki.archiveteam.org/index.php/ArchiveTeam_Warrior doesn't seem to explain that clearly.
It doesn't, but in practice, it's not a problem. The goal is to with minimum turnaround rescue as much data as possible before a site goes down. Getting 1% garbage is more preferable than losing 30% to inefficiency in archiving because you're paranoid.
> For HTTPS, who does the TLS termination?
The Warrior literally just runs wget and saves to warc files that are uploaded to a tracker server.
However to have it more tightly integrated, you'd want an arm image at /, mount the original warrior image at /warrior, then you could do something like `qemu-user-amd64 /warrior/bin/chroot /warrior/entrypoint` - Updating all paths as relevant
Which then contains a README with a link to the source for the new version: https://github.com/ArchiveTeam/warrior4-vm
This new version is also mentioned on the homepage of the Wiki: An updated Warrior virtual appliance (v3.2, v4.0) is now available with better support for newer projects that utilize wget-at.
ArchiveTeam should update the appliance download URL.
https://github.com/ArchiveTeam/warrior-dockerfile makes it pretty easy to setup.
Right now I run three instances. Pretty low resource utilization and they're totally segregated into their own instances so just boot them and they run, shut them down and they stop.
Given they're running arbitrary external commands I wanted them kept on their own machines as much as possible.
Once in a while, a site will tolerate a large number of connections, and since the Warrior VM only supports a concurrency of 6, it can make sense to run multiple warriors on the same project. But this is almost always a bad idea, and many sites will 429 you with a concurrency of anything more than 2 or 3, always check the project-specific IRC channel for concurrency recommendations.