(Note that Archive Team is separate from Internet Archive)
(Note that Archive Team is separate from Internet Archive)
Let me explain where I am coming from. I'm working on an (open source, self-hosted) service to help people migrate from Reddit to Lemmy, called Fediverser [0]. It offers the following:
- A crowdsourced map of "reddit-to-lemmy" alternatives. - Use the list of subreddits and some preferences to find a Lemmy instance that is suitable for you.
- Lets people sign up to a "fediversed" Lemmy instance directly via Reddit OAuth. Simplifies the registration process, can let an admin skip the verification process (e.g, reject redditors whose account are less than 3 years old)
- Using the crowdsourced data, automatically subscribe the user to the communities that correspond to their favorite subreddits.
- If the admin of the Lemmy instance so chooses, it can also set up mirror bots, which will create "shadow accounts" for each reddit author. This shadow account can then be "taken over" by the real redditor if/when they sign up to the instance.
I believe that these features together would lead to a credible threat to Reddit's dominance. My remaining "problem" to solve is, simply put, that I need more people running this, because it's just too much data for a single node. I set up an instance that was mirroring ~100 reddits (posts and comments). In three months, my database was already recording ~3 million "shadow" users and ~10 million posts + submissions.
For this to work, I either need to have more instance admins willing to run the Fediverser software, or I need to move the "shadow" users and the mirrored content straight to the client and only bring to the Lemmy server the content from users who actually migrated.
edit: they already did.
I have a somewhat of a conspiracy theory that deep down they don't implement a search feature on all their content on purpose. Essentially if WB made easy to discover stuff, you end up having to deal more and more with all those shenanigans of people requesting information to be removed. By making the information there, but somewhat unfindable or at least very hard to find, they essentially preserve the information, without having to deal with such problem (I know this happens even nowadays, but if it was easier to find information, it would happen even more).
On related note, Internet Archive backup would probably cost between 20M and 60M USD. Many EU countries would have incentive to do this as public culture preservation projects.
The archive is roughly 70 PB. Decentralized storage projects have achieved 7PB already. Thus attempting a decentralized backup would ALSO work.
There are multiple archive projects around the world, which often isn't understood when the Internet Archive (the biggest and original) is discussed.
Various countries including some EU already have those as "public culture preservation projects", targeting their own nation's web presence.
In that context, and with scarce funding already, there is not really an incentive to back up a load of irrelevant (in the sense it's not their country's) archive material.
Now, notice that the budget of many of these projects are X billions:
https://www.ne-mo.org/cooperation-funding/funding-opportunit...
Putting a 100 million EUR into an European IA backup would be more cost effective than any of these projects.
Alternatively or additionally:
https://en.m.wikipedia.org/wiki/Wikipedia:Fundraising_statis...
Wikipedia could actually also probably – being dependent on IA – invest 50M into the project. In fact, this would probably do more what the donations were meant to do than anything else they could do with the ("excess") funds.
Truth to be spoken, NSA probably has an IA backup. But it still sort of drives me insane to know that political change or natural catastrophes could lead to loss of public access to the IA. No one seems to care about IA enough except IA itself.
Current digital preservation projects are likely a tiny fraction of that 100 million on national levels and will include additional activities like those you ascribe to the IA, but carefully attuned to each nation's priorities while collaborating internationally with each other including organisations in the USA like the IA.
Importantly, they will also be operating strictly according to national and intra-national legislation (which IA has gotten into severe trouble within the USA).
In the context of a complicated international environment with many different local political, cultural, commercial and other factors, it's difficult to see how your proposal to replace local projects with an IA backup would be either more cost effective or legally practical.
The Internet Archive is inspirational and does a terrific job, but canning the many disparate entities that do national equivalents on much more limited budgets in favour of moving those funds to the IA (or any other global corporation or organisation) would risk invoking the classic problems of centralisation with associated detrimental effects on local requirements.
I am not sure if you opened the link I gave. It seems the budget in EU is particularly high.
I don't believe the optimal solution would be to move anything to IA: in fact, a separate legal entity would be much better option due to decreased legal risks. This only needs to be updated perhaps once or twice a decade, or even less.
[1] https://www.pewresearch.org/data-labs/2024/05/17/when-online...