In the end, I've ended up just keeping my own ArchiveBox and it's an all right experience. In the end, it's only useful for things I know I wanted to archive. For almost everything I go to the IA - which has so much.
In the end, I've ended up just keeping my own ArchiveBox and it's an all right experience. In the end, it's only useful for things I know I wanted to archive. For almost everything I go to the IA - which has so much.
It's been dormant / on hiatus for a few years now.
- I think I have seen that AI scrapers create bottleneck in the bandwidth
- To some digital archives you need to create scientific accounts (I think Common Crawl works like that)
- Data quite easily can be very big. The goal is to store many things. We not only store Internet, but with additional dimension of time
- Since there is a lot of data, it is difficult to navigate it, search it, so it easily can become unusable
- For example that is why I created my own meta data link, I needed some information about domains
Link:
Edit:
Would be really neat if you could click on a domain while on IA, and a desktop client downloads as many WAR files in a slower priority download queue, as many as you're interested in, with higher priority pages first, and then you can view it fully offline.
No one uses IPFS. For the average user, it is significantly more difficult to get started. For the experienced user, the ecosystem of tools around IPFS is extremely small.
All in all, IPFS offers very little benefit over torrents in practice and has a much smaller user pool.
https://www.bittorrent.org/beps/bep_0039.html
https://www.bittorrent.org/beps/bep_0046.html
If that updated torrent is a BEP-0052 (v2) torrent it will hash per-file, and so the updated v2 torrent will have identical hashes for files which aren't changed: https://www.bittorrent.org/beps/bep_0052.html
This combines with BEP-0038 so the updated torrent can refer to the infohash of the older torrents with which it shares files, so if you already have an old one you only have to download files that have changed: https://www.bittorrent.org/beps/bep_0038.html
I wish I could find that article!
edit: https://github.com/internetarchive/dweb-archive/blob/master/...
It's based on torrents, and you can easily make a content delivery system on top of this (so people can fetch data from this network).
I emailed a few archiving teams but nobody seemed interested, so I never made it.
Easier to send fiat to IA for them to invest (~$2/GB) and to pay to keep the disks spinning somewhere safe across the world.
(ia volunteer, no affiliation otherwise)
You're right, though, long-term commitment is rare from volunteers. That's why the idea is to make short-term commitment so easy that you have a good enough pool of short-termers that it works out in the aggregate.
(ia source of truth, storage system of last resort -> item index -> torrent index -> global torrent swarm)
https://gist.github.com/skorokithakis/68984ef699437c5129660d...
https://annas-archive.org/torrents
I think I'm misunderstanding you.
But torrent is probably the wrong tech. I’m sure there would be many players willing to host a few TB or more each, which could be fronted via something so it’s transparent to the user.
But a better option might be a subscription model, anything else will be slammed by crawlers.
https://www.bittorrent.org/beps/bep_0039.html https://www.bittorrent.org/beps/bep_0046.html
but unfortunately most foss torrent clients do not support it, partly because at release libtorrent 2.0.x had poor io performance in some cases so torrent clients reverted to the 1.2.x branch