They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to deal with that for the rest of the archive so hopefully that's not an extra complexity for them.
They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to deal with that for the rest of the archive so hopefully that's not an extra complexity for them.
For my needs, I run a report monthly for the content I’ve archived using my IA account to determine archived GBs, and then donate the amount needed to cover those costs.
Consider reaching out to their patron services email address with any questions.
Edit: $2/GB citation: https://help.archive.org/hc/en-us/articles/360014755952-Arch...
[1] https://help.archive.org/hc/en-us/articles/360014755952-Arch...
Do you have source/more info than that?
Lets say the internet archive is 100 PB [1], that's 100,000,000 GB [2], and at that rate it comes out to $200 million [3] for the whole thing forever. That's a lot of money, but also a lot less than I was expecting for something like that.
[1] https://www.protocol.com/internet-archive-preserving-future: "The web archive alone is about 45 petabytes — 4,500 terabytes — and the Internet Archive itself is about double that size (the group has other collections, like a huge database of educational films, music and even long-gone software programs)."
[2] https://www.google.com/search?q=100+petabytes+to+gb: "100 petabyte = 1e+8 gigabytes"
[3] https://www.google.com/search?q=1e%2B8+*+%242: "1e+8 * (US$ 2) = 200 million US$"
I suppose the claim is rather shocking and warrants citations - $200m to host the entire internet archive forever? I don't blame you for the excessive citation.
With a 9% discount rate, that’s only $3.63 million dollars in present value to pay for infinity months.
Of course there are other costs (and cheaper more efficient disks; and cheaper power; and your discount might be less aggressive; and server aren’t free tho you only need like 1 server for 100 disks with SAS expanders since most data is never read; and maintenance), but the $200 million number you got seems very reasonable to me.
Edit: I'm assuming they can deliver reliability and durability similar to modern cloud standards, like AWS S3.
Suffice it to say, physics is probably going to have a lot to say about that assumption in the coming decades.
I guess archiving a static blog is already trivial for their system, but I'd pay $100 a year to have IA host my static blog. The overhead I would consider as a donation to a worthy cause.
Then again, I can just give them $100 a year and find some free static hosting, like GitHub Pages, and call it a day.
I understand they've tried this or things like it a few times but they haven't ever kept the feature.
The Wayback Machine's data is ~20PB..? What is approximately the size of the indexable text (i.e. the text content of html pages, sans tags)? And what would the index size be like, approximately?
I imagine that creating (and maintaining, of course) the index would be the most time-consuming part? Is it at all possible to imagine hosting this index... somewhere... and doing sqlite http range-like queries on it..?
Would it be enough to have an index consist of a list of found words, and the related "document ids"? i.e. "apple" is in doc ids 1000, 2000, 3000, "banana" is in doc ids 2000, 4000, etc.?
And have separate docid -> archive.org url mapping?
The problem, as with any search engine, is the ranking algorithm. Without a sufficent one, the search results are useless. What use is a list of every page in the Wayback Machine containing the word "apple"?
The Wayback Machine possibly would need a much larger index than any normal search engine: not only the present websites, but all the historical versions (though I don't know what proportion of the web they've indexed).
I was very naively "back of the envelope" prototyping a search engine. I both realize that this is not the way that these things are built, and both would really like to have that (or any) search engine to look through the archive..! :-)
As for the index, I agree -- I was trying to guesstimate its size by going from the ~~20PB total Wayback Machine size (which includes all historical versions). Is it 1% of 20PB (for the size of the text content), and then another ~10% of that for the index size? So 20TB...?
> I was imagining that a full text Wayback Machine search engine would mostly be useful to look for words unique enough that sifting through a lot of results (even if those were not ranked "well") could still be useful..?
If we think about use cases, users may often search specific domains. In that case, results ordered by frequency and/or date might be sufficient and even desireable.
> I was very naively "back of the envelope" prototyping a search engine. I both realize that this is not the way that these things are built
It's often the first step!
> So 20TB...?
That doesn't sound so bad.
I have 0 time for this, but I also can't easily let go ;-)
Want to collab on this?
You made my day.
It wouldn't be hard to maintain a hand-crafted database of when domains are reused for something completely different, or even when the same conceptual website has breakages, and use that to choose between Internet Archive or live web accordingly. When one is browsing from an internet archive page, the date is known, when someone is browsing from a live website, heuristics can be used, along with "bisecting" dates when the link is dead.
Ultimately we want more content addressing to avoid this problem entirely (see below), Or DNS -> PubKey, PubKey -> latest content, with some law that the pubkeys shall not be reused for unrelated things. vs DNS which is mere ephemeral Huffman encoding. So see below for the stuff on IPFS. But the trick above is a good stop-gap, and indeed the database itself used to back the extension could be on IPFS.
Alternative content addressed systems may also work better for finding the content than someone hosting your DNS records for a long time but the bulk of the problem space is in guaranteeing active hosting in a way viewable to viewers of the age will be available for many years not addressing the content. On top of hosting and addressing the Internet Archive offers the ability to view old content on modern browsers even if modern browsers have 0 support for such content anymore (or if browsers ever had support at all even). Forward compatibility isn't something solved by a protocol.
This is rather pessimistic thinking. The same money that goes into buying up domains could go into lobbying browsers to add this functionality be default.
> bulk of the problem space is in guaranteeing active hosting in a way viewable to viewers of the age will be available for many years not addressing the content.
You're moving the goal posts. I am not saying "IPFS means we don't need the internet archive". We absolute do need the internet archive. Content addressing helps by making archiving transparent, so the archival copy is not worse than the original.
Fundamentally, consumers producers or archivists may be the party most interested in the continued existence of some information at different moments in the lifespan of that information. Location-based addressing forces the producers to shoulder the burden of hosting, but content-based addressing allows the work to be distributed among those 3 however we see fit. Of course the burdened must still be barred! That doesn't mean the flexibility isn't extremely useful.
> On top of hosting and addressing the Internet Archive offers the ability to view old content on modern browsers even if modern browsers have 0 support for such content anymore (or if browsers ever had support at all even). Forward compatibility isn't something solved by a protocol.
Yeah that's great too, and again not something I am arguing is not good, or not necessary.
Really the way I see it outside a few large banking firms, its kind of hard to be sure any provider of digital services would be around in the 50+ year term for this kind of public archive.
I hope the Internet Archive manages it.
EDIT: I do worry the IA has a bit of a lightning rod effect with skirting issues re: legality of archiving content. IMO its no guarantee it survives any significant time span either.
A library could do it. Perhaps leading institutions like the British Library or Library of Congress. I've thought that IA should be a Library of Congress project, and may eventually end up under their auspices.
They drop stuff that has the least amount of accessing, so if you want to keep something online you have to pay someone to keep accessing it. Makes sense, but I'd argue that it should be a market, meaning the price should go up as more people try to access stuff, but anyone can start to seed popular content and collect revenues also for hosting it (this is better than wasting electricity on accessing stuff or doing proof of work).
But isn't FileCoin exactly that for IPFS?
MaidSAFE goes a step further and has nodes rebalance autonomously and earn the most safecoin, as something gets more popular it gets seeded more.
The feature I most want from IA is a streamlined system to delete content they have archived on domains that I own, including a proper privacy law compliance effort on their part. They have intentionally made it a difficult, manual process to get content removed. They operate as a de facto malicious crawler.
They massively violate GDPR with how they operate and few seem to care about that fact, including all the commenters on HN (which universally give them a free pass on being malicious and violating GDPR very aggressively).
When IA has to comply with laws like GDPR, that's the end of IA.
If you had to choose between the GDPR, & an accurate historical record, which would you prefer?
An author suggesting that the LoC remove their copy of a book/other work (including digital works) because they want to unpublish it would not fly.
The parent comment has an issue with anything on the Web not hidden in some way being considered 'public' and 'published', but that would be something that would require international cooperation to hash out.
Will you be happy when you've burned down that library?
There are very good reasons that archives will not destroy or alter information outside of very clear difficult and manual processes.
And actually, looking at it, I don't think they're necessarily in violation of GDPR [0].
Point 3 says: "Where personal data are processed for archiving purposes in the public interest, Union or Member State law may provide for derogations from the rights referred to in Articles 15, 16, 18, 19, 20 and 21 subject to the conditions and safeguards referred to in paragraph 1 of this Article in so far as such rights are likely to render impossible or seriously impair the achievement of the specific purposes, and such derogations are necessary for the fulfilment of those purposes."
According to GDPR, national law of EU parties overrules GDPR when it comes to personal data being used in archival context. I don't know every EU country's stance, but most of the bigger economies would allow for this.
There is also a difference between deleting the data and rendering it inaccessible to the public. Keeping something under wraps is generally more 'acceptable', but active destruction of the item (digital or not) and its providence is much more limited. Also there's a difference between personally identifying data (covered by GDPR), your content (which would be covered under copyright and not GDPR), and connections people can make if that content is available (not covered at all because it's not anybody else's issue if you write something terrible and people keep recognizing you over it so long as you did actually write it).
> Article 3(2), a new feature of the GDPR, creates extraterritorial jurisdiction over companies that have nothing but an internet presence in the EU and offer goods or services to EU residents[1]. While the GDPR requires these companies[2] to follow its data processing rules, it leaves the question of enforcement unanswered. Regulations that cannot be enforced do little to protect the personal data of EU citizens.
> This article discusses how U.S. law affects the enforcement of Article 3(2). In reality, enforcing the GDPR on U.S. companies may be almost impossible. First, the U.S. prohibits enforcing of foreign-country fines. Thus, the EU enforcement power of fines for noncompliance is negligible. Second, enforcing the GDPR through the designated representative can be easily circumvented. Finally, a private lawsuit brought by in the EU may be impossible to enforce under U.S. law.
[snip]
> Currently, there is a hole in the GDPR wall that protects European Union personal data. Even with extraterritorial jurisdiction over U.S. companies with only an internet presence in the EU, the GDPR gives little in the way of tools to enforce it. Fines from supervisory authorities would be stopped by the prohibition on enforcing foreign fines. The company can evade enforcement through a representative simply by not designating one. Finally, private actions may be stalled on issues of personal jurisdiction. If a U.S. company completely disregards the GDPR while targeting customers in the EU, it can use the personal data of EU citizens without much fear of the consequences. While the extraterritorial jurisdiction created by Article 3(2) may have seemed like a good way to solve the problem of foreign companies who do not have a physical presence in the EU, it turns out to be practically useless.