Internet Archive as a default host-of-record for startups
twitter.com
twitter.com
I manage the Wayback Machine at the Internet Archive.
Very happy so many people here care about preserving, and making available, our cultural heritage!
Please know a dedicated, and talented, team of engineers works every day to do a better job of archiving more of the public Web, and making it available via the Wayback Machine.
As noted the Internet Archive is experimenting with filecoin.io and storj.io and is always open to suggestions about how we might do our jobs better, and improve our service. We also host regular meetups (and have hosted summits and a camp) related to the Decentralized Web. See: https://blog.archive.org/tag/dweb/
The Internet Archive also offers archive-it.org, a subscription service, for those who want a higher level of support and more features.
We appreciate any support you can offer, financial and otherwise. Please share any bug reports, feature suggestions and other feedback with us via email to info@archive.org
Oh/and… checkout the new PDF Search feature we just launched at the bottom of web.archive.org. More to come like that in 2022.
Finally, you might also find some of the things I wrote here of interest: https://gijn.org/2021/05/05/tips-for-using-the-internet-arch...
That said.. damn I really wish y'all would revisit some of your fundamentals like recursive scraping and making sure your scraping is whole and complete before working on filecoin and other needlessly flashy systems. I'm genuinely worried y'all are digging yourselves into a technical pit that you can't get out of and it will hurt or even kill your goals.
Cheers.
1) can consume a lot of storage really fast
2) makes IA a bot rather than a user service
3) is already done on a case by case basis by ArchiveTeam allowing IA to stay away from the previous two problems.
I realize it's a hard problem, but I really wish there were a way to automate more of it. Some of these communities are too small for anyone to bother preserving the pages manually, and I don't imagine we'd even show up on ArchiveTeam's radar. But they're not large pages, and basically static. I don't think they'd be a huge burden to store and maintain. It seems like some sort of a coverage + size metric would be pretty effective at guiding an automated scan such that you'd be able to preserve things liked this without needing humans to go and manually archive each and every page.
AT just asks to be given a few weeks/months of notice depending on how many TBs (MBs?) the crawl needs to be and not to throttle/ban their clients.
I know about the web.archive.org/https://... "hack" but it's probably draining your servers unnecessarily when there are a lot of 301 redirects that weren't archived and on top of that are only able to be validated client side after receiving the whole HTML response.
A REST API would help thirdparty clients to know about this in advance, and the API documentations I found were super unclear in whether something like this exists or not.
Context: I'm building a web browser and I'm trying to offer a feature for error cases when the server or URL isn't available anymore, so that users can see the web archived version of it.
On top of that I have no idea how to "un-UI" the web archived versions. I know that the wget user agent somehow leads to this, but it's also kind of undocumented how the webserver of IA does this in the background and when exactly the UI is injected and all the URLs are rewritten. Something like maybe a http request header to get the raw actual source would be nice.
Have you looked at the headers that they send?
GET /web/20100330210402/https://arxiv.org/abs/0911.1112 HTTP/2
[...]
HTTP/2 200 OK
[...]
link: <http://arxiv.org/abs/0911.1112>; rel="original",
<https://web.archive.org/web/timemap/link/http://arxiv.org/abs/0911.1112>;
rel="timemap"; type="application/link-format",
<https://web.archive.org/web/http://arxiv.org/abs/0911.1112>;
rel="timegate",
<https://web.archive.org/web/20100330210402/http://arxiv.org/abs/0911.1112>;
rel="first memento"; datetime="Tue, 30 Mar 2010 21:04:02 GMT",
<https://web.archive.org/web/20100330210402/http://arxiv.org/abs/0911.1112>;
rel="memento"; datetime="Tue, 30 Mar 2010 21:04:02 GMT",
<https://web.archive.org/web/20110101203756/http://arxiv.org/abs/0911.1112>;
rel="next memento"; datetime="Sat, 01 Jan 2011 20:37:56 GMT",
<https://web.archive.org/web/20211123040625/https://arxiv.org/abs/0911.1112>;
rel="last memento"; datetime="Tue, 23 Nov 2021 04:06:25 GMT"
[...]
The Wayback Machine implements RFC 7089. <http://mementoweb.org/guide/quick-intro/>Brave browser does that, and I've seen Firefox add-ons that do it. Maybe you can look at their code.
Doesn’t IA already partner with Cloudflare to do exactly what Carmack is suggesting.
https://blog.cloudflare.com/cloudflares-always-online-and-th...
They would also need to be able to handle legal administration costs of things like DMCA take-down notices, but I assume they already have to deal with that for the rest of the archive so hopefully that's not an extra complexity for them.
They drop stuff that has the least amount of accessing, so if you want to keep something online you have to pay someone to keep accessing it. Makes sense, but I'd argue that it should be a market, meaning the price should go up as more people try to access stuff, but anyone can start to seed popular content and collect revenues also for hosting it (this is better than wasting electricity on accessing stuff or doing proof of work).
But isn't FileCoin exactly that for IPFS?
MaidSAFE goes a step further and has nodes rebalance autonomously and earn the most safecoin, as something gets more popular it gets seeded more.
It wouldn't be hard to maintain a hand-crafted database of when domains are reused for something completely different, or even when the same conceptual website has breakages, and use that to choose between Internet Archive or live web accordingly. When one is browsing from an internet archive page, the date is known, when someone is browsing from a live website, heuristics can be used, along with "bisecting" dates when the link is dead.
Ultimately we want more content addressing to avoid this problem entirely (see below), Or DNS -> PubKey, PubKey -> latest content, with some law that the pubkeys shall not be reused for unrelated things. vs DNS which is mere ephemeral Huffman encoding. So see below for the stuff on IPFS. But the trick above is a good stop-gap, and indeed the database itself used to back the extension could be on IPFS.
Alternative content addressed systems may also work better for finding the content than someone hosting your DNS records for a long time but the bulk of the problem space is in guaranteeing active hosting in a way viewable to viewers of the age will be available for many years not addressing the content. On top of hosting and addressing the Internet Archive offers the ability to view old content on modern browsers even if modern browsers have 0 support for such content anymore (or if browsers ever had support at all even). Forward compatibility isn't something solved by a protocol.
This is rather pessimistic thinking. The same money that goes into buying up domains could go into lobbying browsers to add this functionality be default.
> bulk of the problem space is in guaranteeing active hosting in a way viewable to viewers of the age will be available for many years not addressing the content.
You're moving the goal posts. I am not saying "IPFS means we don't need the internet archive". We absolute do need the internet archive. Content addressing helps by making archiving transparent, so the archival copy is not worse than the original.
Fundamentally, consumers producers or archivists may be the party most interested in the continued existence of some information at different moments in the lifespan of that information. Location-based addressing forces the producers to shoulder the burden of hosting, but content-based addressing allows the work to be distributed among those 3 however we see fit. Of course the burdened must still be barred! That doesn't mean the flexibility isn't extremely useful.
> On top of hosting and addressing the Internet Archive offers the ability to view old content on modern browsers even if modern browsers have 0 support for such content anymore (or if browsers ever had support at all even). Forward compatibility isn't something solved by a protocol.
Yeah that's great too, and again not something I am arguing is not good, or not necessary.
I understand they've tried this or things like it a few times but they haven't ever kept the feature.
The Wayback Machine's data is ~20PB..? What is approximately the size of the indexable text (i.e. the text content of html pages, sans tags)? And what would the index size be like, approximately?
I imagine that creating (and maintaining, of course) the index would be the most time-consuming part? Is it at all possible to imagine hosting this index... somewhere... and doing sqlite http range-like queries on it..?
Would it be enough to have an index consist of a list of found words, and the related "document ids"? i.e. "apple" is in doc ids 1000, 2000, 3000, "banana" is in doc ids 2000, 4000, etc.?
And have separate docid -> archive.org url mapping?
The problem, as with any search engine, is the ranking algorithm. Without a sufficent one, the search results are useless. What use is a list of every page in the Wayback Machine containing the word "apple"?
The Wayback Machine possibly would need a much larger index than any normal search engine: not only the present websites, but all the historical versions (though I don't know what proportion of the web they've indexed).
I was very naively "back of the envelope" prototyping a search engine. I both realize that this is not the way that these things are built, and both would really like to have that (or any) search engine to look through the archive..! :-)
As for the index, I agree -- I was trying to guesstimate its size by going from the ~~20PB total Wayback Machine size (which includes all historical versions). Is it 1% of 20PB (for the size of the text content), and then another ~10% of that for the index size? So 20TB...?
> I was imagining that a full text Wayback Machine search engine would mostly be useful to look for words unique enough that sifting through a lot of results (even if those were not ranked "well") could still be useful..?
If we think about use cases, users may often search specific domains. In that case, results ordered by frequency and/or date might be sufficient and even desireable.
> I was very naively "back of the envelope" prototyping a search engine. I both realize that this is not the way that these things are built
It's often the first step!
> So 20TB...?
That doesn't sound so bad.
I have 0 time for this, but I also can't easily let go ;-)
Want to collab on this?
You made my day.
Really the way I see it outside a few large banking firms, its kind of hard to be sure any provider of digital services would be around in the 50+ year term for this kind of public archive.
I hope the Internet Archive manages it.
EDIT: I do worry the IA has a bit of a lightning rod effect with skirting issues re: legality of archiving content. IMO its no guarantee it survives any significant time span either.
A library could do it. Perhaps leading institutions like the British Library or Library of Congress. I've thought that IA should be a Library of Congress project, and may eventually end up under their auspices.
For my needs, I run a report monthly for the content I’ve archived using my IA account to determine archived GBs, and then donate the amount needed to cover those costs.
Consider reaching out to their patron services email address with any questions.
Edit: $2/GB citation: https://help.archive.org/hc/en-us/articles/360014755952-Arch...
[1] https://help.archive.org/hc/en-us/articles/360014755952-Arch...
Do you have source/more info than that?
Lets say the internet archive is 100 PB [1], that's 100,000,000 GB [2], and at that rate it comes out to $200 million [3] for the whole thing forever. That's a lot of money, but also a lot less than I was expecting for something like that.
[1] https://www.protocol.com/internet-archive-preserving-future: "The web archive alone is about 45 petabytes — 4,500 terabytes — and the Internet Archive itself is about double that size (the group has other collections, like a huge database of educational films, music and even long-gone software programs)."
[2] https://www.google.com/search?q=100+petabytes+to+gb: "100 petabyte = 1e+8 gigabytes"
[3] https://www.google.com/search?q=1e%2B8+*+%242: "1e+8 * (US$ 2) = 200 million US$"
I suppose the claim is rather shocking and warrants citations - $200m to host the entire internet archive forever? I don't blame you for the excessive citation.
With a 9% discount rate, that’s only $3.63 million dollars in present value to pay for infinity months.
Of course there are other costs (and cheaper more efficient disks; and cheaper power; and your discount might be less aggressive; and server aren’t free tho you only need like 1 server for 100 disks with SAS expanders since most data is never read; and maintenance), but the $200 million number you got seems very reasonable to me.
Edit: I'm assuming they can deliver reliability and durability similar to modern cloud standards, like AWS S3.
Suffice it to say, physics is probably going to have a lot to say about that assumption in the coming decades.
I guess archiving a static blog is already trivial for their system, but I'd pay $100 a year to have IA host my static blog. The overhead I would consider as a donation to a worthy cause.
Then again, I can just give them $100 a year and find some free static hosting, like GitHub Pages, and call it a day.
The feature I most want from IA is a streamlined system to delete content they have archived on domains that I own, including a proper privacy law compliance effort on their part. They have intentionally made it a difficult, manual process to get content removed. They operate as a de facto malicious crawler.
They massively violate GDPR with how they operate and few seem to care about that fact, including all the commenters on HN (which universally give them a free pass on being malicious and violating GDPR very aggressively).
When IA has to comply with laws like GDPR, that's the end of IA.
If you had to choose between the GDPR, & an accurate historical record, which would you prefer?
An author suggesting that the LoC remove their copy of a book/other work (including digital works) because they want to unpublish it would not fly.
The parent comment has an issue with anything on the Web not hidden in some way being considered 'public' and 'published', but that would be something that would require international cooperation to hash out.
Will you be happy when you've burned down that library?
There are very good reasons that archives will not destroy or alter information outside of very clear difficult and manual processes.
And actually, looking at it, I don't think they're necessarily in violation of GDPR [0].
Point 3 says: "Where personal data are processed for archiving purposes in the public interest, Union or Member State law may provide for derogations from the rights referred to in Articles 15, 16, 18, 19, 20 and 21 subject to the conditions and safeguards referred to in paragraph 1 of this Article in so far as such rights are likely to render impossible or seriously impair the achievement of the specific purposes, and such derogations are necessary for the fulfilment of those purposes."
According to GDPR, national law of EU parties overrules GDPR when it comes to personal data being used in archival context. I don't know every EU country's stance, but most of the bigger economies would allow for this.
There is also a difference between deleting the data and rendering it inaccessible to the public. Keeping something under wraps is generally more 'acceptable', but active destruction of the item (digital or not) and its providence is much more limited. Also there's a difference between personally identifying data (covered by GDPR), your content (which would be covered under copyright and not GDPR), and connections people can make if that content is available (not covered at all because it's not anybody else's issue if you write something terrible and people keep recognizing you over it so long as you did actually write it).
> Article 3(2), a new feature of the GDPR, creates extraterritorial jurisdiction over companies that have nothing but an internet presence in the EU and offer goods or services to EU residents[1]. While the GDPR requires these companies[2] to follow its data processing rules, it leaves the question of enforcement unanswered. Regulations that cannot be enforced do little to protect the personal data of EU citizens.
> This article discusses how U.S. law affects the enforcement of Article 3(2). In reality, enforcing the GDPR on U.S. companies may be almost impossible. First, the U.S. prohibits enforcing of foreign-country fines. Thus, the EU enforcement power of fines for noncompliance is negligible. Second, enforcing the GDPR through the designated representative can be easily circumvented. Finally, a private lawsuit brought by in the EU may be impossible to enforce under U.S. law.
[snip]
> Currently, there is a hole in the GDPR wall that protects European Union personal data. Even with extraterritorial jurisdiction over U.S. companies with only an internet presence in the EU, the GDPR gives little in the way of tools to enforce it. Fines from supervisory authorities would be stopped by the prohibition on enforcing foreign fines. The company can evade enforcement through a representative simply by not designating one. Finally, private actions may be stalled on issues of personal jurisdiction. If a U.S. company completely disregards the GDPR while targeting customers in the EU, it can use the personal data of EU citizens without much fear of the consequences. While the extraterritorial jurisdiction created by Article 3(2) may have seemed like a good way to solve the problem of foreign companies who do not have a physical presence in the EU, it turns out to be practically useless.
https://en.wikipedia.org/wiki/InterPlanetary_File_System
You can put the storage costs on the nodes because storage at archive.org's scale adds up, especially when it's run by volunteers.
I guess after some time the nodes would agree on the hash and throw away the data because it would cost too much to store.
car analogy time: It is the same as reading a post about "how to lift my car to do work in the garage", and the the second paragraph starts with "using energy harvested from my perpetual motion machine"
I see blockchain as a technology that may develop useful applications, but-- in terms of current day usage-- I'm extremely skeptical when it's referenced in conjunction with applications that might achieve the same goals without it.
“Any sufficiently long Internet discussion will propose blockchain as a solution.”
rather than
“Blockchain is eating the world.”
So every time someone suggest to put some content on a blockchain I wonder if they realize that there are people that want to erase/remove their content from the internet. I also think it is dangerous to keep everything someone or some company created on the internet. It is too easy now to internet judge some adult about things they did while being young or to keep people accounted for mistakes they did and paid for them their duties to society.
I think if we ever build this feature on a blockchain I hope it is opt-in and people realize what that mean.
I actually think using blockchain for things like ensuring providence is interesting, since in archives being able to have a clean record of what happened to a piece is VERY useful. It just won't earn a ton of money, so we'll need to wait for the capitalism to burn off to see more not-for-profit uses.
Indeed the real challenge of archival is not loosing the stuff, by making sure that people can still find the stuff. "Orphaned" information that no one knows exists, or is bothering to interact with, isn't that valuable compared to resources that are actively being used and still "live" in the culture.
Of course, the archive can never serve the same amount of bandwidth, but the goal is a) interested parties can mirror the stuff they care about in a higher bandwidth / item way after some huge disruption c) random viewers never notice something going down, nor who is serving the info, but just a temporary drop in connection quality.
Ultimately, location-based addressing is a stupid way to run society, needlessly fragile by baking in very property claims (IPs, DNS, etc.) that are incidental to the task at hand. Content-based addressing, with location based hints to avoid trying to solve really hard problems all at once, is the only way to make culture more robust.
Of course, ensuring that there's persistence of attention as well is a tougher problem. But one only needs to look at sites like https://reddit.com/r/tumblr to realize that there is immense societal interest in "meme archaeology." Reducing the barriers to entry to would-be archaeologists, giving them a "chain" of breadcrumbs that lead to content, and building communities that will socially reward people for their archaeology work, is the best thing we can possibly do.
Erm, to me this sounds like putting up with link rot as hack around bad IP law? There are already IP exceptions for preservation. And if content-addressing was the norm, geocities-type sites might bow to market pressure to not "own" the content, but merely have some some sort of license for being the exclusive pinning service and running the ads or whatever. This is like avoiding the problem where your the rent on your current apartment doesn't fall as much as the market writ large because your landlord knows moving is not free.
That is to say, IPFS doesn't help if the desire blooms after the nodes dry up. Things could still be lost.
Concretely, this would be to skip the "many people on encountering a dead URL don't bother to try the internet archive" problem.
The incentive challenges are making sure that the average number of nodes is more than one, because, as Brewster likes to say, "libraries burn; it's what they do", plus all the traditional challenges of maintaining a commons at high levels of resilience. Once you have data on a network like IPFS, we can use a number of incentive models to make sure it stays there, including charitable projects like the Archive, government support (archives are traditionally state projects -- if every country's archive was pinning this content, it would be far more resilient), and decentralized incentive frameworks like Filecoin.
(Disclosure: I work for the Filecoin Foundation; in our decentralized preservation work, we've funded the Internet Archive's work in this area, though I should emphasise that IA works with a lot of different decentralizing technologies through their https://getdweb.net/ community.)
Why is IA not globally distributed, like a CDN?
I use IA for "problem" websites, e.g., ones that rely on SNI, i.e., ones hosted at certain CDNs. I simply add add these sites to a list and the local proxy does the rest.
http-request set-uri https://web.archive.org/web/1if_/http://%[req.hdr(host)]%[pathq] if { hdr(host) -m str -f list }
IA "hosts" an enormous number of sites without the need for SNI (plaintext hostnames sent over the wire).EDIT: @sebow the way they (re)format the HTML is less friendly to the text-only browser I use.
IA is more of an curated internet archive + explorer(which granted is very good).
What is the point of so-called "DNS privacy/Private DNS" if "anyone can tell which site you are visiting" simply by observing IP addresses, without any need to see domainnames.
If SNI (plaintext hostnames sent over the wire) is a non-issue, then why are people working on encrypted Client Hello in TLS1.3.
This is a neat shortcut to simply get the very first archived version! I often have to go to /*/ and manually click on one of them, which is very tiring.
Is there one to get the latest?
To get the link to the latest, can use memento. For example,
usage: echo example.com|1.sh
#! /bin/sh
read x;
curl -A "" https://web.archive.org/web/$(curl -A "" -s "https://web.archive.org/cdx/search/cdx?url=$x&fl=timestamp,original&limit=-1"|tr \\40 /)Off the top of my head, we lost a lot of common wisdom in dealing with the flu pandemic of 1918 because personal letters and most newspapers were not preserved. I think 100 years from now they might wish we had preserved more from marginal and/or world communities. What folk wisdom is being lost? Perhaps we need to expand our definition of what is worth saving.
That's what interests me. For example, there's a cool repository of 12 step speaker meeting talks hosted in Iceland [1] and frankly, some of the talks are junk, but there's a lot of wisdom. What I find interesting is how it showcases how ordinary citizens talk to each other. The words they use, the accents, the little gems of folk wisdom contained, along with some uncommon stories.
This will be valuable 100 years from now if, for example, you want to build a virtual world based in the mid to late 20th century and you want to get the accents correct. What phrases did people use? What were some common misconceptions? Maybe 100 years from now addiction will no longer be a problem. If my virtual world is to be accurate I need to know what it was like for ordinary people when it was a problem. Etc...
Its fascinating when you start looking into any historical time period (you wouldn't even need to go far back), before a lot of details are educated guesses. Since no one chose to record the mundane in detail or it failed to preserve over time.
Too add an example I know well.
I grew up on a small farm.
We have plenty of images of Christmas parties etc, but almost none showing actual work being done which is what I think my kids would appreciate the most.
Luckily YouTube for all its warts exist and I can look up the "motorized tea spoon", the U-9 Motostandard for them when I need to explain it: https://www.youtube.com/results?search_query=motostandard+u9...
(We had the one with front-mounted wagon and a stick for steering. And yes we had another slightly larger tractor as well, the AEBI Transporter TP50: )
The failure mode I see very often is that the frontend apparently doesn't know what the backend's doing: The part which ingests URLs and tells you what URLs have been archived does not know what archives the backend has, so it will tell you a page has been archived and give you a link to the archive, but when you click the link, it tells you it does not have the page archived, oh, look, it exists online, would you like to archive it now? Archive it again, and it will tell you that you can only archive a page once every 45 minutes. If you're a weird little obsessive like myself, you go through this process a half-dozen times for one page before it acknowledges that, yes, it does have the page archived (once, mind you) and you can actually see it.
While I'm filing bitch reports...
The Wayback Machine apparently loves setting cookies. It will set cookies until it has exceeded its own ability to accept cookies, at which point it will give you a blank page and you have to look in the developer console to figure out that it sent you a "too many cookies" error in the response header. I've had to force my browsers to not accept any cookies from the Internet Archive to fix this.
I call it “Beating the Samson Option”, of pulling the temple down upon your head.
https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
Perhaps other people still have a need for it.
Thank goodness the IA doesn't blindly obey robots.txt.
It turns out that the internet is not always forever. I find that comforting.
"The Internet is forever" isn't some natural law by which all content abides as though it can never disappear. It's a warning that you don't control the content once it's accessible on the Internet.
And requiring an explicit opt-in would basically mean no IA.
To be clear, the IA is a positive, maybe even a great one. But it skirts by because most people don't care. (They did as you say post whatever on the public internet.) Add the facts that they're a non-profit, aren't trying to monetize their hosting, and will generally take things if the owner asks.
Libraries and other archives have some very limited special rights (which mostly relate to making physical backups of physical books). But invoking "library" isn't some general get out of jail free card with respect to copyright.
These rights are also under constant attack: It's normal to charge libraries exorbitant prices for digital materials compared to their analogue counterparts, for example.
It effectively is. Your consent is not required, and people are doing far worse than just keeping it available (Clearview; there are also reports of people hoovering up encrypted data to crack in the coming decades when we're post-quantum).
This is no different than demanding people not keep track of anything else, and attacking archive.org might make you feel better, but that won't make anyone else stop.
1.) Most people won't opt-in because a significant majority accept defaults and don't opt into most things.
2.) For people like yourself probing a bit deeper, you might well ask whether you really want to give up your ability to decide you don't want something you thought was so funny when you wrote it at 20 now that you're a politician running for office or up for a political appointment.
It's fairly unethical.
Archives are exempt from being forbidden to create copies due to copyright infringement. The Library of Congress can make all the copies it wants, it just can't SELL them.
Now, there is a question whether a private company should legally be able to BE an archive of record, but as of now there's no legal reason they can't be, I believe. So it's legal.
Trying to muddle in privacy concerns with the accumulation of public knowledge undermines the whole concept of a shared society, or there being even the potential of accumulating "progress" in the first place.
Also, historically speaking, people with means used to save their letters for posterity, which proved to be a very valuable resource for future academics, so the idea of, what, deleting all your proton mails and signal messages as encouraged is arguably overshooting the return to some pre-internet norm.
Your example is an example of choice or consent. They also had the option to burn their hand written books and scrolls down periodically. Systems like IA take that choice away.
Source: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
You made it public on the damn web, people doing what they wanted with it is fair game per the original ethos of the web. If you wanted it to be private, you should have some auth or encryption.
The actual real world harm is in actually-common-used means of personal communication were network effects preclude on carrying out they business with technology appropriate to desired privacy levels. Chatting on my custom website was always niche, and therefore the network effects argument doesn't carry water.
We should refrain from worrying about robots.txt minute until the elephant in the room is put to rset.
Thanks.
Their documentation about it is rather crap right now, but its in several of their FAQs about it
I do wonder how many startups actually want to be archived, rather than just ditch everything with unseemly speed as soon as they get acquishutdown.
eg.
http://web.archive.org/web/20180901110658/https://www.snowfl...
http://web.archive.org/web/20140701061721/https://databricks...
A website could handle tons of scrappers without having high bandwidth, only the provider will need high bandwidth.
The issue is that scrapers often play with cookies and dynamic websites and such solution wouldn't work on these cases
I am doing some compression research, and would love to help IA in any way I can. There are some amazing SOTA compression algorithms available now.
And if IA ignores images/video, and focuses only on text, they can store an insane amount of websites at a very low cost.
that would be a huge and expensive paradigm shift for internet service backends that traditionally have been only designed to be run by one entity and typically are a mix of custom, open source and proprietary software that is run in a specific way.
i suppose it could be done, but there hasn't been any reason to make that investment. service backends also tend to be more "living software" where part of the system is the team that continuously builds, updates and operates it.
basically, it would look something like java applets, but for entire internet service backends. one step and all the services, databases, everything would spin up and start serving. that would be great but is probably a ways out.
that's not static assets, that's a multiuser game server.
i think he's envisioning entire internet service backends that can be packaged up like java applets and re-run on demand paired with some kind of decentralized serving infrastructure that any user can insert coins and resurrect a sophisticated web service from the past.
more likely i suspect we'll see more efforts by hobbyists to resurrect these things and more releases of backends from failed projects into the public domain.
with so much physical gear that requires service backends being made today, we may even see regulation that requires release of the source for a service when a service is shut down. crazy to think that if the company who made your car or tractor fails, that your perfectly good car or tractor could cease function when they shut down the service backend.
Complex repercussions obviously around acquisition, IP, and other business dimensions however. Maybe unworkable even. But I think there's a world where this actually exists and lowers the barrier to building business-critical software and selling to companies that need a 50-year commitment to risk you.
Not everyone fully supports the IA mission and they need to respect that view as much as they respect their supporters.
Removing data from the web in 2021? Hmm… https://web.archive.org/web/20211222032633/https://news.ycom... oops!
I don’t think that is correct. A lot of it was added via automated methods?
For example, I host my own ArchiveBox at home (you get fulltext search as a bonus) and it is configured to submit every URL I save to IA: https://imgur.com/a/Yhnxo1W IA considers that to be manual submission not subject to robots.txt rules.
Could someone ELI5?
This may not be the best example, but while I'm sure the Rosetta Stone, being a treaty, likely would have seemed like something worth preserving, would anyone have imagined that it would be the pivotal document in understanding ancient Egyptian? That it would be one of the most important documents of all time?
I would rather revert to the internet where quality content was published openly on the web. But presently web content, at least that upranked by Google, is very low quality (seo, clickbait, biased, trivial)