Donate to the Internet Archive
archive.org
archive.org
This has ruined many supposedly permanent links. The infamous "She's a Flight Risk" blog from a decade ago is down.[1] My first website is missing. Even public domain stuff like NASA's report on nuclear propulsion is gone.[2]
With just a small rule change (obey robots.txt at the time of crawling), they could eliminate the risk of a page disappearing. Instead, we're stuck with a slower version of the link rot we're used to. It doesn't stop me from supporting them, but it's incredibly frustrating.
1. https://web.archive.org/web/*/http://www.aflightrisk.blogspo...
2. http://web.archive.org/web/20121029225832/http://ntrs.nasa.g...
From asking around they do retain the original data, but exclude it from public results following a robots.txt exclusion. It's something I do wish they would re-consider though as it limits the purpose.
While the distinction may be useful to someone with access to the data, it doesn't matter to me. From my point of view, the URL goes from available to unavailable. It's indistinguishable from typical link rot.
The Internet Archive has kept their robots.txt policy for over a decade now, despite constant requests to change it. I doubt they'll change it any time soon.
https://archive.org/post/406632/why-does-the-wayback-machine...
Maybe if it gets enough attention, and/or the Internet Archive people get enough e-mails about this problem (Wayback Machine obeying robots.txt) then they'll change their mind.
https://archive.org/about/bios.php
Here's a link with the director's email:
Does the Wayback Machine retain data excluded by new robots.txt rules? (In other words: If you change your policy in the future, can the change be retroactive?)
Why does archive.org keep this policy? It drastically limits what The Wayback Machine could be. I've searched quite a bit, but I haven't found a satisfying answer.[1]
1. https://archive.org/post/1019415/retroactive-robotstxt-remov... contains links to previous discussion on archive.org.
The irony that the policy can only be viewed through the wayback machine is not lost on me.
I don't think you can claim that when second-parties dictate your retention of existent data[0].
An archive is a place to which I can confidently go to retrieve a document in line with the retainer's retention and access policy. If the retainer doesn't control the retention of existing documents then... it's not really an archive. It's just an ephemeral store which may or may not still have the document in which I'm interested.
[0] in terms of raw etymology you are correct, in that arkheia just meant 'public records'. But having to regress to the Greek origin is a bit of a stretch.
Archive.org has mostly gotten away with what they do based on the fact that they try to be comprehensive, they don't post ads or charge for access, AND that they won't display your site if you ask them (through robots.txt). Yes it's frustrating but barring some legal ruling that cements their right to archive and offer access to copyrighted works, I don't expect it to change.
So much content has been lost and sadly the Internet Archive is not an archive whilst this policy exists which evicts all historical content when a updated robots.txt is found.
I have been for quite some time been donating 1TB-2TB/month of bandwidth in support of Jason Scott and the ArchiveTeam. This has been my way of supporting the Archive.org project and would consider increasing the bandwidth donation in exchange for resolution to the robots.txt bit-rot.
I won't link to the leaderboard but for reference I'm currently working on the TwitPic project and am within the top 10; same username as HN.
The retroactive-robots.txt policy made sense originally as a way of reducing risk from angry-rightsholders, while minimizing burdens on staff time, and had little downside when the history-of-the-web was short, and most domains were still under their original ownership. It was a toggle any webmaster could throw, with no support/maintenance effort required at IA.
It obviously sucks now, more than a decade after it was adopted as a quick fix, but would take some dedicated policy and technical design to gracefully replace. For example, many rightsholders may be relying on the old behavior. But, the IA hasn't yet been able to prioritize creation of a new scheme.
A new process could involve a policy where someone claiming to be the original site/content rightsholder asserts that as of some boundary date, for example when they ceded or sold the domain, later robots.txt should not affect earlier content. (Such a boundary could also be clearly-indicated in Wayback summary/calendar pages.) Then, presumption would flip to showing the earlier content, unless some other rightsholder (such as the current domain-holder) formally claims ("under penalty of perjury", etc) that it's their material and they wish the block to stay in place.
It'd be a bit in overall shape like the DMCA takedown/counter-notify procedure, as if the robots.txt was a sloppy takedown request. Squatters making a false claim to ownership of older content would be forced to go on record and take some risk. Ideally any dispute could then proceed between the other two parties, leaving the IA out of it. Again, this would be similar to how the DMCA tries to leave ISPs/caches/hosters out of takedown legal battles.
I doubt most new domain-owners are strongly or consciously trying to hide the past; it's usually an automatic choice made for other reasons. A few might be holding the history hostage to make ownership of the domain more valuable – "buy it to re-access the past". Some others might be embarrassed of the previous unaffiliated content – but a clearer Wayback UI indicating changes-of-management could help allay that concern. Unfortunately, for both good and bad reasons, there is no easy/reliable/canonical source of all domain-ownership info over time.
I love you, Internet Archive.
I'm not sure whether that means they need more funding or whether they're simply unresponsive, but it definitely didn't help me solve my problems with a trademark troll.
[1] http://www.faqs.org/tax-exempt/CA/Internet-Archive.html#anal...
https://archive.org/details/byte-magazine
http://en.wikipedia.org/w/index.php?title=CHIP-8&diff=585383...
http://en.wikipedia.org/w/index.php?title=Halt_and_Catch_Fir...
It really is a different beast from your normal panopticon internet tracking. This is more akin to showing up in the newspaper, rather than a surveillance camera.