They say they "declined to take down the archives"-- but they didn't in fact do this at all, they just insisted a request to take down the archives come in the form of a robots.txt, and they automatically and without review comply with all such requests in the form of a robots.txt. They don't in fact ever decline to take down any archives, if the request is properly given as a robots.txt.
I don't know why they bothered making statements about "declining to take down the archives" in the first place (to the journalist or to us), making comments about "Reid’s being a journalist (a very high-profile one, at that) and the journalistic nature of the blog archive" -- they did not in fact "decline to take down the archives" at all. The "journalistic nature of the archive" was in fact irrelevant. They took em down. They are down.
If that file were to be removed, presumably the archive would again be served up upon request.
The lawyers were asking for the archive store itself to be wiped.
The Internet Archive has a mechanism for doing this, as I understand it. It involves asserting copyright over the material in question and essentially "making a case" for removal. IA decided the case they made didn't pass muster, and denied specific removal on those grounds, which is why they mention "journalistic nature of the archive" and so forth.
But that's entirely orthogonal to their policy of treating active maintenance of robots.txt as indicative of positive copyright assertion over the contents of an entire domain -- which Ms. Reid's team appears to have taken as a fallback position. They couldn't get the sanitized archive they wanted, so they just made the whole thing invisible.
That seems like a pretty glaring flaw in something designed to create an enduring record.
[1]: Sure, some people just want to hide embarrassing or incriminating content, but there’s also cases where someone is being stalked or harassed based on things they shared online, and hiding those things from Archive users may mitigate that.
I don't think it's mentioned in an official document, but it's usually referred to as "darking".
It probably safe to assume that the same concept applies to the Wayback Machine as to the rest of IA.
Edit: Here's a page that indirectly conveys some information about it: https://archive.org/details/IA_books_QA_codes
[0]https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
> We are now looking to do this more broadly.
That's the part I'm asking about.
from the faq: http://netarkivet.dk/in-english/faq/#anchor8
8. Do you respect robots.txt? No, we do not. When we collect the Danish part of the internet we ignore the so-called robots.txt directives. Studies from 2003-2004 showed that many of the truly important web sites (e.g. news media, political parties) had very stringent robots.txt directives. If we follow these directives, very little or nothing at all will be archived from those websites. Therefore, ignoring robots.txt is explicitly mentioned in the commentary to the law as being necessary in order to collect all relevant material
I wonder if there are any other national archives of the internet that do the same.
https://www.bl.uk/collection-guides/uk-web-archive describe the much more limited aproach taken by the british library much later in time, but might extend to a similar scope.
I think if I were running a national or internationally mandated archiving initiative I would basically want to take in content from Internet Archive, and not remove things, and probably it would be less expensive that way than having my own crawler.
The key is it works both ways. By respecting the live robots.txt, and only the live one, data hiding must be an active process requested on an ongoing basis by a live entity. As soon as the entity goes defunct, any previously scraped data is automatically republished. Thus archive.org is protected from lawsuit by any extant organisaton, yet in the long run still archives everything it reasonably can.
The solution (which the Internet Archive really needs to implement) is to look at the domain registration data or something, and then only remove content if the same owner updated the robots.txt file. If not, then just disallow archiving any new content, since the new domain owner usually has no right to decide what happens to the old site content.
Anything otherwise obviates the entire mission of the Wayback Machine.
She's a prominent left-wing television personality. Would you be so accommodating if it were, say, Tucker Carlson trying to scrub embarrassing information about himself from the wayback machine?
I don't know what "start from scratch" would mean – the point is that each site is sampled many times throughout history. That said, it is very odd that a current change in robots.txt would prevent looking at old samples. And that's indeed what it looks like [1]:
> Page cannot be displayed due to robots.txt.
The robots.txt shows a positive assertion that parts of a site should be excluded from being used by automated systems.
In most cases I imagine WBM does not have permission of the owner to keep a duplicate of the site, it's certainly tortuous in UK law.
Sites that don't change their robots.txt are probably highly correlated with sites that don't sue for the infringement.
""To remove your site from the Wayback Machine, place a robots.txt file at the top level of your site (e.g. www.yourdomain.com/robots.txt).
The robots.txt file will do two things:
1. It will remove documents from your domain from the Wayback Machine. 2. It will tell us not to crawl your site in the future.
To exclude the Internet Archive’s crawler (and remove documents from the Wayback Machine) while allowing all other robots to crawl your site, your robots.txt file should say:
User-agent: ia_archiver Disallow: /""
Their current documentation no longer says that they stop displaying old archives automatically in the presence of an ia_archiver Disallow directive, but I have not experimented about whether they still actually do this anyway.
Just guessing wildly about similar issues, I know that news organizations which have a publicly documented "unpublish" policy tend to get that policy used aggressively by reputation management firms and the like.
I don’t think that’s it...it’s not a technical thing. Deleting all archives must be a courtesy they extend to anyone that specifically denies access to Wayback Machine in their robots.txt. Does anyone know if this is documented? If so, why didn’t her lawyers just carry out the robots.txt technique and not even bother contacting them? Most importantly, why would they have such a policy? This is all very odd.
I hang around in emulation circles and there's been some talk in the past few weeks because some Nintendo ROM archives had been taken offline from archive.org but people soon figured out that they could still access them by tinkering with the URL. The situation is a bit different here though.