ROBOTS.TXT is a suicide note
archiveteam.org
archiveteam.org
Yes, robots.txt is no magic bullet against ill-behaving crawlers (as proven by ArchiveTeam) but it never was supposed to be that.
You choose to ignore my specific wish not to be crawled by you? Fair enough, I'll return the favour and simply block your useragent
ArchiveTeam ArchiveBot/[DATECODE] (wpull [VERSION]) and not Mozilla/5.0 (Windows NT 6.1; WOW64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/[VERSION] Safari/537.36
In Apache for example # Remove the ^ anchor in case text gets prepended
RewriteCond %{HTTP_USER_AGENT} ^ArchiveTeam
RewriteRule .* - [F,L]
and if possible your IP ranges.I know they often don't... and I wonder why they don't, but there's no reason they should.
Disallow: /harming/humans
Disallow: /ignoring/human/orders
Disallow: /harm/to/selfResource Limit Is Reached
The website is temporarily unable to service your request as it exceeded resource limit. Please try again later.
The irony! Cache link:
http://webcache.googleusercontent.com/search?ei=h0J3WMKKA4er...
So blocking the archive blocks new content added after robots.txt is made available, and adding a certain few lines in the robots.txt file explicitly hides older content on the domain in the archive.
That way, everyone wins. Sites that really want to remove content for some insane reason can do so, older sites usually aren't lost if the domain expires and domain holding page owners/cybersquatters don't accidentally cause older content to be hidden (since hey, they don't really want to block it, just stop their holding page from being archived).
I think the crux of the matter is found here:
> If you don't want people to have your data, don't put it online.
As much as I agree in principal with this, because of the way web requests work, I don't want to be associated with this group.
You cannot ignore copyright, and robots.txt is exactly what I would use if I didn't want something archived by an organisation I have nothing to do with.
AFAICT this page is a reaction to an archive.org policy of respecting robots.txt retroactively - e.g. oldwebsite.com runs from 1999-2009, domain expires in 2010, gets bought in 2011 and the new owners add a robots.txt disallowing IA. The archive.org copies for 10 years are now inaccessible.
One group has respect for authorship, and one does not.
It may not be the most palatable solution, but hardly a need for a tantrum, and intent to ignore well established rights.
Also makes me wonder if a solution could be implemented where domain owners can explicitly give permission for their work to be archived, with the assumption that all content from the current 'iteration' of the site remains accessible, even if the domain changes ownership. So I could "I give the Internet Archive full permission to archive DomainA.com", so the archive keeps said info accessible even if the domain is sold to someone else/expires.
Then again, Kafka wanted his unpublished works burned after his death - and the world is arguably a better place for having ignored his wishes.
SEO is where robots.txt shines right now. It's not that people are trying to hide something it's because we don't want it to conflict with the content we actually want to promote.
This file doesn't "block" anything, it simply asks the robot to do something, which implies that it is probably being abnormally careful: truly annoying, unwanted, careless robots might follow these guidelines, but that seems like a stretch. In reality, this file exists so that extra careful robots are able to get feedback from websites that have extremely narrow bandwidth availability or extremely high generation cost... concepts which this article makes a pretty compelling argument for "that doesn't make sense". In practice, this file then makes the owners of websites sometimes think "I can build something weirdly broken (such as a procedurally generated content tarpit, or mapping anonymous GET requests to database insertions) and just rely on this file to explain what I did along with enforcing rate limits and boundaries"... and then an "unwanted careless webcrawler" comes along and causes them serious issues. It is akin to having your entire webserver crash if someone sends you a non-ASCII character in a form field, but thinking "this will work out: I have a little flag in my HTML file that makes it clear I only accept ASCII". If you absolutely feel like you need to block something, then actually block it: any robot gracious enough to pay attention to this file is also going to send a useful user agent, and you can use that to return a legitimate 403.
One important use case to exclude sections of your website is to not pollute the sitemap which Google crawls or to be more precise--the daily crawl volume Google allocates to your site. If you let every page be crawled more important pages get crawled less. Example: In the past, you created a content category which didn't turn out successful. Before you remove this category with plenty of links which would result in crawl errors it would be smarter to exclude them in the ROBOTS file and focus on your core categories.
This is very similar to, taking photos/videos of people on street without their consent, and archiving and publishing. Even more, like taking photo of someone and publishing, who is wearing a t-shirt saying please don't take photo of me.
Sorry but if you will use my server resources, you will be bound with my rules.
If anything, a robots.txt would encourage archiving because people would be annoyed with the asocial attitude.
User-agent: *
Disallow: /secret/Archived at
https://archive.fo/http://www.archiveteam.org/index.php?titl...
The lesson from this is that ROBOTS.txt works as long as everyone follows a line set in sand.