Analyzing One Million Robots.txt Files
intoli.com
intoli.com
The Internet Archive (archive.org) is currently running their end-of-year donation drive, if you value the work they do it's a good time to donate: https://archive.org/donate/
(and on the topic of robots.txt, it sounds like they're moving in the direction of disallowing people from using them indiscriminately to block access to valuable archival materials: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea... )
I ended up analyzing very different things from this article though, so this article was still pretty interesting to me.
“traditionally used for
vague attempts at humor
which signal to twenty-something
white males that this is
a “cool” place to work.”
WTF with the casual sexism/ageism?That being said, I think it's a little hypocritical to insult someone for dropping casually racist contents into their technical work when you're doing something very similar.
s/DNS server/third party open resolver/
IME, querying an authoritative server for the desired name triggers no such limitations.
One does not even need to use DNS to get the IP addresses for those authoritative servers, if the zone file is made available for free to the public as most are, under the ICANN rules.
I have thought about building a database of robots.txt many times. IMO, robots.txt has an important role besides thwarting "bots". It can thwart humans as well. It can be used to make entire websites "disappear" from the Internet Archive Wayback Machine.
Perhaps others are making mirrors of the IA.
However, I have thought it could be useful to monitor the robots.txt of important websites on a more frequent basis than IA, in order to (if possible) preemptively archive the IA's collections if robots.txt changes are ever detected that would effectively "erase" them from the IA.
Perhaps the greatest thing about robots.txt is that it is "plain text". This "rule" seems to be ubiquitously honoured. Did the author ever find any html, css, javascript or other surprises in any robots.txt file?
The Internet Archive is also modifying its policy on retroactive blocking using robots.txt, although I don’t have the blog post link handy at the moment.
If you’d like to mirror certain Internet Archive contents, every item is served as a torrent.
But hey, I guess it's one of those cases where the law and basic ethics clash a bit; with certain laws saying 'unauthorised' access to a server is illegal, then ignoring that would leave them under fire for that instead.