Wikipedia's robots.txt
en.wikipedia.org
en.wikipedia.org
At Reddit we originally blocked a couple of crawlers but then realized how pointless that was. The entire robots file[0] is basically now just for google. All of the restrictions are enforced on the server side because there were so many bad bots it didn’t matter if we listed them or not.
For sure, but that's great, shame can be very effective.
$ curl -sIXGET 'https://www.reddit.com/robots.txt' -H 'cookie: rseor3=true;' | head -1
HTTP/2 501
Compare with: $ curl -sIXGET 'https://www.reddit.com/robots.txt' | head -1
HTTP/2 200He explained it a bit on a Twitter thread [1] where I brought it up a while back and it sounds like they didn't like the fact that he was using GitHub to host XML files for another service because of the traffic from crawlers it created.
Allow: /humans.txt
Disallow: /
Nice!
# Observed spamming large amounts of https://en.wikipedia.org/?curid=NNNNNN
# and ignoring 429 ratelimit responses, claims to respect robots:
# http://mj12bot.com/
User-agent: MJ12bot
Disallow: /
Coincidentally, I've just read more negative things about MJ12bot last week: http://boston.conman.org/2019/07/09.1To be honest, robots.txt is not for these kinds of bots. These kinds of bots are either malicious or incompetent. But more importantly, they're 100% useless to you as a website operator. They offer no SEO benefit, drive no significant traffic and simply consume resources.
The answer, sadly, is to hit them at the web server / load balancer / reverse proxy layer and just bruteforce all these bad actors away.
They'll never stop trying, though. Checking some NGINX logs for some of these bots that have been blocked for years, they still knock on the door over and over again.
iptables -A INPUT -s 207.244.157.10/32 -j DROPEdit: Added count of class-A blocks.
# Folks get annoyed when XfD discussions end up the number 1 google hit for their name.
I remember using it to download whole (small) websites on dialup and then to read offline.
https://meta.wikimedia.org/wiki/Data_dump_torrents
I, somewhat fancifully, keep two flashdrives with a wikipedia data dump on them "just in case"
I think both robots.txt and security.txt are great ideas. However, they will always only be useful to those who follow the wishes of the website (which hopefully outweight those who do not).
If a client or IP range is misbehaving in the server logs, it goes into robots.txt. If it's ignoring robots.txt, it gets added to the firewall's deny list.
I've tried to automate that process a few times but haven't ever gotten far. It's unending, though. Feels like all it seems to take is a handful of cash and a few days to start an SEO marketing buzzword company with its own crawler, all to build yet another thing for us to block.
If you write a crawler, you probably don't want it to waste time indexing a list of articles in every possible sort order, trying all "reply" buttons, things like that.
For me, a "Disallow" line in robots.txt means "don't bother, nothing interesting here". It is a suggestion that benefits everyone when followed, not an access control list.
On the other hand, many websites (like wikipedia here) hide interesting pages behind a Disallow.
Does anyone know what this comment means? It is towards the end of the file.
In wikisyntax, starting a line with a space ensure's that its in a pre tag, so this was a way to make the onwiki page show the entire thing in a pre tag and not use normal formatting. This seems to be broken by actually starting a literal pre tag inside the first line, but I'm pretty sure this used to work as a way to display the entire thing in a pre tag.
For the curious, the code generating the robots.txt file is at https://github.com/wikimedia/operations-mediawiki-config/blo...