You are going to hit a lot more false positives with this one than actual bots
You are going to hit a lot more false positives with this one than actual bots
What in the flying **. Is this a common thing?
For example: https://bright-sdk.com/
> Bright SDK is approved by Apple, Amazon, LG, Huawei, Samsung app stores, and is whitelisted by top Antivirus companies.
humans are a truly horrible species and this kind of thing is a great example of why I believe that.
So is that such a bad thing? If OP is going to use this to provide data about bots, blocking mass amounts of the internet could actually be a terrific example of how many people are at least tangentially connected to bots.
And "a lot" of false positives?? Recall, robots.txt is set to ignore this, so only malicious web scanners will hit it.
Personally I sometimes do a quick request to /wp-admin to check if a site is WordPress, so I guess that has a nonzero chance of affecting me. And when I mirror a website I almost always ignore robots.txt (I'm not a robot and I do it for myself). And when I randomly open robots.txt and see a weird url I often visit it. And these are just my quirks. Not a problem for a fun website, but please don't ban a whole IP - or even whole ISP - because of this.
So that is a balance between a bad actor and even "stop it" blocks, and auto expire means transitory denial.
add a captcha by limiting IP requests or return 429 to rate limit by IP. Using popular solutions like cloudflare could help reduce the load. Restrict by country. Alternatively, put in a login page which only solves the captcha and issues a session.
of course people look at this. it's not an everyday thing for the prototypical web user, but some of us look at those a lot.
I... I do... sometimes. Mostly curiosity when the thought randomly pops on my head. I mean, I know I might be flagged by the website as someone weird/unusual/suspicious, but sometimes I do it anyway.
Btw, do you know if there's any easter egg on Hacker News' own robots.txt? Because there might be.