Detecting Bots in Apache and Nginx Logs Using Python
tech.marksblogg.com
tech.marksblogg.com
I think the OP needs to break the problem down into 1) access from bots that identify themselves (e.g. googlebot, bingbot, etc) 2) access from bots that masquerade as humans and 3) activity that is from humans.
#1 should be trivial by looking at the user-agents. #2 and #3 could utilise machine learning to categorise the behaviour of the connecting party, etc
It's not particularly clear why the OP wants to detect bots... I suspect it's to get a clearer signal of what assets are being accessed by humans.
The best detections for human vs non-human so far have been (1) introducing captcha and (2) based on view time and interaction (hotspot). With both we can have a reasonable criteria to build a somewhat simple model.
If based on server access log, then you need to group resources together (css, js, images, html) or ignore most of the resources, and calculate view time. Based on some reasonable expectation, a user who views multiple pages at the same time or within +X seconds would likely be a robot than a human despite we humans are used to opening up several tabs once we are used to how to use the website. As many tests do put a sleep in between (since there's always lag) so for the non-intrusive robot, we will have a difficult time distinguishing within some reasonable confidence.
If we enabled tracing / tracking, then this becomes slightly easier as well since we can learn the behavior of "new" and "veteran" users. This is why tracking is such as a privacy issue for many.
* Did they request robots.txt?
* User-agent
* Logged in vs. guest account (if applicable)
* Number of requests/time period
* Patterns of requests
The real value of this kind of analysis (to me) is to bucket the types of visitors: * bot
* well-behaved (e.g. googlebot, etc.)
* valuable
* not-valuable
* badly behaved
* visitor
The valuable vs. non-valuable traffic is interesting to me, I've seen sites where the bot traffic was 25-35% of the total, and that some of the bots, though well-behaved, didn't really bring any business.I run a VPN through Hetzner, so requests from my IP are not a bot (I hope!). Really you want to look at the paths (filtering out all the /w00tw00t requests) and the user agents above all, which the author touches on. However a whitelist approach is better than a blacklist IMO.
Also in the `in_block` you really want to hoist the `IPAddress(ip)` call out of the `any()` loop!
def in_block(ip, block):
ip_addr = IPAddress(ip)
return any(ip_addr in IPNetwork(cidr) for cidr in block)
Or just be explicit: def in_block(ip, block):
ip_addr = IPAddress(ip)
for cidr in block:
if ip_addr in IPNetwork(cidr):
return True
return FalseInstead of `open(filename, 'r+b').read().split('\n')` you can use `for line in open(filename):`, which avoids loading large files into memory. (small gotcha: `line` will contain the newline character(s) which can be stripped via e.g. `rstrip`)
You can also drop the square brackets from calls to methods that take an iterable e.g. `any`/`all`, `set`, and `join`. So `join([...])` becomes simply `join(...)`. Python will use a generator expressions instead of constructing and passing a new list list. To quote PEP 289: "generator expressions [are] a high performance, memory efficient generalization of list comprehensions"
These can really make a difference with big files/lists, but are a good habit in any case. I hope it helps in the future!