How to detect web bots?
antoinevastel.com
antoinevastel.com
What if the bot is making individual requests from unique IP addresses? What if it's scraping many pages over time rather than a single smash-and-grab approach?
The article admits that this is a hard thing to solve. In most cases, it's probably not very worthwhile to try to detect bots in the ways that are suggested. Focusing on patterns in the server logs to find out what's being targeted might be more beneficial. Then slap a login or captcha around anything valuable that bots shouldn't have access to
In the latter category I have heterogenous load balancing, segregated by traffic type. For example, if you have a read-mostly web application, often the most expensive parts are in the administrative functions. If they run on separate hardware, then a regression doesn’t affect everyone, and a client side bug or flood of traffic doesn’t automatically lock people out of the admin workflow.
Similarly, you have server side rendering for web crawlers and anyone with JavaScript disabled for Angular and other interactive frameworks, which generally involves running a small quantity of specialized services on a few machines.
Shunting all users who smell like a bot to a small set of machines doesn’t lock them out, but it does preserve the experience for everybody else.
And putting captchas or logins on anything you have to protect from bots also "protects" from real people, most of them simply leave and don't bother with those thing.
Congrats, you just identified a visually impaired user with a screen reader who uses a VPN through AWS/GCP/etc.
I'm surprised there are so many trolling comments here.
I would absolutely love to see that as someone who scrapes sites. It would make me smile immensely to see the trolling.
https://en.m.wikipedia.org/wiki/Honeytoken
I have also done this as an anti-botting tactic myself though and it can get pretty complicated to setup properly unless you're doing the low-tech "feed these networks BS data" list.
Every few months someone discovers the correlation between our log messages and user agents and we rehash the same discussion about a particular badly behaved crawler that produces enough log noise that it distracts from identifying new regressions.
I coordinated an effort to fix a problem with bad encoding of query parameters ages ago, but we still see broken ones from this one bot.
Don't worry about backlinks - they know you've registered a domain so already checking A records in domain DNS. They are much more clever that 10 years ago when we could expect 0 requests to site that was just set up.
Discriminating against bots just makes the internet worse for everyone. In the future, everyone will have personal bots and agents to help automate many things.
If corporations can have personhood rights, why can't non-human agents?
It doesn't make much sense to me for someone to have the 'right' to demand me to respond to them nor their Python loop.
I wrote something that leverages AWS Lambdas to get around rate limiting, this solution would've tagged me instantly.
I hate this sentiment that every one who doesn't want to be tracked is a criminal.
It's the digital version of "You don't have anything to hide if you aren't doing bad".
---
I don't understand the fear about bots or scraping. As long as bots are behaving nicely (not slamming servers), they are just as much web citizens as humans.
The web is about sharing of information and having an entire company about exterminating the viability of bots is horrifying.
> I hate this sentiment that every one who doesn't want to be tracked is a criminal.
The statement you are referring to is more accurately paraphrased as “everyone who is a criminal does not want to be tracked”, which is very different from “everyone who does not want to be tracked is a criminal.”
> It's the digital version of "You don't have anything to hide if you aren't doing bad".
For the reason noted above, the statement at issue is actually the digital version of “if you are doing something bad, you do have something to hide,” which is a near-universal truth (I mean, it not true if you have actual or practical immunity from any accountability for your wrongdoing, but otherwise it's true.)
The sentiment is "and therefore we can discriminate against people who use user-agents[0] that don't actively help us violate their privacy rights".
0: such as screen readers, for that zesty ADA-lawsuit flavour.
I have a website I use to make money to feed my family. Shitty bots mess up my website, cost me money. A SaaS offers a solution, which saves me money.
What about that line of logic is horrifying to you?
More bots, more noise, bad analytics
More bots, IP theft at scale, competitors now have my content
I don't understand how this is hard to wrap one's head around. This is why we can't have nice things.
Edit - I offer 3 points to refute parent and I'm getting downvoted. This site is moderated by a bunch of bots!
If you don't force botters to take extreme camouflage measures, bots are easily filtered out of logs (and offer a potentially useful metric of their own).
If a business is threatened by "competitors" simply scraping published information, it's probably already doomed from the start. And I would posit that most site owners with this mindset originally got their data by scraping other sites, which is why they feel vulnerable. Compete over elements that are actually valuable!
Sure, you can try, but it's my right (and sport) to make it as hard as I can for you. :)
People like to imagine feel-good abstract ideas behind both bots and Tor like "the freedom/openness of information" or some journalist trying to visit your website under some strict regime. In reality, 99% of it is just abuse or someone trying to make a buck off you. It's not really as romantic as you think.
And FYI the bots are destroying the in-game economy in my game by doing extreme grinding no normal human being is capable of. It's a very hard problem to tackle with mechanics.
Who said that? You can ban bots for bad behaviour.
>As long as bots are behaving nicely (not slamming servers), they are just as much web citizens as humans.
I enthusiaticly support banning bots for abusive behaviour. I also enthusiaticly support banning humans for abusive behaviour.
> extreme grinding no normal human being is capable of
> a very hard problem to tackle with mechanics.
- diminishing returns for over-farming particular areas (per-player, maybe also total)
- stat penalties for playing too long / bonuses for rest ("You have worked a fourty-hour week in this game, don't you have a life you should be getting back to?")
- moving resources[0] around to break pathfinding, and/or adding poisoned ones to discourage blind grabbing.
- or just ban people who regularly dump extreme quantities of grind-farmed stuff into the economy (since that is the behaviour you're trying to stop), bots or no bots.
0: including enemies, who might, for example, run away from someone who's been slaughtering them for the past six hours
I still don't understand why you think bots should be allowed in my game at all. I don't want them and the players certainly don't want them (bot accusation is the number one drama).
Could anybody link me up with something to read on how to detect it is a VM based on TLS fingerprint?
Non-passive: It's also possible to read a GPU name like "VMware SVGA" from JS, or watch for mismatched/wrong hardwareConcurrency.
I hate credential stuffing and malicious activity as much as the next person, but I would never sacrifice user privacy like that.
Also, you can spoof a webcam (but that's not the real problem).
Scrapers that know what they're doing could implement this in a day. There are already easier ways to detect scrapers that don't know what they're doing.
And that's just the start of a long cat & mouse game that'd inevitably end up even more user-hostile and unaccessible than requiring a webcam to browse a website.
Will be interesting to see what happens to other industry players like PerimeterX. Distil was eaten by the corpse of Imperva, and I don't see Akamai making strong headway with Botman.
Google is going after this too with reCAPTCHA. The HN reaction to that has been interesting.
It's interesting to me how many comments in these threads talk about scraping as the issue with bots. Every sale I saw when I worked on this problem was related to credential stuffing. Seems the enterprise dollars are in the fraud space, but the HN sentiment is in scraping.
Funny how disconnected the community here can be from what I saw first-hand as the "real" issue. Makes me wonder what other topics it gets wrong. Surely my area of expertise isn't special.
No one credential stuffs for innocent fun.