Bad bots account for most internet traffic? Analysis
securityweek.com
securityweek.com
From the top...
> Bad Bots are increasing dramatically — Arkose estimates that 73% of all internet traffic currently (Q3, 2023) comprises Bad Bots and related fraud farm traffic.
Note that "Bad Bots" are never defined, and that 73% will never be explained, either to say how they mean it (73% of connections is plausible, 73% of packets is not, not in a world with Netflix and YouTube), or how they would measure such a thing.
Then we get into little gems like
> Scraping, it must be said, is a legally murky area. It is not specifically illegal; but if it defies a website’s published terms of use, it is certainly immoral.
Which... that's a claim you can make, but if you just throw out terms like "certainly immoral" without bothering to substantiate them I get to laugh at you.
The tendency to cast aspersions without substantiating them gets substantially more pronounced later:
> This is a website you can use to make sure your bots aren’t getting prevented by a website,” [...] “You can purchase this software. It has enterprise support and so on. But it is purpose built to commit crime. That is what it does. And there are many other different websites like this, but they look like legitimate businesses. It is a good example of a product purpose built to commit fraud.”
I'll grant that businesses don't like people bypassing their protections, but that doesn't make it automatically illegal or fraud to scrape or even to programmatically interact with sites. I'm not naive enough to think that it's not being used for illegal stuff, but that doesn't mean you get to paint the tools with a single brush.
This stuck out to me as well. What are "bad bots" and how do they differ from "good bots", and how are the authors determining which is which?
Not in a world where pron exists.
- second paragraph, first sentence
That said, the language around the “73% of traffic” figure is still a little ambiguous but most plausibly means “73% of requests/sessions to the significant commercial properties we studied”
Assuming that’s what they mean, it’s reasonable that bots are banging on major sites three times more often than legitimate users.
I think the linked article is doing the sloppy sensationalizing and the report appears to be much less so.
> Scraping, it must be said, is a legally murky area. It is not specifically illegal
There's already a jump happening here from "it's not specifically illegal" to "it's a gray area because it's immoral." Even if it was certainly and unambiguously immoral, that does not make it automatically a legally murky area. In fact, scraping seems to be reasonably established as specifically legal.
So the article is already kind of edging away from reality, where they acknowledge "this is not a crime" but they want to blur that and say that it's legally dubious, even though what they really mean is that they believe it's immoral.
Which, fine, I don't respect it but whatever. But then we jump to:
> But it is purpose built to commit crime.
Wait, no. Scraping is not specifically illegal. This is the emulator debate again; we have something that really is strongly implied to be legal and has won court cases before. It's not like it's an untested gray area, we know that there are scenarios where scraping websites against the permission of the website owner is not illegal.
So the article pretends it's more of a gray area than it actually is, and then once it has thrown doubt on it, drops the pretense entirely and calls it illegal. Not all scraping is fraud and not all scraping is illegal, and it's pretty galling for the article to admit that scraping is not illegal and then to turn around and pretend that they didn't just admit that.
> It is also a good example of crime-as-a-service.
No, it's not. The article admitted that scraping is not specifically illegal. Tools to help scrape websites are not "crime-as-a-service" because objectively (and there is legal precedent backing this up) scraping is not a crime.
----
The propaganda strategy being employed here is to admit a piece of factual information (scraping is not illegal), however hesitantly, and then to pretend that it was never admitted.
If anyone calls the propagandist out, they can point to their admission and say, "no, I acknowledged that it's not illegal, see". But then for every other part of the conversation, they just ignore that and act like it is illegal, even though at best all they've put forward is that they think it's more dubious than courts currently recognize.
They concede the point only to the degree they're forced to and only to the degree that would prevent someone from saying that they're lying or ignoring reality. But as soon as the context changes, they go back to acting like the point was never conceded.
This has the benefit of not only allowing them to make the same arguments they wanted to make before, but also sort of trains listeners to think of the original conceded point as more murky than it actually is. "Sure," the reader thinks, "it's technically not illegal, but at the point we're talking about crime-as-a-service and illegal tools, surely there's something legally dubious going on, this must be more of a gray area" -- even though not only has the article not backed up that idea, it's actually admitted that the idea is wrong.
> There's already a jump happening here from "it's not specifically illegal" to "it's a gray area because it's immoral. Even if it was certainly immoral, that does not make it automatically a legally murky area. In fact, scraping seems to be reasonably established as specifically legal.
Yeah, I didn't want to get distracted from my main arguments, but that stuck out to me too - in my non-lawyer amateur understanding, the LinkedIn case[0] means that in the US scraping very much is explicitly legal.
It's very: "technically speaking, the Constitution doesn't explicitly require people to quarter troops in their own homes, but..."
Not outright wrong, just kind of missing a lot of context and quietly implying a lot of stuff that is wrong.
Arkose runs a bot detection service, OpenAI being one of their customers (check your requests on chat.openai.com). I’m guessing the 73% are the number of requests/sessions that fail their JavaScript fingerprinting.
Not sure how reliable these numbers are as I have a Go library specifically for passing their fingerprinting in pure Go (No JS) and it has been working for a while without much change
(My personal favorite was being blocked by cloudflare, and then getting an email from cloudflare telling me how great they are because they block so many (alleged) bots)
I mean, it would not surprise me at all if bot traffic on Youtube is larger as in bytes/s of video, than actual visitors.
Remember that Google has no interest in disclosing how many bot views there are.
There was this news article about Spotify banning some musicians accounts due to listen farming.
https://www.musicbusinessworldwide.com/great-big-spotify-sca...
Since the payouts are not tied to the individual watcher or listener, scams run rampant. I guess the temptation to manipulate the money flow is way higher than having a "fair" distribution of each subscribers money to the artists.
Yt definitely makes an effort to prevent that. You can read up on how botguard's virtual machine works in some older writeups.
I watched today a couple hours of Youtube videos. How many bots are needed to match that much traffic?
if we're talking about requests it makes sense. I just have to open my apache/nginx logs to see how many bots are attempting to get in.
However I agree, on its face 73% sound ridiculous.
Producing a RSS feed and clean scrapable page is always a good thing to do to boost your page recognition and popularity.
You data does not have to be scrapable, but you page should always be bot friendly.
Toll Fraud is rampant lately. I can tell because people have attempted it on nearly every consumer-facing project I've worked on. Additionally, the ill-advised "text the app to your phone" gimmicks have disappeared (fraudsters have cleaned up on direct, public endpoints to send SMS).
Shameless plug, but I've written a little about Toll Fraud and how to prevent it [0]
[0] https://koptional.com/article/how-to-stop-twilio-toll-fraud/
https://www.arkoselabs.com/wp-content/uploads/Breaking-Bad-B...
It's also stupid and massively wastes computation and traffic on both ends.
Then people would think "but then I'd have to pay money to offer a free API", but they're already doing that in a more expensive way via the web interface anyway.
Legitimate scrapers, maybe. Everyone else does it to circumvent the API limitations, by posing as real traffic. APIs imply API keys which can be traced and banned.
Same thing goes for email. I had one particular domain I managed spam filtering on and for every 10,000 messages we received, 1 was delivered to users. The amount of crap requests on the internet is insane.
My low-traffic personal projects get 90%+ malicious traffic.
Many sites intentionally make scraping as painful as possible because they don't want anyone scraping them.
Tools like this help: https://github.com/mitchellkrogza/nginx-ultimate-bad-bot-blo...
Not to mention that user agents are easily spoofable.
Anyone malicious who knows what they're doing will just send you a statistically common user agent string.
Under most http libraries this is a one-liner:
headers = {'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/119.0.0.0 Safari/537.36 Edg/119.0.0.0'}We analyzed tens of billions of sessions worldwide across industries, [...]
I think they mean: 73% of all commercial sessions are malicious.
Very sensationalist, at the end of the day they are selling something
You mean like wifi?
its kinda expensive and dangerously addictive
the resolution and refresh rate are amazing
When Netflix was a cheap stock but it made up a significant part of internet traffic, I assumed it could be the next big thing.
So how do you invest in "bad bots", whatever this is supposed to mean. Honestly, I don't trust the numbers.