47% of all internet traffic came from bots in 2022?
securitymagazine.com
securitymagazine.com
It used to be that the cost of scraping came with the benefit of being search engine listed which drove traffic, but that feels less true than it used to (for a lot of reasons).
But now the cost of scraping doesn't feel in the favour of a website.
Scraping and bots are for search engines listing, technology tests / experiments, advert / audience measurements, brand protection, IP tracking, copyright enforcement, screenshots for links on other websites (i.e. Facebook), Pinterest linkbacks, training of LLMs (my hypothesis on Bing's massive increase), spam, etc, etc.
With the search engine value lowered by less traffic, yet a solid community still growing via word of mouth... the rest of those things offer no value to me or the community. So I asked the community, what do you want to do here? Leave them all? Ban some? Ban all? Some midway thing?
Almost unanimously the community (who fund the costs by donations, and at least 30% of all traffic and costs were known to be associated to bots) chose to block every bot.
So that's what we've done.
We've blocked every major hosting and Cloud ASN, or put a challenge up to the few known to be proxies (i.e. Google Data Saver), and we've blocked hundreds of bot user agents, we've blocked requests where no Accept header was present where it should be, we've blocked TLS ciphers that aren't modern web browsers — I looked at requests by Python, Go, Curl, Wget, etc... and blocked everything that obviously differed from a valid browser.
In the end we blocked about 40% of our traffic, and so far not a single real human has said (and it's a tight-knit but large community with lots of ways of contacting me) that they've had any issue at all.
We appear to have reduced our traffic and associated costs, with no loss to us at all.
When the day comes that I shutter another I'll ask the active members at the time what they want to happen to their data. They may desire to leave it as a resource, they may want to delete it, if there's a clear majority in the decision I'll go with whatever they desire. I value the choice of those whose data it is, who contributed to creating it, over anyone else's hypothetical needs.
You might not have that chance; unless you have a co-admin with full access to everything, the reason for the forum to shut down might be because you're no longer there.
Respectfully, on most forums, I don't care about the community, I care about the content, that's why I'm there, to have discourse and generate meaningful value in the form of knowledge. If someone passes, yes, that sucks, but that's life, we're all snuffing it at some point. However, the world carries on spinning, and that information should continue to be available, especially if the forum is for a niche and frequently generates useful information.
If a forum is becoming "toxic" then that sounds like a moderation problem.
Last time I checked, Archive.org et al weren't a "corpo behemoth", but consuming server resources is exactly what a normal user does.
Site owners should get with the times and serve up cached static pages to users who aren't logged in. Even then, they should be serving up cached static pages and rebuilding cache for relevant pages when someone posts new content when it comes to forums. Not being able to handle a few crawlers is an administration problem. Why should the community/public suffer for someone's inability to configure a server appropriately?
My opinion is that "screw crawlers and scrapers" is a valid opinion. If i'm hosting a playground, it's my playground and my rules. If you want to play elsewhere, please do. If you want to preserve data, please do, but not at my expense. Disagree with that? feel free to, but don't think that you are somehow in the right, because if you go with this shit to court, you will be laughed out of the door.
If people wished for their content to be available to all forever they'd run a blog and pay to ensure it is available, and would proactively seek to get it archived.
People on forums aren't doing that, and the data of any given individual is a contextless collection of semi-random mumblings on different topics because without the fullness of a conversation involving others none of it makes sense.
It is within that context that a forum admin can decide what to do, they have been granted right (by T&C) to the collection of all the forum members comments which restores the context and gives meaning to the content. Every individual on the forums I operate can obtain their own data, but it would be meaningless by itself.
As the operator of the collection of content I get to determine what best to do with that, and sometimes that may be to delete it all. Sometimes that may be to seek to archive it. And on this occasion it is to treat this knowledge as having valuable to those already participating in the community and to not be shared beyond that.
Elsewhere you said this:
> Call it what most forums are: an ad-supported business. People generate content for the owner for free because they too derive value from the information that others share. The middleman is just a middleman
But the 300+ forums I run have no adverts, they are not a business, they are non-profit. Their value (if you want to measure everything in a capitalist way) is social, to help those in the community.
The purpose of the forums I run isn't to expand the sum of human knowledge, or to make myself personally wealthy of the back of the efforts of others, the purpose is to help be a remedy to adult loneliness by connecting people by their shared interests in geographically small areas such that it builds relationships and forms bonds.
Yes there is a hell of a lot of expertise captured here around those interests... but no-one has any inherent right to it.
That forum was around a music band in the UK, and the audience of the forum turned out to be lower than expected - University age. They were emotionally immature, over-shared online, slept with each other, had relationships and break-ups... all in public. The music forum did have lots of music info on it, but it was intertwined with a lot of very highly personal information posted at a time when a reasonable expectation of the internet was ephemerality.
It was totally right to protect the individuals future selves from their past selves, and I would delete again.
This idea that all information must be preserved for forever is also at odds with privacy. See, e.g., the right to be forgotten.
What if we made archivism more fashionable?..
Do you maintain a freely-available repository of all of your knowledge and experience, in case someone else wants to consult it one day?
While the openness of the (now-ending) early days of the internet was liberating and allowed knowledge sharing on an unprecedented scale, the downside is the huge devaluing of that knowledge and skills.
No, but I try to maintain some of it, and I see the value in maintaining as much of it as possible.
The real value of knowledge doesn't change if you duplicate it or make it widely available. On the long term, blocking access and rent seeking doesn't create value, it destroys it. It seems useful for the individual who wants to pay their bills or for the one with insatiable greed but in the end it will makes us stupid.
For example: I would like a high quality UV-B lamp that isn't INSANELY expensive. They are pretty ordinary lamps but developing the coating is very expensive. The work has been done tho, lots of times, over and over again. Most results are just bad.
About 35% of the US and about 1 billion globally have vitamin D deficiency, 50% has an insufficiency: Fatigue, Not sleeping well, Bone pain or achiness, Depression or feelings of sadness, Hair loss, Muscle weakness, Loss of appetite, Getting sick more easily, etc
Great loss of economic productivity or more opportunity for me? You decide!
Have we really come to a phase of internet use, where everytime you see something, you have to manually save it, and on every post (even here or on reddit, facebook r wherever) a link is not good enough, but you have to copy-paste the whole block of text just to make it a bit future-proof?
Call it what most forums are: an ad-supported business. People generate content for the owner for free because they too derive value from the information that others share. The middleman is just a middleman.
To not allow that content to be indexed/cached/archived/mirrored whilst making money off of it is pretty scummy in the long-term. There's tons of forums I used to visit whose information is now forever lost, that included a lot of very useful programs for niche bits of kit, which is now otherwise very expensive e-waste.
Because otherwise their work was wasted.
> Do you maintain a freely-available repository of all of your knowledge and experience, in case someone else wants to consult it one day?
I would if I could, I’ve already contributed what knowledge, bandwidth, and money I can to the Internet Archive. What about you?
> While the openness of the (now-ending) early days of the internet was liberating and allowed knowledge sharing on an unprecedented scale, the downside is the huge devaluing of that knowledge and skills.
I cannot even process how wrong this is. Objectively the preservation of knowledge and skills is a good thing, and you cannot devalue knowledge, which is itself priceless.
This argument really makes no sense. If I tell Bob how to fix his transmission down at the local diner, but nobody records the conversion, that wasn't wasted work. But fixed his transmission: mission accomplished.
I've also noticed that my cache hit rate is extraordinary now, which I assume is because humans read recent stuff and bots read the long-tail of old stuff.
If you use Cloudflare, turn off their anti-bot stuff. It is far more efficient to let them just serve bots from the cache than having scrapers use tricks to bypass them and go directly to your origin server.
Bots end running around CF would guarantee turning on authenticated origin pull.
40% is 40%. Maybe 40% of their cost isn't enough to warrant whatever time these efforts cost them, but for many people out there it will be.
If you are running on i.e. EC2 and RDS instances, you’re not saving anything by using 40% less of the CPU, unless you can actually downsize the instance as a result. Read-only traffic is also not that hard to scale out, but with forums etc, you can be stuck with some legacy systems for sure.
But we'll, not with the bing bot. It ignored my timeouts and queried hundreds of thousands, for him identical, pages every single week. Not one connection, not two or three but about 10 IPs hammering my servers at once. No second between request, not even pausing when the server is going down. Something even 'bad bots' usually do.
I assumed it was just any bot calling itself Bing. But no, it was their IP ranges.
I blocked nearly all of their IPs. Which appears to be the only way to make sure it doesn't ddos me again. Bing is like 1% of my traffic, not even worth the hustle.
Otherwise, if someone alt+clicks a bunch of a category's threads as they look interesting then you're going to have a bad time.
The bot was never ment to query the same pages thousands of times. These pages were identical to him. There were already bot specific rules programmed in.
This website already is heavily cached and optimized. Even thought it only has 16 database connections there maybe is one timeout every few months. Users usually don't open tabs much faster than the short requests take, only bing does.
Really the time that went into optimizing it makes me kinda sad when someone questions that is was a effort Vs payoff thing. The bot ignored all rules I gave him and barely brought any benefit for me in terms of traffic. There is no payoff here, only effort.
How did you do that ? Years ago there was a script to block AWS and I made substantial savings by running it.
With that context, I used bgp.he.net to look up the big ones I know and then wrote the rules.
The paid databases comes with AS type (hosting, ISP, business etc.) and we have a VPN detection database as well.
[0] https://ipinfo.io/developers/ip-to-country-asn-database
[1] https://community.ipinfo.io/t/filtering-asn-database/395
The browser has the goal of being light and downloading only the text and pictures (no css, no js). So we have the same goal here.
So I am tired and only make a few posts every few days now on HN. I am sure while my activity has dropped, the bots are getting more active so nothing is lost. Maybe some quality and how the traffic share look like, but I don't know.
Both of these things will happen (old web getting spammed, old web being distilled and crystalized), and the future will be weird and unpredictable to us now.
> the future will be weird and unpredictable to us now
I'm going to play devil's advocate for those people who always drop by saying LLM is pretty much like a human and human is pretty much like an LLM anyway, and say it would be no different to now
An internet saturated by bots is like reading reviews on Amazon without pictures. Pointless, intentionally misleading, and often confidently wrong.
Also, I think voting and moderation will upvote and downvote the AI generated comments in such a way that they don't poison Redditch as a training data source.
What I hate the most about today's internet is how search engines allow blatant scrapers to feature so high in search results. So many times I Google for something to find Stack Overflow as the main search hits, and right next to it there are a couple of sites that copied Stack Overflow's questions verbatim. Once I googled for FLOSS projects I had on GitHub and lo and behold there were half a dozen obscure sites that also claim to host my project, with everything copied verbatim from git repo to project descriptions.
you can already see where this is going. Sites with 6 pages of boilerplate that sounds like an 6th grader padded an essay around a 2 word answer they've scraped from somewhere else. Worst of the 2 words of content aren't even all that accurate most of the time. At least sites that copy the answer verbatim still give you the answer!
The nice thing is the Discord API doesn’t seem to ban you for archiving servers as long as you do so responsibly.
Now the code is a hot mess, sure and I am partially to blame for that but that is kind of besides the point. We don’t need to serve any more than thousands of concurrent users which the website can handle but we have to basically ban Chinese traffic to stay online.
Maybe people at bigger companies already know this but it was a revelation to me how much it takes just to stay alive in production.
I anal and I definitely don’t know what they get out of crawling every single product detail page on our website multiple times a day. Nothing here changes that often. Maybe they have some bad/overzealous code? Are they looking to attack take over our servers to them attack others with our machines? If it is an attack, why use Chinese IP addresses? Why not use their bit farms? If it is legitimate search engine, why not respect robots.txt?
Because there is nothing, you can do about it anyway?
Even if you could proof, it is an attack, do you really would consider sueing some chinese IP adresses?
But I rather suspect, it is just bad crawler code.
There are a handful of bots that mirror/archive multimedia content that is anonymously accessible. There is no way those bots have the storage capacity to mirror even a single pass of all the anonymous content.
What I have seen increasing exponentially is port scanning but that takes up almost no bandwidth. Even the broken scanners that in effect look like an amateur DDoS only utilize about 15kb/s using dozens of CIDR blocks at the same time. That does not even remotely hold a candle to streaming.
The link below [1] is talking about Netflix as a percentage of internet downstream traffic and this is only Netflix. There are now hundreds of streaming providers and according to Sandvine streaming accounts for 65% of internet traffic. [2] This does not include torrents and other file sharing.
Here [3] are some fun stats. One of them backs up the submission but it isn't clear if they mean requests or bandwidth. Given that Netflix or streaming alone is 65% of the bandwidth that would lead me to believe the issue of this thread is a lack of clarity around bandwidth vs requests. The wording on all of these sites is too Wibbly Wobbly.
Every day, the internet generates more than 2,183,908 tons of CO2 emissions.
Internet traffic statistics show that 51.8% of all traffic is generated by bots, while humans account for only 48.2%. I think they mean requests, not bandwidth.
[1] - https://www.makeuseof.com/tag/how-much-of-the-internets-band...
[2] - https://www.tubefilter.com/2023/01/20/sandvine-video-data-ba...
And I can certainly imagine bots fetching Netflix or Porn or whatever video for personal archival purposes (I use youtube-dl to protect a few videos I really like from the vicissitudes of Google myself).
[1] $9.45/GB at brightdata.com/proxy-types/residential-proxies
Our data show in the first half of 2021 bandwidth traffic was dominated by streaming video, accounting for 53.72% of overall traffic, with YouTube, Netflix, and Facebook video in the top three.
https://www.sandvine.com/hubfs/Sandvine_Redesign_2019/Downlo...
And looks like Sandvine is very credible as countries relies upon it to censor internet in their countries.
American Technology Is Used to Censor the Web From Algeria to Uzbekistan
https://www.bloomberg.com/news/articles/2020-10-08/sandvine-...
Sandvine's primary business is traffic shaping and bandwidth management for cellular networks, shipboard networks, and other places where you need to do intelligent QoS. They are the reason Google Maps still works well on your cellular connection while you pass a house using cellular broadband to download a torrent.
The fact that some people use it to censor is an abuse of technology, not the intention of it.
When cost to broadcast for anyone on the net falls to 0 everything turns to shit.
I'm guessing it's half of "connections", including DDoS attacks. But even then, I wonder how reliable their methodology is. Like are they including port scanning here, when a connection isn't even made?
> Of all internet traffic in 2022, 47.4% was automated traffic, also commonly referred to as bots. [...] Of that automated traffic, 30.2% were bad bots, a 2.5% increase from 27.7% in 2021
This is a bit misleading, according to the accompanying pie chart, 30.2% of all traffic were bad bots, not 30% of the 47.4%.
What is sorely lacking (from a quick skim of the PDF) is a detailed description of how the data was measured, what protocol it includes, what the error margins are etc.
[1] https://www.imperva.com/resources/reports/2023-Imperva-Bad-B...
47% of website requests might be more accurate.
Also, I'd be surprised if less than 50% of SMTP requests were malicious in some way (spam, scams, phishing, ...).
They were ostensibly Chinese bots, and they overloaded the hosting, not the bandwith, just requests
2 million queries per day are confirmed bots, 20k are not.
The PDF report isn't explicit, but since their assertion is "based on data collected from the company’s global network throughout 2022, which includes 6 trillion blocked bad bot requests", it means the "internet traffic" is measured in number of HTTP requests. As noted in other comments, results would have been very different with network bandwidth.
BTW, I haven't checked if the search bots behaviour has recently changed, but I remember that most of them ignored the directives in robots.txt asking for a slower crawl. And I couldn't find a way to declare that the content almost never changed.
[1] https://developers.google.com/search/docs/crawling-indexing/...
Being a neophyte. How does this work exactly? Is a JS payload that does something bad or an app using web view. Genuinely curious on this one
No. Why would it?
The last time I heard it cited (which, to be fair, was a few years back) the actual number was 33%—for every three dollars spent on digital advertising, a dollar gets lost to fraud. Not too far off.
it isn’t enough to point to real events, they’ve got to produce stuff that rivals the most hyperbolic gartner report ever.
I'm pretty sure GDPR says nothing about cookies that are needed for the site to work, such as session cookies when you're logged in, or cookies to hold settings you set. Am I wrong? Declining on this form sends me to the / root page. Weird.