61.5% of Web Traffic Is Not Human
theatlantic.com
theatlantic.com
I've pretty much stopped looking at "number of impressions" when purchasing advertising too. It's always untrue. A recent blog I was going to advertise on said it had 35,000 unique visitors a month. But yet just a few weeks ago they had a "free give away contest for a jewelry brand (no purchase necessary, international readers allowed, free to apply, no shipping costs)" and only 100 people applied. That pretty much said it right there.
One more example: A site I was on allowed unregistered visitors to give their entire collection of corporate logos a rating between 1 and 10 stars. And they used radio buttons to vote for the numbers 1-10. As soon as you select a radio button it auto-submits and gives the logo that rating. EVERY SINGLE LOGO had a massive number of "1 star" ratings because all the bots would see a radio button, check it, and trigger the vote. All the logos on their site had a rating of 1.
The internet conditioned me to ignore such claims as too-good-to-be-true. I would expect loads of people to treat it like "you're the 1000000th visitor! you won!" and other "free stuff here (if you register and fill out hundreds of forms)" offers.
The report [1] does not include any non-http traffic, which excludes all heavy content like torrent, Netflix, music streaming, etc.
[1] http://www.incapsula.com/the-incapsula-blog/item/820-bot-tra...
I, for one, am glad that human trafficking over the web is down to only 38.5%.
Also, I wonder what the definition for "traffic" is within the context of this research.
"For the purpose of this report we observed 1.45 Billion bot visits, which occurred over a 90 day period. The data was collected from a group of 20,000 sites on Incapsula’s network"
So, in the first place, not necessarily representative of the web generally. And it does not say how they detect bots. Of course some declare themselves [2]; otherwise I guess one would have to rely on captchas [3], which are known to be of limited reliability and subject to an arms race.
1. http://www.incapsula.com/the-incapsula-blog/item/820-bot-tra...
2. kristopolous on this page, '(bot|spider|crawl|aggregator)'
3. which Incapsula does, as a condition of access to its summary of its previous similar research.
For me, this seems like a success of the whole teach-everyone-to-code thing; the author's picked up a bit of simple scripting and it's made some things dramatically easier for her. It's certainly not professional-level, but that's normal; for example, many people think programmers should know how to write, but few think they should all be able to churn out high-caliber journalism. Is this a failure of setting expectations, or is that really where some people are setting the bar?
I had no idea what he was talking about.
Per spider.io, bots are ripping off advertisers left and right. That being said, scared advertisers are the target market for their products.
That means that bots who declare themselves make about 13.5% of requests to this site. The robots.txt has no restrictions.
I don't know if that's helpful.
Bots that don't trigger js, however, are more a drain on bandwidth than anything else.
My company right now won't purchase RTB ads that display outside of walled gardens so that we can 1. manually verify that the ads are showing. 2. Minmize the number of bots seeing the ads (presumably it's hard for bots to infiltrate Facebook or the WSJ than Buzzfeed).