> - Identification of well known legitimate bots;
What about non-well-known legitimate bots? If I run my own web crawler, am I at risk of falling into the tarpit (and having my IP address reported)?
> - Identification of well known legitimate bots;
What about non-well-known legitimate bots? If I run my own web crawler, am I at risk of falling into the tarpit (and having my IP address reported)?
In my dream world, Cloudflare would make this a two way street: you get protection against malicious bots, but you have to allow equitable treatment of all well behaved search engines (i.e. you can use robots.txt to ban lesser search engines, but doing that requires you to ban Google as well).
Disclaimer: I have no current economic interest, but kinda want to run a personal search engine.
> Identification of well known legitimate bots;
> Hand written rules for simple bots that, however simple, get used day in, day out;
> Our Bot Activity Detector model that spots the behavior of bots based on past traffic and blocks them; and
> Our Trusted Client model that spots whether an HTTP User-Agent is what it says it is.
Use CommonCrawl https://commoncrawl.org/
Then, after you can use that corpus and provide useful results and have people coming to your search site, then you can have a discussion with CloudFlare that you will be creating your own crawler.
A big part of the net neutrality fight was "allowing the little guy a chance to play with the big boys".
Part of the service says:- for Bandwidth Alliance partners, we’re going to hand the IP of the bot to the partner and get the bot kicked offline... If the infrastructure provider hosting the bot is part of the Bandwidth Alliance, we’ll share the bot’s IP address so they can shutdown the bot completely. The Bandwidth Alliance allows us to reduce transit costs with partners and, with this launch, also helps us work together with them to make the Internet safer for legitimate users.
My reading of that is if CF decide your IP is bad, they can leverage the providers Bandwidth Alliance status to shutdown the providers customer. What if CF's systems misfire? Will there be a grace period? Will there be an appeals process? Will CF compensate anyone effected and have their hosting withdrawn from a misfire?
No one likes bad bots, but I'm feeling more and more uneasy allowing CF decide which is which.
They bring real visitors and customers.
Random bots running on AWS or TOR bring me headaches and bigger bills.
if this is to be believed [0] less then 50% of google searches result in a click though to the site, so one could argue that because of things like snippets in search results Google don't always drive traffic to your site.
[0] https://sparktoro.com/blog/less-than-half-of-google-searches...
Respecting robots.txt, using a well-known ua, coming from publicly declared, google owned IP block.
So if you're a small business that wants to build the next Google, you can't? Because obviously, your User Agent won't be well known when you start.
I would love for them to prove me wrong though and be open about such things.
I'm pretty sympathetic to this line of thought, so don't take this as 'you're wrong', but I notice that we would never apply this logic to laws, or environmental standards, or contracts.
You'd never hear someone say, "if we have a clearly defined tax code, that will just make it easier for people to find loopholes."
Where I do hear this argument come up is explicitly in contexts of moderation and abuse policies. And it's something that sounds very reasonable, but it's hard for me to get away from the fact that in most other contexts it sounds problematic to me.
Maybe part of the problem is Cloudflare's scale? Maybe the reason it feels bad to have a police officer pull me over because, "we think you're going too fast" instead of "you went over a posted limit", is because that's a critical infrastructure that I can't avoid.
Cloudflare is a private company, and even if it wasn't, I don't know if it would be big enough presence that I would worry about their policies. But when I hear things like this:
> Our goal is nothing short of making it no longer viable to run a malicious bot on the Internet. And we think, with our scale, we can do exactly that.
That maybe shifts the situation a tiny bit farther away from "moderation policy on a personal blog" towards "policeman pulling me over because I broke a law I didn't know existed."
I dunno. I'm not 100% sure how to feel about it. I do think that Cloudflare should be able to filter traffic however they see fit, but that doesn't mean that every strategy they choose is inherently good, or that it might not be problematic for them to lack transparency about their standards.
Also for fairness - Does Akamai publicize the internals of BotManager? Imperva gives you much transparency into how DistillNetworks works?
Then fraud detection system cannot work ethically.
Can't become a well-known UA if your shutdown before you can make a name for yourself. And a Google IP block doesn't mean legitimate.
I'm not trying to say that Google is bad, but why can't OP's crawler be legitimate?
It does. 2-way DNS lookup is how you verify legit indexers [0] (as ua itself is meaningless).
But what are CF (and others like akamai) using to determine you are a bad bots? My objection is more the threat they will use their position to kick bad bots offline.
Now don't get me wrong, I had bad bots as much as the next person. But what about misfires? They will happen. I hope CF have a plan in place to deal with such cases.
Personally, due to family issues I've been working remote-remote (I'm a remote worker anyways, but I've been working remotely away from my normal office). As such I've been using a 4G Mobile broadband connection and the amount of CloudFlare "We are just checking you are not a bot" hoops I've had to jump though recently has opened my eyes on how much of the net is actually behind CF's network because CF still associate "IP as Identity" and my 4G connection is CG-NAT'ed.
I just fear at times we forget that there is the rest of the world out there, not just our little bubble and the choices we make have a huge impact on the usability of the net for others.
EDIT: LOL.. Just checked your profile. The Reg is one of the sites that doesn't like my 4G connection and as such I visit less. What are the chances. (Though you only make me sit there and wait to access the content....)
Sorry if that's not the answer you wanted, but that's how I feel after seeing how the other side of the net is treated. IF the urge takes me, I'll help, but I dunno when that urge will take me though.
<EDIT>Fun Fact: My public IP address appears to have changed in the time it took me to write this comment. Atleast to the reg, How my ISP handles routing is another black box upon itself.</EDIT>
Don't worry, their proprietary AI will figure it all out.