Also, what a piece of zero-trust shit the web is becoming thanks to a couple of shit heads who really need to extract monetary value out of everything. Even if this non-solution were to work, the prospect of putting every website behind Cloudsnare is not a good one anyway.
What the web needs right now, to be honest, is machetes. In ample quantity. Tell me who's running that crawler that is bothering you and I will put them to the sword. They won't even need to present a JWK in the header.
The standard response to a crawler is a 402 Payment Required response, probably as a result of an aggressive bot detection.
So essentially, it's turning a site's entire content into an API: Either sign up for an API key or get blocked.
The question remains though how well they will be able to distinguish bot traffic from humans - also, will they make an exception for search engines?
> Each time an AI crawler requests content, they either present payment intent via request headers for successful access (HTTP response code 200), or receive a 402 Payment Required response with pricing.
I don't see how it would make sense otherwise, as the requirements for crawlers include applying for a registration with Cloudflare.
Who in their right mind would jump through registration hoops only so they can not access a site? This wouldn't even keep away the crawlers that are operating today.
I agree there has to be some way to distinguish crawlers from regular users, but the only way I can see how this could be done is with bot detection algorithms.
...which are imperfect and will likely flag some legitimate human users as bots. So yes, this will probably leading to web browsing becoming even more unpleasant.
The idea behind the headers is to allow bots to bypass automatic bot filtering, not blockade all regular traffic. In other words:
- we block bots (the website owner can configure how aggressively we block) - unless they say they're from an AI crawler we've vetted, as attested by the signature headers - in which case we let them pay - and then they get to access the content
(Disclosure: I wrote the web bot auth implementation Cloudflare uses for pay per crawl)
So I'm cautiously optimistic. Well, I suppose pessimistic too: if this works what this will mean is that all contents will end up moving into big player hosting like CF.