Let's not conflate crawlers with the traffic that bot protection services block. A crawler that respects robots.txt is a good internet citizen and can provide a vital service.
This Web Search API, unlike an AI crawler, only fetches periodically. It feels like a step in the right direction for managing resource strain across the internet. If only the LLM giants could do something similar.
But that's meaningless because 99% of AI crawlers are "bad bots" which ignore robots.txt and use domestic IPs to circumvent blocks.
Some other reports of this: https://github.com/TecharoHQ/anubis/issues/1565
I ended up just fully closing connections with no response from these assholes on any URL.
This comment will be flagged because it challenges conventional wisdom.