HN
Hacker News
Top
New
Best
Ask
Show
Jobs
Comment by d4rkn0d3z | Hacker News Reader
Full thread
d4rkn0d3z
·
"When we detect unauthorized crawling..."
How did you do that?
View on HN
mog_dev
·
Simple, you add the trapped paths to robots.txt Well behaved robots will not crawl them.
d4rkn0d3z
·
Nifty trick.
CaffeineLD50
·
And the misbehaved bots follow the path right into the pit and then...the Void of Infinite AI Abyss.
ccgreg
·
Cloudflare's documentation says that Labyrinth is not based on robots.txt.
Epskampie
·
In line 1 of of the linked page: "waste the resources of AI Crawlers and other bots that don’t respect “no crawl” directives".
ccgreg
·
Does that indicate the robots.txt is how "no crawl" is indicated? robots.txt doesn't have "no crawl", it has allow and disallow.
hombre_fatal
·
Just consider how you click around HN versus how your crawler would behave if you wanted to crawl every page of HN starting from the homepage.
Reply on news.ycombinator.com