I do a similar thing with web crawlers that do not respect the robots.txt
https://github.com/cl-test-grid/cl-test-grid/blob/873b2fa978...
I don't know if this snippet is really effective, can be improved a little, especially that I noticed a couple of new crawlers that ignore `User-agent: * Disallow: /path` in robots.txt, and do not fix that even after reported.