Exactly this. It's no different from a bot pretending to be Googlebot. I've tried reporting abusive IPs to various foreign hosts, but nothing every comes to it. I've settled for just blacklisting excessively abusive IP ranges.
Years ago, I built a sort of camo-proxy like, in Elixir. It was doing full passthrough of the User-Agent because the project that I originally built it for needed it for the upstreams (i don't remember why). Anyway, I ended up pulling it into pleroma -- a couple of months later, we had been informed by some instance owners that google itself were sending them DMCA notices, because it was proxying googlebot's request, with its user-agent!
What is your way of detecting them? Just cat your way through your logs?
fail2ban
Almost lol: grep, sort and uniq. If I notice someone is hammering my employer's ecommerce site, I'll block them. It isn't required often so I've been reluctant to spend the time setting up fail2ban.
Is it a multi-server setup? If so, do you ssh into each machine and look at the logs?
It's 3 servers so it's not too much hassle to ssh into them and check it manually.
Google (and other “legitimate” scrapers) publish the ip ranges they crawl from, anyone claiming to be googlebot (or whatever) but not in the ip range can safely be black holed.