As others mentioned, you can block their user-agent. That will stop most crawling. They have some bots that use browser agents to see if you serve different content to non-google bots.
To block those, you can feed their IP's to haproxy or iptables ipset.
fetch_google()
{
for line in $(dig +short txt _cloud-netblocks.googleusercontent.com | tr " " "\n" | grep include | cut -f 2 -d :)
do
dig +short txt "${line}"
done | tr " " "\n" | grep ip4 | cut -f 2 -d : | sort -n | uniq | xz -9ecv > _GOOGLE.netset.xz
};
In HAProxy, you could include the decompressed version of that IP list with something like
acl BOTS src -f bot.txt
Then either redirect, reject, silent-drop.