What do these crawlers gather? Just make this data accessible via API calls or direct database download, like Wikipedia did (https://en.wikipedia.org/wiki/Wikipedia:Database_download).
Even wikipedia begged for those damn bot about stopping doing this, the data is already accessible in archive here.