Ahrefs could probably save another few hundred million if they didn't repeatedly visit the same links over and over indefinitely to find the same error codes or a binary file that they surely don't care about. I see them in my logs for my podcast hosting service, hitting the same 404s for months and months. They end up hitting audio files and downloading (or attempting to download) many gigabytes of content each day. They don't do anything with audio! As soon as they get the HTTP headers, they could say "oops, I don't care about this" and disconnect. And in many cases, they got those URLs from the enclosure tags of RSS feeds... I'm not sure what they were expecting!
In the cloud, not in the cloud, frankly it seems to me like their business would be far more efficient if they tuned what their crawler actually crawled.