Websites, including Reddit, are happy to give all their data to Google bots, but if anyone else tries to crawl them they will get black-listed very fast even if they follow the robots.txt rules.
Websites, including Reddit, are happy to give all their data to Google bots, but if anyone else tries to crawl them they will get black-listed very fast even if they follow the robots.txt rules.
So I tried this myself and quickly realized how hard it was. I stumbled upon thousands of devices that had TCP port 80 (and 443) open so I had to devise various ways of removing these devices.
By the end of my project, I had run out of disk space so many times it was laughable. And tuning the crawling and the resulting mountain of data was daunting and started to affect my day job so I eventually gave up.
A couple of months after my "project", and with enough warnings from IT and our network security folks, our company decided to purchase a couple google 1U servers.