I'm sympathetic to this. I built a search engine for my senior project and my half baked scraper ended up taking down duke law's site during their registration period. Ended up getting a not so kindly worded email from them, but honestly this wasn't an especially hard problem to solve. All of my traffic was coming from the cluster that was on my university's subnet, it wouldn't have been that hard to for them to IP address timeouts when my crawler started scraping thousands of pages a second on their site. Not to victim blame, this was totally my fault, but I was a bit surprised that they hadn't experienced this before with how much automated scraping goes on.