I've done this before for large scraping projects. I find the datacenter the target website is hosted in, then get a dedicated server right next to it. I've never gotten better performance.
I spent way too much time pruning stupid crap such as slashdot and started to learn this 'Bayesian classifier' thing.
Your idea is much better.
(Don't downvote him, it's a valid question)
300TB is quite a lot, even today.
I offered to help them set up an API instead of scraping, but they decided scraping was easier in the short term.