Putting aside legal issues, you don't have any moral problems doing volumes of scraping content that is not yours?
(Don't downvote him, it's a valid question)
300TB is quite a lot, even today.
I spent way too much time pruning stupid crap such as slashdot and started to learn this 'Bayesian classifier' thing.
Your idea is much better.