This reminds me a lot of the (I think) very underrated YaCy project: http://yacy.net/en/index.html
I know there are a ton of technical reasons why this is difficult, but web-scraping and the like always seemed like a problem that could be done fairly elegantly with distributed systems.