One way is we send crawl_task to different to N random nodes and accept one that most similar?
another way could be build messy network to solve messy problem. What we do is build reputation bashed graph network and you accept index from nodes you trust. so people will start un following misbehaving nodes. there is not universalroot view of network instead its dynamic and different from prospective of each node. or it could have one root view if we store reputation data in bchain and with some type quadratic voting to modify the chain. ?
yeah Bitcoin showed us way to build mathematically secure system without any trusted party but it could do that cz problem it was solving is mathematically provable. Problem like collecting indexing crawl data you have to trust somebody.
If the miner, they would be able to deliver any crap so content deliveries would have to be judged in some way and awarded differently.
If the pool provides addresses to crawl, the miner could be given crafted/dedicated URIs from time to time and lack of delivery of Proof Of Crawl could result in a penalty chosen in a way rendering "cheating" unprofitable. But then fresh URIs have to come from somewhere.
The second associated problem is how would one prevent them from appropriating the work of others that would just re-sign it.
One way would be to allow the worker to introduce a few voluntary errors but have a secret joker that allow him to pass the challenge of a failed verification.
One alternative way is based on data malleability. The worker pick a secret one way function and compute is F( data + secretFunction(data,epsilon) ) ~ F(data) and publish the values of the secretFunction(data,epsilon) but not the secretFunction. Only someone with knowledge of the secretFunction can make a claim on the work done. If there is a challenge only the real worker will be able to publish the secret of the secretFunction (Or use some zero knowledge proof to convince you they know it).
A web page isn't an immutable piece of text. It can change on every visit and it can sometimes returns errors.
Once a reference snapshot has been crawled, the indexing task is more easily verifiable.
The crawling task is harder to verify, because external website could lie to the crawler. So the sensible thing to do is have multiple people crawl the same site and compare their results. Every crawler will publish its snapshots (which may contain some errors or not), and then that's the job of the indexer to combine multiple snapshots of various crawler and filter the errors out and do the de-duplication.
The crawling task is less necessary now than it was a few years ago, because there is already plenty of available data. Also most of the valuable data is locked in walled garden, and companies like Cloudflare make the crawling difficult for the rest of the fat tail. So it's better to only have data submitted to you, and outsource the crawling.
If every contributor maintained their own index, then you could reward contributors based on how many hits their index generated.
This would open up the possibility of people maintaining indices for specialized topics that they were experts in, and give the federated search engine a shot at taking on Google.
For most people, the cost of creating and maintaining a website is high. This is why products like Wix and Squarespace exist (and are not cheap).
I am thinking a simple dashboard where anyone could go and curate a list of content they find useful. They could share this with the world.
The interface should be so simple that my parents could use it - and they aren't going to be putting up websites anytime soon.
(But of course it wasn't a volunteer public benefit effort like you describe.)
I also have been wondering how this would play out with some kind of decentralized indexes. The nodes could automatically cluster with other nodes of users sharing the same interests, using some notion of distances between query distributions. The caching and crawling tasks could then be distributed between neighbors.
Crawling is also not as resource consuming as you might think. Sure you can distribute it, but there isn't a huge benefit to this.