edit / full disclosure: I work at google but nothing related to search
I think OP's point is, assume you only have 10/100 terabytes of space and limited compute ability - how would you approach the problem? I assume 90% of google's searches probably come from less than 1% of their total index, not to mention that Google is also keeping full cached versions of the whole website including images.
The challenge isn't to index the entire web, it's to index the useful parts of it, and I think an index covering most of the useful web can be seeded quite easily with some community effort.
The challenge then moves to the curation, but it's no longer infeasible.
It's got a fairly small index, but yeah, it's not particularly hardware-hungry.
I'm confident 100mn is doable with the current code, maybe .5bn if I did some additional space optimization. There are some low hanging fruit that seem very promising. Sorted integers are highly compressable, and right now I'm not doing that at all.
> What is your current max rps?
It depends on the complexity of the request, and repeated retrievals are cached, so I'm not even sure there is a good answer to this.
I honestly don't know what the actual limit is, all I know is it dealt with 2 QPS without affecting response times. But 2 QPS for a search engine is actually kind of a lot. Most people don't actually search that much. Like you get a few queries per day. Put it this way: 2 QPS is what you'd expect if you had around a million regular users. That's not half bad for consumer hardware.
How is it to be funded?
You'd have to find a way to verify reputation to make sure no bad actors could contribute.
Not knowing their implementation details, I’m guessing it could be doable without reinventing much. An oracle could dispatch a P2P archive job to a pool of clients randomly assigned tasks, with both the first to archive and the first to validate being recognized by the swarm somehow, with periodic re-archiving and re-verification, rate adjusted by popularity of site and of search keywords.
my dream, is a distributed/p2p index. each browser contribute to storing part of the overall index, and handle queries coming from other users so that how to fund huge data centers never become a question.
> is the size of the common Web already way too large to play catch up against google/bing at this point?
Probably. But I would prefer a search engine that didn't search the whole web. I would prefer a search engine that searched the sites related to fields that I'm interested in.So I would pay for or donate to a search engine that provided me good results in e.g. software development. They could add additional fields as demand warrants, so long as quality as maintained. I would even like to see a faceting feature, so I could search for e.g. Matrix and get results on the mathematical concept when need be without having to wade through movie review or fiddle with magic search keywords.
> should indexing skip a blog page because a the author usually write about algorithm but here goes on and on about business while sporadically mentioning algorithmic technical aspects?
Yes, because that author's pages wouldn't even be fetched at the point of development that we are discussing. > and what if you want to search about pottery that Sunday morning, turn to Google?
Yes, Google still exists. Why not?I'm somewhat fine with very commercial search engines when searching for something within a field I'm familiar with, e.g I if search for "kubernetes configmap best practice", skimming infomercial sites, ads, junk is trivial. if I search for "baby milk safety" I have no clue how to digest the results list, I'm totally unfamiliar with the media outlets catering to parenting, nutrition or food health, I know absolutely no authors in the field, no renowned media outlets for tips, not even brand reputations to make a somewhat informed decision on what to skim and what to read with attention.
But I don't disagree an engine focusing on specific fields provides great value. and perhaps that's where it should start to have a chạnce to win.