It would be even better, if the web would be treated, at least for parts, as a digital library and non-profit organizations would recognize the value of access to such a resource and provide it (just as maintenance for roads or public schools).
It would be even better, if the web would be treated, at least for parts, as a digital library and non-profit organizations would recognize the value of access to such a resource and provide it (just as maintenance for roads or public schools).
We probably won't expose the underlying crawl data ourselves but being able to reference it just like [1] does is indeed important.
I can think of many other use cases as well, where a product needs to built from larger, but very selected set of raw input or pages.
BTW: Thanks for Hackday Paris 2011! Loved that event and venue :)
One problem with what you'd like to do is the pagination. Because we have to send queries to all the shards and then re-rank them, it becomes increasingly hard (and useless for most users) to build pages with p > ~20 (which is why all search engines heavily limit their pagination). So our main infrastructure definitely won't be optimized for that :/
However sending the top ~500 pages with a keyword + a domain filter should be doable pretty easily!
You do realize that you are talking about potentially a LOT of data?
To give you an example: The word "work" occurs on about 4% of all web-pages. So even if there were only about 2bn pages in an index, that would mean 80 million matching pages. Even if you only need their URLs that would be about 2.4gb of data assuming an average URL length of 30 bytes. Ok, compression can make that smaller, but still...
It would also mean that the server would need to make 80 million random reads to get the URLs. Even with SSDs that would take some time. Hmm, actually in this case it may be faster to just read all URL-data sequentially, than doing random reads. But in both cases we would be talking about minutes needed to get all that data from disk.
I currently have a search-index with about 1.2bn pages - I expect to reach 2bn pages by mid-May - that could be used to get the kind of data you need. But not in a realtime API. Not that amount of result-data.
> that could be used to get the kind of data you need.
Cool. Would you be interested in sharing or exchanging data?
What would be more useful to you, the raw data - meaning for each page a list of the keywords on it - or the reverse-word-index?
Raw-data may be better for batch-processing or running multiple queries at the same time.
My crawler currently outputs about 40-45gb of raw-data per day (about 30 million pages). Full crawl will be 2bn pages, updated every 2-3 months.
The reverse-word-index would be about 18gb per day for the same number of pages.
Reverse-word-index is already compressed, raw-data isn't.
There is a small problem with the crawl though, as it does not always handle non-ascii characters on pages correctly. I'm working on that.
BTW: I also currently have a list of about 8.5bn URLs from the crawl. About 600gb uncompressed. These are the links on the crawled pages. Obviously not all of those will end up being crawled.
Worst case we may build our own specialized crawler just for this purpose, but it would be nice if there was a useful search engine API we could leverage. And, of course, we'd be happy to pay for access to such an API.