We use a combination of polite massive-scale web crawling and user contribution. We had an in-house web crawler for years until recently when we released it as a stand-alone service.
58 karma · joined August 7, 2014
https://outpan.mixnode.com
We use a combination of polite massive-scale web crawling and user contribution. We had an in-house web crawler for years until recently when we released it as a stand-alone service.
I expect this to work on a larger key space as well. It is interesting to see how the expansion works out in terms of usage patterns.
I will send you an email :)
As for curation, it is intended to provide examples of what key, attr and values are regardless of their content. This is an experimental feature and might be removed/tweaked...
We did a benchmark on Nutch and couldn't really pass the 10-14 M(B)ps on a $1200/month machine. Even though we hired a professional to optimize the setup. The same is roughly true about Heritrix.
Just wondering if there is something missing in his setup, such as domain/ip rate limiting.
The problem is the huge contrast with https://www.quora.com/How-much-would-it-cost-to-crawl-1-bill...
Even taking into account the drop in prices on AWS. Also, if you take a quick look at companies that provide such services the prices are orders of magnitude higher than deusu's costs.
For the life of me I can't figure out how you manage to crawl over a billion web pages (even in 2-3 months), index the data and run the server with €300 per month. Especially the crawler part...