Crawling Billions of Pages: Building Large Scale Crawling Cluster, part 1
engineering.bloomreach.com
engineering.bloomreach.com
Once upon a time I wrote my thesis on building a web crawler. The (tiny) blog post with an embedded preview:
http://blog.marc-seeger.de/2010/12/09/my-thesis-building-blo...
The PDF itself:
http://blog.marc-seeger.de/assets/papers/thesis_seeger-build...
It's mostly a "this is what I learned and the things I had to take into consideration" with a few "this is how you identify a CMS" bits sprinkled into it. These days I would probably change a thing or two, but people told me it's still an entertaining read. (Not a native speaker though, so the English might have some stylistic kinks)
What? select()'s biggest issue is if you have lots of idle connections, which shouldn't be an issue when crawling (you can send more requests while waiting for responses). epoll() is available since 2003. What bottlenecks?
> Currently, more than 60 percent of global internet traffic consists of requests from crawlers or some type of automated Web discovery system.
Where is this number from and how accurate can you make it?
Have a beer and read a more insightful blog about this subject: http://blog.databigbang.com
Also headers and sitemap.XML can tell you how often the pages change.
* Initial pull * Secondary pulls x time later, where x doubles each time, up to a maximum value, y
y is the one that's tricky to define. For us, it's a value computed based on the frequency of update of similar URLs for that domain, the domain as a whole, similar content, and a few other bits and pieces. Essentially, our thinking is that if we can understand how alike any page is to another cluster of pages, we can use their average frequency of update to give reasonably likely initial values for x, and sensible thresholds for y. We also temper this with how much change there is, to determine whether the differences are something we care about.
Obviously, should the system notice that if its change timings are particularly outside where it'd expect given the cohort assigned, it's then able to start moving around its comparison. An example would be a blog category page which updates so infrequently that it's particularly unusual, or a page with a lot of social feeds on it where there's a lot of flux constantly.
Works pretty well, but if anyone's got a better solution I'd love to hear of it.