I just finished crawling 5.19B web pages, Ask Me Anything
I WAS JUST RATE LIMITED BY HN, SO IM GOING TO ANSWER YOUR QUESTIONS UNDER A NEW ACCOUNT: dor_jack_2
As for processing the data we crawled, we are using ArchiveSpark (https://github.com/helgeho/ArchiveSpark)
Also, Mixnode defaults on Amazon S3 for storage which was ok with us since we're using EC2 for processing the results.
The data does not follow a DFS or BFS pattern so pages/site varies greatly by a host's server capacity and anti-crawling configs.
There was a minimum of 10 seconds between followup requests to the same website unless robots.txt had a lower delay. Pretty polite...
Will update this in a few days with more data.
As one would expect the vast majority of data recorded is text/* (html,...)