Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?
Does any one know how does this compare to let say Yahoo BOSS? Is it even comparable?
As far as comparisons to Yahoo BOSS are concerned, no, we are definitely not comparable to Yahoo BOSS or other such APIs that run on top of an already built (and properly ranked) inverted index of the web. At this stage we only produce bulk snapshots of what we crawl, and we are focusing our engineering resources on improving the frequency and coverage of crawl (the results of which will hopefully start to bear fruit in early 2012). Perhaps at some point in the near future, we can partner with the community to build a rudimentary full-text inverted index of the Web that we can make available in bulk via S3 as well.
From what I have heard BOSS continues to do very well and is pointed at internally as how to turn an API into a real business and product.
One more note, I am now at Factual where we are very happy consumers of the CommonCrawl service.