Cuil Crawl Data: 310 terabytes of compressed data, snapshot from 2007-8
archive.org
archive.org
http://www.reddit.com/r/worldnews/comments/7da5i/police_raid...
Then, the user that was linked to, goes on to describe a technique to use the term "Cuil" as a unit of measurement, for how disjointed something is from something else.
http://static.businessinsider.com/image/4af8564f0000000000e6...
http://static.businessinsider.com/image/4af84e320000000000b7...
http://static4.businessinsider.com/image/4af880fe00000000009...
http://www.businessinsider.com/cuil-office-tour-2009-11#cuil...
That's freaking awesome.
YouTube alone will guarantee you could never store the Internet on a small hand held disk. They're adding 72+ hours of video to the service every minute. And ten years from now, it'll all be HD+ content, and they'll probably be adding 500 hours per minute or something similarly crazy.
I think Wikipedia is around 42gb right now, uncompressed (just the content pages). I don't think that includes the images. So right now we're just to the point where you can store a text Wikipedia on your smart phone with a $30 or $40 sd card. I'd guess in eight to ten years we'll have 500gb to 1tb smart phone equivalents depending on how that all evolves. You might be able to store a dozen plus dual layer full blu ray discs on your smart phone in a decade.
In other words, it's easy to say that we're not going to get anywhere near storing even a sliver of the Web / Internet on a small device in our lifetimes.
If they had rolled their technology with less fanfare they may have made a minor dent in the market, but even then its unlikely. But at least then they could have taken their IP and gone into the Enterprise Search space and perhaps have gotten bought out by someone evil.
At cuil we were told by mgmt of the launch date and our very talented PR team kicked into gear and generated a ton of buzz via a nationwide press tour.
Despite popular opinion, the ops team kept the site up during the massive traffic surge. A combination of poor mgmt, initial deference to user testing, last minute commits and a nasty indexing bug were the reasons the relevance sucked on day 1.
It could be a copyright issue or to keep bandwidth low, but I don't know for sure.