Greplin: 1.5 Billion Documents Indexed, Six Engineers
techcrunch.com
techcrunch.com
Is it fair to say that the size of the "private" web (what Greplin aims to index) is, in aggregate, larger than the public web? And are there any amazing things that become possible once you've indexed a large portion of that private web?
Scares me a bit too much to sign up for the convenience.
Therefore, the comparison with Google’s web-wide index in 2001 is a little misleading (in terms of the amount of data), given that the average size of a web-page is greater than all of these.
Of course average size of a file on Dropbox is likely to be larger than a webpage. I wonder what percentage of those 1.5 billion documents are files on Dropbox.
I am building a startup that does that i.e. it indexes your doc/pdf files (more formats coming), and allow you to instantly search through them. It's called grepfiles.com, but is in very early stage (pre-alpha), so go easy on it since I am not sure how well it scales. Mail me at mail@asif.in if you have any feedback. Would really appreciate it.
Think it is a public beta and anyone can signup if not ping me and I'll send you an invite.
My main concerns with the service are: + Centralized risk - keys to very valuable kingdom + No two factor - but they tell me its coming + No word on whether they encrypt in storage - although it should only be an index to the information rather than the actual info + Standard SAAS / Cloud risks - internal abuse, legal turnover etc.
Any others? All of these could be mitigated to a reasonable degree. What do you think? Is there a future for this type of service (or big buyout for Google / Bing) or is it just too scary?
I'd be okay using Greplin if I knew Google was going to acquire them. I trust Google. I figure when Google goes bad, there will be much bigger issues facing humanity and our internet pasts will all be damning anyways.
I think it is a sign of the times that I read this paragraph, though it was a subtle joke, reread it and decided it was serious, and then did some more thinking about whether the author is serious or not.
Just the logistics of _handling_ 1.5B docs would keep six people pretty damn busy.
It is pretty impressive, though saying that it launched in February is misleading. I signed up last year, ran into a bunch of problems with it not indexing anything, and haven't opened it since. Now it looks like everything actually has been indexed, which is cool. I'm deleting my account for now though, as it doesn't yet seem easy enough to be useful for my purposes.
Security aside, one of the fears I have isn't necessarily against hackers, but against legal entities making use of the private information illegally, in addition to Greplin selling "me" in a very compact and precises manner to whoever they want.
So the question even for the seasoned computer security expert that want to use a distributed Greplin variant is: Do you trust your friends and colleagues to have better security on their home or work computers than Greplin can achieve with dedicated work?
With a distributed system it would still be a non-trivial task to protect against a dedicated worm or trojan that infest the network and traces paths to other Greplin users after stealing all the data from each instance.
Since the data is social and each document in many cases concerns more than one person, it might actually be a less complicated task to achieve sufficient security in a central location.
Likewise, someone installing something on their own machines for privacy concerns can be said to have more vested in keeping things secure than the person who's only doing it for their job, maintaining a server with thousands of bits of data on it.
With cloud providers like Amazon providing computing power on the pay as you go basis I am not sure why this is a news now days.
Some ridiculous comparisons are thrown about in the article -
same size as Google’s web-wide index in 2001
60 times the size of Google’s original 1998 index
I am not sure how to process and make sense of these comparisons.
(a) It's a proxy for traction. Greplin indexes data that can't be crawled; users have to authorize it to index their data. So aside from how hard of an engineering feat it is, the fact that they've indexed this much data probably means that they have a sizable number of users.
(b) While you're right that the technical challenge of indexing that many documents is easier now than in 2001 thanks to things like AWS (and numerous open source projects), to do it with a team of six is still impressive.
they have an engineering blog as well http://tech.blog.greplin.com/
Granted, given the 30MM/day number they must be growing that index very quickly and they've likely hit that 1.5 mark pretty darn quickly.
Solve?
Greplin has probably not built their own search technology. I'd guess they're simply running Lucene or Sphinx like everyone else.
Their index is still small by search standards, as you can tell from TechCrunch having to reach 10 years back to make an "impressive" analogy.
Today, 1.5 billion documents translates to a couple terabytes of data (probably high single digit). 30 million indexed/day translates to about ~400/sec. You could store and process all that on a single, beefy box. Or you can spread it out over a couple amazon instances.
But yes, in 2001 this would have been impressive. In 2001 you'd pay $150 for a 40 GB harddrive...
Google's global index: 1 billion documents. Searchable by 1 million users. Need to support 1B x 1M search capacity.
Greplin's individual indices: 1000 documents/user for each individual index. With 1 million users, there are 1B documents total. Each user only searches his 1K index. Only need to support 1K x 1M search capacity.
It's orders of magnitude difference.
I'm sure we'll be hearing much more from these guys.