The Implications of Facebook Indexing a Trillion Posts
techcrunch.com
techcrunch.com
Apart from that I have no answer. However, I find it important to realize the question shouldn't be about what Facebook/whateverelse can "do" now, but what will "they" be able to "do" in the future.
Imagine if there was some Superior Entity which knew everything about every single person born after some year X. What could that Superior Entity do? What could it entail if the knowledge of this Superior Entity got leaked to some hands? I mean, everything about politicians, press, authorities, laymen, teachers, children... Everything.
What are the actual practical worst-case implications of total ubiquitous surveillance for an individual? For society? This certainly depends on the society. Implications are certainly different for people living in, say Iran, Sweden, USA, Canada, Brazil, Russia, China, Germany, North Korea ... For example in some countries religious matters, although very personal, can be very serious. Go being an open atheist in say Indonesia.
I really have no clue, but everything seems like a huge mess with surveillance.
Exactly. It's not about Facebook or the authorities mining data for criminal activity. The problem occurs once the authorities have you in their sights, for whatever reason. With reams of private data available and no context, they can (and will) use it to paint any picture of you they need to.
I was pretty sure Facebook was already doing this, even before the public search. No?
> Never talk about (...) babies? Facebook could eventually filter those out of your feed.
That would be a very welcome change.
Maybe they partition the graph into cliques which they index, search and then fix up with a second pass to add and remove noncliqued friends?
Off-topic: Wasn't there a start-up a while back that was allowing people to build their own personal search index for their social media content?
In fact it was probably already developed long before this feature was deployed to the general populace, and this new feature is just a token gesture of "giving back to users."
I doubt it'll see much use from end-users themselves.
Yes the search terms from users will be useful, but I don't see this as being a replacement for Google.
Edit: I don't think Twitter do nearly the same filtering with this and their social graph so maybe not as close as I first thought.
https://www.facebook.com/notes/facebook-engineering/under-th...
https://www.facebook.com/notes/facebook-engineering/under-th...
Indexing is an exercise in denormalization. At the time I post X, the data is also stored in an index keyed by eg every phrase I wrote and maybe even stemmed by synonyms. This makes search fast. It is a memory-time tradeoff which basically precomputes the answer for every query, limited to my feed.
As for searching 500 indexes in parallel - this is hardly unreasonable. The alternative -- updating N indexes for every post where N is the number of friends -- is more unreasonable, since it does the maximum amount of work, when the vast majority of it is wasted. No friend is going to search for every single word. On the contrary, some friends will eventually search for a large number of words in YOUR index.
Facebook can easily do parallel queries and combine them. Even with a simple MySQL partitioned / sharded setup a site can do it. We do it at Qbix. Let alone Facebook which has improved on mapreduce in their architecture to be more dynamic (http://thenextweb.com/facebook/2012/11/08/facebook-engineeri...) and then there's this: http://www.semantikoz.com/blog/hadoop-2-0-beyond-mapreduce-w...
So yeah. That's how they probably do it.
On a side note: accounts are persistent through time, and it's not hard to match multiple old deleted copies (resignups) and create a 'super-account'
PHP sometimes doesn't get the credit it deserves for what can be accomplished with it, particularly modern PHP... but i'm not entire sure it deserves credit for this in any reasonable way.