I call bull on this. I seriously doubt Facebook is collecting this browsing history and not associating it with users.
I call bull on this. I seriously doubt Facebook is collecting this browsing history and not associating it with users.
It seems more likely to me that yes, they could do a DELETE FROM BROWSING_HISTORY WHERE USER = "%S", but that's not their business model. Instead they have to add a column to the database that flags data that will be shown to the user or not, and then go through every part of the site where the data might be visible and add a check for that flag. It's a lot more work, and worse they have to get it right the first time because people will be angry if there is a leak somewhere and their supposedly deleted data is visible.
Like, we know Facebook is evil, but they can’t be that evil, right?
People keep saying this and they keep being surprised.
Blowing vile, antisemitic dog whistles from the very top level also doesn't do much in order to improve their image
At this rate, I predict a John Oliver segment taking apart their abuse of the English language in about a month.
ALICE: Column A is the URL. Column B is just a token we use to ensure it's a valid entry, and it's the default sort. As you can see, it's gibberish. No personally identifiable information.
JANE: We deriv...
[JAKE hits JANE with his elbow, discreetly]
JAKE: We derive great pleasure from helping people connect with each other, is what Jane's saying. And that's something we truly believe in.
So how you store the data is important in a post-sql world.
The overall point is that of course they can access the data by individual and not just by time of access.
> For the purposes of this discussion, a flat file is essentially a database. A CSV has fields and is easily made searchable
absolutely doesn't apply to the type of data I'm thinking about. (Data in the petabytes+ is not easily made searchable, indexing isn't common, joins are basically impossible at scale. Instead the data is partitioned on a specific field which makes it fast to access by that field but slow to access on others. Usually for this type of data that field would be a timestamp, but it could also technically be partitioned by a userid which makes it easier to split on user at the cost of splitting by time, since you now need to search all partitions in order to find a time range. )
So either I'm mistaken about the nature of facebook's data (I assumed it was clickstream data, like "User X clicked Box Y at Time Z with Metadata A"), or I'm misunderstanding what you mean by that statement.
Also
> The overall point is that of course they can access the data by individual and not just by time of access.
I can access clickstream data by a specific individual user with a 10 minute Mapreduce job, but re-partitioning the data is going to be a production.
Apologize if you already know all of this, but I saw a lot of people making references to SQL in this extended thread which strikes me as strange because even sql-frontends to analytics backends like HIVE get compiled down to 10 minute mapreduce jobs, it's an entirely different world.
The implementation details are fairly unimportant. It’s what they do with the data, and the satement that it is “stored by” time of access is intended to be misleading to non-experts, in my opinion. Apparently they are even managing to distract experts and technical people.
Usually when you train a model you extract features for each event and then you apply your model at prediction time. There is no step in which you have the user’s entire history set aside in an easy to query fashion. Instead the model learns aggregate user behavior.
Out of curiosity do you work with analytics scale data? Nothing in their statements seemed odd to me, it would easily take a few months at my current company for those reasons.
For analogy it would be like if fb said “oh we support IE 7 to 11 but it’s taking us longer to support IE6”, and people said “it’s just one lower than 7 how hard can it be? Facebook is clearly hiding something, I write for IE7 all the time”. And I’m not quite sure how to convey the fact that IE6 is a different beast altogether.
What do I mean by ‘Facebook tracking individual user behavior’? Um... I don’t know what to say to that. What a strange question. It’s pretty clear what is meant by that. It’s the primary focus of their business.
Next, “usually when you train a model you extract features for each event and then you apply your model at prediction time”. That’s nice, but when did we start discussing ‘training models’?
Facebook has a list of internet activity. It’s stored by user. They’re in the business of selling personal data. Nothing you said has anything to do with that. I have no idea what we are even discussing.
Whether someone can make an ad hoc a query by user is completely not the point and has nothing to do with anything. I would assume that they have extracted the information into other reports or data stores, obviously, not process trillions of lines of entries whenever they want to access it by individual.
Yea I figured. I'm not talking about Facebook at this point, this is a general engineering comment.
My point with the modeling comment is that the might not access it by individual, because models can be trained on statistical averages and time-ordered instead of by individual. Also the models are trained offline (meaning not in real-time).
If you have an analytics table sitting in HDFS ordered by timestamp and you want to produce a realtime UI for users, that'll take some engineering work, right?