Facebook’s Clear History privacy feature is still months from launching
recode.net
recode.net
"Why can’t Facebook just stop collecting your browsing history entirely? Well, it could, but a large part of Facebook’s business depends on collecting this kind of browsing data, so it would cripple a big revenue stream."
As long as Facebook's business model involves monetizing your data, they are going to slow walk any feature that might impact this.
It is a great approach for marketing teams. They grow, look good, claim to drive revenue, try to impact consumer behavior, but it all is dependent on Facebook's (insert any ad platform partner) attribution. So yeah, Facebook will not mess with its attribution. It would make a big mess for both its business and it's client's business.
Marketers's should measure the incremental impact of their marketing buys (get their data science team to manage the study). They will be surprised to see how much display banners or video ads are impacting consumer behavior. Stay away from modeled approaches, try to grade the reporting yourself.
> It is difficult to get a man to understand something, when his salary depends on his not understanding it.
"Facebook Clear History doesn't fully erase user data, information later sold to a group of Zombie Nazis bent on world domination."
> There’s a reason that Clear History isn’t called “Delete History”: Using the feature will disassociate browsing data that Facebook collects from your specific account but it won’t be erased from Facebook’s servers completely, Baser said. Instead it’s just “de-identified,” which means it’s stored by Facebook but no longer tied to the user who created it.
I call bull on this. I seriously doubt Facebook is collecting this browsing history and not associating it with users.
ALICE: Column A is the URL. Column B is just a token we use to ensure it's a valid entry, and it's the default sort. As you can see, it's gibberish. No personally identifiable information.
JANE: We deriv...
[JAKE hits JANE with his elbow, discreetly]
JAKE: We derive great pleasure from helping people connect with each other, is what Jane's saying. And that's something we truly believe in.
It seems more likely to me that yes, they could do a DELETE FROM BROWSING_HISTORY WHERE USER = "%S", but that's not their business model. Instead they have to add a column to the database that flags data that will be shown to the user or not, and then go through every part of the site where the data might be visible and add a check for that flag. It's a lot more work, and worse they have to get it right the first time because people will be angry if there is a leak somewhere and their supposedly deleted data is visible.
Like, we know Facebook is evil, but they can’t be that evil, right?
People keep saying this and they keep being surprised.
Blowing vile, antisemitic dog whistles from the very top level also doesn't do much in order to improve their image
At this rate, I predict a John Oliver segment taking apart their abuse of the English language in about a month.
So how you store the data is important in a post-sql world.
The overall point is that of course they can access the data by individual and not just by time of access.
> For the purposes of this discussion, a flat file is essentially a database. A CSV has fields and is easily made searchable
absolutely doesn't apply to the type of data I'm thinking about. (Data in the petabytes+ is not easily made searchable, indexing isn't common, joins are basically impossible at scale. Instead the data is partitioned on a specific field which makes it fast to access by that field but slow to access on others. Usually for this type of data that field would be a timestamp, but it could also technically be partitioned by a userid which makes it easier to split on user at the cost of splitting by time, since you now need to search all partitions in order to find a time range. )
So either I'm mistaken about the nature of facebook's data (I assumed it was clickstream data, like "User X clicked Box Y at Time Z with Metadata A"), or I'm misunderstanding what you mean by that statement.
Also
> The overall point is that of course they can access the data by individual and not just by time of access.
I can access clickstream data by a specific individual user with a 10 minute Mapreduce job, but re-partitioning the data is going to be a production.
Apologize if you already know all of this, but I saw a lot of people making references to SQL in this extended thread which strikes me as strange because even sql-frontends to analytics backends like HIVE get compiled down to 10 minute mapreduce jobs, it's an entirely different world.
The implementation details are fairly unimportant. It’s what they do with the data, and the satement that it is “stored by” time of access is intended to be misleading to non-experts, in my opinion. Apparently they are even managing to distract experts and technical people.
Usually when you train a model you extract features for each event and then you apply your model at prediction time. There is no step in which you have the user’s entire history set aside in an easy to query fashion. Instead the model learns aggregate user behavior.
Out of curiosity do you work with analytics scale data? Nothing in their statements seemed odd to me, it would easily take a few months at my current company for those reasons.
For analogy it would be like if fb said “oh we support IE 7 to 11 but it’s taking us longer to support IE6”, and people said “it’s just one lower than 7 how hard can it be? Facebook is clearly hiding something, I write for IE7 all the time”. And I’m not quite sure how to convey the fact that IE6 is a different beast altogether.
What do I mean by ‘Facebook tracking individual user behavior’? Um... I don’t know what to say to that. What a strange question. It’s pretty clear what is meant by that. It’s the primary focus of their business.
Next, “usually when you train a model you extract features for each event and then you apply your model at prediction time”. That’s nice, but when did we start discussing ‘training models’?
Facebook has a list of internet activity. It’s stored by user. They’re in the business of selling personal data. Nothing you said has anything to do with that. I have no idea what we are even discussing.
Whether someone can make an ad hoc a query by user is completely not the point and has nothing to do with anything. I would assume that they have extracted the information into other reports or data stores, obviously, not process trillions of lines of entries whenever they want to access it by individual.
Yea I figured. I'm not talking about Facebook at this point, this is a general engineering comment.
My point with the modeling comment is that the might not access it by individual, because models can be trained on statistical averages and time-ordered instead of by individual. Also the models are trained offline (meaning not in real-time).
If you have an analytics table sitting in HDFS ordered by timestamp and you want to produce a realtime UI for users, that'll take some engineering work, right?
I think this is a very important part of the article. History will still exist but not tied to a particular user.
I suppose that's the subtext of what you wrote, though.
1. Does your app collect anything on a mobile device's clipboard and transmit this information to your servers?
2. Does your app send thumbnails of every image accessible from a device, even if the user hasn't explicitly selected the image for upload?
I'm sure there are other examples, and perhaps device permissions prevent these in ways I'm not familiar with but it'd amaze me if companies weren't grabbing every bit of data they could, including the above two scenarios.
[0] https://www.washingtonpost.com/technology/2018/12/19/dc-atto...
Assume Facebook brings in $10/user/month in ad revenue.
But the distribution is uneven; $50/user/month for rich westerners, $6/user/month for poor people in poor countries who outnumber westerners 10:1.
If you offer everyone the chance to go ad-free for $10/month, your income from rich westerners drops from $50 to $10 and your total revenue drops to $6.36/user/month.
The users that are most valuable to advertisers are those with a lot of disposable income.
If you offer an ad-free experience for the value of the average user, far more of these high-income, high-value users would subscribe, both because they have the money and because they tend to be more privacy-conscious.
Imagine (simplified) losing ad income from the top 50% of your users for the price of your average user, while being left with only the bottom 50% to show ads to.