How Cambridge Analytica’s Facebook targeting model really worked
niemanlab.org
niemanlab.org
"The whole point of a dimension reduction model is to mathematically represent the data in simpler form. It’s as if Cambridge Analytica took a very high-resolution photograph, resized it to be smaller, and then deleted the original. The photo still exists — and as long as Cambridge Analytica’s models exist, the data effectively does too."
That's an eloquent piece of explanation of a very important point. And apropos the discussion about privacy legislation, it's also going to be a very interesting point. Will the Cambridge Analyticas of the world be able to claim they have held on to no personal data, when strictly speaking the raw data has indeed been deleted after being used to create a derivative work that can for all important purposes be used to recreate the original? Assuming I find out I'm being profiled and demand to have my data removed, will society grant me rights to have derivative forms removed or adjusted too? I'm somewhat pessimistic that legal hairsplitting about matters like these will make enforcement very difficult.
I am in favor of no. Imagine I build a gender classification model off public tweets, and then you later delete your twitter account and demand my model not be used because it was trained off 'your data'.
I am in the camp that so long as the data isn't traceable back to you specifically, then don't put any information out there you are not OK with sticking around.
To be precise, you almost certainly cannot use this data to recreate anything remotely resembling the original dataset. This type of dimensionality reduction would throw away enormous volumes of data. There is no meaningful sense in which you can reconstruct the data from it.
What they have done is distill some insights about people from this data. It's arguable whether they should be allowed to keep those insights, but there's no privacy risk there really.
It's honestly kind of disingenuous to describe dimensionality reduction in the way that they do here. It is like reducing the resolution of a photo, but it'd best be described as reducing that resolution to say, the 20 most representative pixels. There's no real sense in which the photo still exists.
So if Google has distilled someone's emails over the years into "closeted homosexual with a deeply repressed leather fetish", that's not an invasion of their privacy as long as they throw away the source materials?
I use ML models every day in my work, and understand how they function. It is true that individuals information is probabilistically encoded into the parameters of the model. However, if the model is any good, the people they trained on's information is encoded only a bit more than that of the entire population.
There is sort of a privacy issue in the following sense: The models they've built have learned relationships between preferences and personalities that they wouldn't otherwise have been able to learn. But these relationships are abstract. They are not tethered to any particular, identifiable individual.
A reasonable argument can be made that those learned relationships are, in a sense, stolen property. And I think arguments along those lines are interesting things that we'll have to explore as this sort of thing becomes more common. But the idea that this model invades individuals privacy just isn't really true.
Edit: is it that the model is then applied to only strictly public data about the person? If so I guess the interesting question then becomes whether the model is definitely not anything near overfitting (i.e. containing enough information to match a person's public data directly since it was trained on it (amongst other data))? (I'm not an ML developer.)
Edit 2: also, going with your comparison with the "20 most representative pixels", it seems interesting then that 'this much' (although not exactly sure how much) information can be inferred from a public profile when just also knowing enough about the whole Facebook population. OK, so perhaps a human would be able to infer about as much, but doesn't scale, and that's why the model becomes valuable?
I don't know exactly what they were modeling, but from the published reports, it sounds like they were trying to predict big 5 personality characteristics (conscientousness, neuroticism, openness, extraversion, agreeableness) from FB profile data (e.g. likes, dislikes, bio, post content, etc.). So in that case, the model would contain weights that measure the strength of relationship between characteristics like "likes punk rock music" and "openness". That description really only literally applies to a linear model - but nonlinear models are, for these purposes, the same.
People very much don't want these models to exist. They don't want a predictive model which will guess their affiliation just by providing unrelated Activity bread crumbs.
That's why I assumed this whole issue has exploded recently.
Not the privacy, but the implications.
What reason do you have to think their data set consisted of only what has been reported?
How do you know anything about the models they used?
It would not be possible to make inferences about the income of any particular person from the slope and intercept, so it would be ok to share those values in, say, a journal article, even though disclosing income of a particular person would not be ok.
What is more relevant is a model which, given characteristics such as "closeted homosexual with a deeply repressed leather fetish", they would be able to infer other characteristics, such as support of particular political candidates, responsiveness towards targeted political or commercial ad campaigns, etc. That's what's relevant here.
My concern would be, how granular is too granular? What if we added "and live in zip code 12355 and is registered Green Party"? This now gets eerily specific, and might be sufficient to identify an individual.
In the US, you're not allowed to benefit directly from a crime you committed. For example, if you rob a bank, you can't buy your mother a car with the money and say "sorry, it's gone!" when the police come knocking.
With that line of reasoning and if there was a legal, privacy, or at least a TOS breach in collecting the data, the derivative machine learning models may be tainted also. Then again, it's likely impossible to prove exactly what data went into the model, so hard to establish which models might be tainted.
If you have a million points that largely fall on a 3-dimensional line and you project that into 2 dimensions, you can easily recover that lost dimension with losses relative to the deviation. And that loss may not even matter depending on the kinds of data and margins of error you're working in.
However, it concisely represents a manifold in a much larger dimensional space and effectively captures most of the information in it.
It may be (and is) lossy, but don't underestimate the expressive power of a deep neural network.
And, for example, where someone's proclivity on the exploration/exploitation spectrum, if you will, (IE, how strongly do they respond to fear-based messaging) falls is probably quite predictable from a spectrum of likes.
Cat pictures may be less informative, but not all of these people clicked exclusively on feline fuzzy photos.
Is this falsifiable? It reads like a tautology to me.
It's dimensionality reduction. You cannot recover the original object. It's like using a shadow to reconstruct the face of the person casting the shadow.
Note this has nothing to do with the expressive power of a deep neural network. You are by definition trying to throw away noisy aspects of the data and generalize a lower dimensional manifold from a high dimensional space. If it's not lossy, it won't generalize.
[Edit: and that the salient characteristics are likely contained in the model.]
There is a real issue here of whether or not they should be allowed to keep a model trained from ill-gotten data. But the way I would think about it is: If you steal a million dollars and invest it in the stock market, and make a 10% return, what happens to that 10% return if you then return the original million? That's a much better analogy for what's going on here. They stole an asset, and made something from it, and it's unclear who owns that thing or what to do with it.
Regardless, I still think having the most relevant features already extracted is all they need to ask many of the questions they might want to. The point is that that’s still quite bad.
Makes me think of the Simulacrum[1]. "The map is not the territory."[2]
1. https://en.wikipedia.org/wiki/Simulacra_and_Simulation
2. https://en.wikipedia.org/wiki/Map%E2%80%93territory_relation
SSN is a lookup key into the raw data. Dimensionality reduction is by definition lossy since it's used in scenarios where: rows of data = n <<< m = number of features
On this tangent, IP ownership for deep learning models is interesting - how to you prove (in court) someone has/hasn't copied model/stolen a training set? If you fed someone else's training/model into your system, how easy is it to prove? Will we see the equivalent of map 'trap streets' in trained CNN models?
Which led me to: https://medium.com/@dtunkelang/the-end-of-intellectual-prope...
i'd argue that insight is the bit that's important, and the bit that's the privacy risk.
This by itself may be mostly true perhaps - and many of the comments get into ways of playing with this dataset to make it better, I don't have experience with those methods, but,
what I have not seen anyone mention, if you have this dumbed down dataset, the original is gone.. you can still combine with other data sets that are either public or previously created and likely fine tune;
dumbed down set + public voter records + public arrest records + previous whatever records - sort, match, what's left over.
and pretty much recreate what you needed from the original, maybe not 100%, but I would guess you could get really close.
First off, I think that's wrong. The idea is after all to keep the information that will result in the smallest error compared to the original on the dimensions one cares about. Within what the model emphasizes a reconstruction can be not only "remotely resembling the original dataset" but as closely resembling the original dataset as is possible with the capacity of the representation.
Next, I'm really not talking only about the particular method described in the post. It's definitely possible to choose to make a light enough reduction to preserve the aspects of the information one is interested in, and to optimize for recall rather than generalization. A more realistic context is going to be that some information about the affected individuals is still exposed or kept (maybe in a compact derived form), which would in many cases give excellent possibilities to restore information accurately enough that claims to have the removed the data are effectively deceptive.
Even for cases where the models are in good faith created only to "distill some insights" I'm skeptical that they really are useless for recovering individual information. I'm by no means an expert in differential privacy but I do listen when it comes up, and a lot of what we see from that field seems to come down to being able to trade off the relation between keeping the data useful and how many pieces of additional information (or assumptions and brute force) are needed to break the integrity protections. With surprises that tend to be on the side of 'Oops. Turns out this clever trick can recover the originals easier than we thought.'
> It's honestly kind of disingenuous to describe dimensionality reduction in the way that they do here. It is like reducing the resolution of a photo, but it'd best be described as reducing that resolution to say, the 20 most representative pixels. There's no real sense in which the photo still exists.
In my honest opinion the original analogy does an excellent job of intuitively explaining that most of the informative aspects of the data are kept (we can still see just fine what's in the image) while irrelevant details are discarded, and that is probably what was intended.
If anything comes off as disingenuous in that context it's your representation that it's like a strong reduction in the pixel domain (where it does indeed destroy a lot of the information). What can be done is much more like running the picture through a high-performance Imagenet classifier and keeping the 20 (or 2048, or whatever's needed) most informative values at a level that corresponds strongly to semantic content of the picture, and holding on the model. We could probably generate images that people would have a hard time distinguishing from the original with that.
The pixel analogy is bad, but to use it anyway -- you get to choose how many pixels you keep. You could keep literally all of them.
To be more extreme there are many compression/extraction methods that can perfectly reconstruct the original data with very high compression ratios. GIF/PNG can reproduce many images exactly. Certainly, they are derivative works?
Hmmm
Let me make that a bit more convoluted for you:
Let's say, very hypothetically, that you, like most of the general population suck at things technological and just switch off mentally when someone mentions phrases like "social graph" or "Javascript", and I'm a semi-psychopath who plans to make a fortune by scamming dumb fucks like you. I take on my mask of sanity and most endearing nerd T-shirt and appear at your door to give you a FREE robot servant. Except to keep it FREE the robot servant is going to pause what it's doing sometimes and whisper subliminal messages to you from my sponsors. But you're not afraid of stuff like that are you? And all you have to do is sign my brick of legal documents in complexified Legalese, which of course you don't have mental stamina to read through. "But, hey", you think "people are nice and trustworthy and if there was something really bad going on here that would be illegal and punished, and besides what's the worst thing that could happen", and I get your signature and you got your FREE robot.
If you had bothered to learn complexified Legalese and do your reading you'd have noticed you also just approved that the robot spends it's spare computation cycles surreptitiously watching you and getting to know you, and one of the many things it does is it glances over your shoulder when you get your porn fix, and in collaboration with my team of highly trained robot masters it concludes not only that you're into furry porn but also which particular furries really push your buttons. We catalogue this away for future use. Years later, in spite of your technical ineptitude, but maybe let's say because you're really a good people person, you've risen to become the highly respected mayor of your town. Then a business partner of mine who used to work covert operations over at CI6 but now has switched to lucrative private contracting, comes to me asking for the files collected from your robot for a influence gig he's taken on from Toxico, and I sell them to him the data for a suitably juicy sum for a man of you stature. Ex-CI6-guy studies the file with interest and uses it to select two skilled furries he can tell you'll be incapable of resisting and sends them to cross your path at a representation dinner. They very convincingly persuade you to have them over the next weekend your wife's away, and it's WILD(!).
And of course "your" robot is carefully documenting the whole thing, which I also sell to my partner, for an additional cut of the profits from the Toxico job. While you still wrestle with your conscience about whether last weekend was really a Good Thing, someone appears at your office door to propose you use your influence to switch the city energy supply over to a 90-year contract on Toxico's patented owl-burning power plant (with levels of carcinogen emissions they'll never be able to get us for!). Incidentally that someone at your office door has probably also been chosen according to your robot file to be a kind of person you'd have a harder-than-usual time to say no to for one reason or the other, but in the end you still refuse because despite some personal weaknesses you're a decent man, who values and protects your town, it's people and it's environmental surroundings. And then you're informed that someone may have videos you'd really rather not become public. Unless you agree to the Toxico proposal and do so generously, your comfy little life may meet with a sudden and radical change of fortune.
So who's to "blame" in this scenario? I'm sure there's a point where ignorance should be illegal, but I think generally society looks with some leniency on getting deceived. And winding back to the early parts, what does this do for your moral rights to have me remove your data? It was right there on page 200 in crisp, clear letters that you allowed me to watch your porn surfing habits and share that data and derived works with selected partners. You agreed to this! We haven't done anything that we're not allowed to do according to the contract.
...
Say we have a bunch of profile images, and then describe them in text. "Blonde, caucasian, large nose, curls, receding hairline, strong jaw, big ears", or perhaps even more specific stuff like "has a mole on the left cheek at the same height as the right earlobe" and "right nostril is larger than the left" and "dimple in chink".
Based on a description like this, we could identify an individual in probably a short paragraph. Nonetheless, on the data side, this is a lot less information than is represented by the raw pixels.
When it comes to the topic of CA's tools, and 'psychosocial' targeting, we can't separate the broader context and the way in which one single term can encode tons of data ("looks like George Clooney with a bigger forehead"). I'd argue the same princple applies to political views, and personality.
To be fair, it does depend on the amount of compression before it is not recognizable, but if you can still squint and see the Mona Lisa (when you also have her phone #)... have you not violated her privacy?
But I also wonder what exactly you are going to do with the predictions. What exactly do you show to someone to make them more likely to go and vote if they are inclined to vote your way, or make them stay at home otherwise? Is there evidence that whatever you're showing actually works? Or do you try to change people's minds? What do you do?
Knowing how the state of things -in this case, people's voting inclinations- is not the same as knowing what to do, ie a strategy.
I don't know how effective it is, I'd like to learn more. But I smell the possibility that these CA type firms are simply selling snakeoil to desperate political activists.
Qualitatively: show things that get them angry.
Quantitatively: test and control pop splits.
Or don't look at votes, look at candidate likes and shares over time, especially as they shift.
The defined metric doesn't have to be "propensity for this individual to vote for a candidate." It can be "percentage delta over untreated markets compared to prior campaigns."
How do you actually do this? Presidential elections come once every 4 years.
What if all the sensitivities are dependent on the length of the candidates' hair? It seems the total hair length of the two candidates was a maximum at the last election. Another time you might be sampling more towards the middle.
I could imagine this working on Dem voters who are wavering on Hillary with leads like "she thinks the TPP is the gold standard" etc.
If you understand someone's mentality on the subject you can decide if they see:
1) An ad with someone breaking into a home and the homeowner defending themselves with a firearm (sell insurance?)
2) A grandfather and grandson on a hunting trip (hunting supplies?)
3) Or maybe gun violence hotline with powerful images.
The people seeing these ads are under the assumption that everyone else sees them, not that it's specifically targeted at their personality type. These affect if you think other people understand your issue or not. Thus affecting your motivation and attitude.
If you see an ad that fits your mindset, you think you're on the majority side. This was powerful in classic media, it's just as powerful now.
How long will that be true? Do people make that assumption about search results?
I agree that it's pretty obvious you're being retargeted when ads for camping supplies start showing up three days after you search for them on Amazon. But the practice of "personalization" of results and ads is far larger and deeper, to a degree that most people never seem to think about.
The truth is WAY worse of course, but she immediately knows the ads she saw won’t show for me as well.
Even when I explain how ads can be different, I don't think people really want to believe it, or understand it, and they certainly do not realize the power of these targeting abilities..
According to the article:
"The accuracy he claims suggests it works about as well as established voter-targeting methods based on demographics like race, age, and gender....the digital modeling Cambridge Analytica used was hardly the virtual crystal ball a few have claimed."
It's pretty clear that they were selling snakeoil. In fact, the use of CA wasn't particularly helpful to anyone [1]...hiring them was just a prerequisite for obtaining campaign contributions from the Mercer family, who had put up the money behind CA [2].
[1] http://www.businessinsider.com/cambridge-analytica-facebook-...
Here is an interesting Ted Talk which discusses an FB experiment that details how effective minor UI changes can be on voter turnout (13:40)
https://www.ted.com/talks/zeynep_tufekci_we_re_building_a_dy...
One year they will have been round and had a lengthy discussion with Mrs X, but Mr X slammed the door in their face another time. This was somewhat lower tech: the information was printed out and attached to a clipboard.
Most of the time this information is correct. It's more interesting when it's really incorrect. That said, some of the best sessions I've been involved in were where there was no information.
Demand? That seems like a great way to get arrested or shot for trespassing.
In Facebook campaigns you can use certain things, such as user's interest, to select who sees your message.
I'm not an expert on Facebook analytics, but I believe you can get pretty good stats on how your campaigns are working, how much promoted posts get shared etc.
This sounds like the holy grail for advertising. You get to write your message for certain profile and get quick feedback how it worked. Even if the system is not perfect, you would have an advantage compared to somebody else who is spending the same amount of money and not using similar targeting.
Maybe their model also allowed them to find social influencers with many followers. Being able to targer these people and get them to share your message would be really good.
The article compares this to the effectiveness of traditional voter targeting methods. I'm not sure what the parameters used on those are, but maybe all of them are not available on FB, justifying the need for something else.
Besides that they apparently also used similar trickery in their consultancy for the Brexit side.
https://www.theguardian.com/politics/2018/mar/26/pressure-gr...
This is far from over.
https://www.theguardian.com/commentisfree/2018/mar/23/plenty...
It's often not hard to convince people of something they want to believe.
The fact that people were willing to spend an amount of money that breached electoral law in the UK, and presumably even more in the US, suggests that there was some reason for them to do so. This happened only because experts in this field believed it would influence the outcome of the election.
That's your evidence.
I thought Clinton spent large amounts of money on data and the Democrats admitted the data was bad or at least that was their excuse. How much did CA pay for this data? I still find it crazy that Trump campaign spent 30% of what Hillary did and still won. The Russians used 100k$ worth of ads to sway the election. This stuff doesn't t add up.
https://www.lrb.co.uk/v40/n01/jackson-lears/what-we-dont-tal...
> Trump and Western Allies Expel Scores of Russians in Sweeping Rebuke Over U.K. Poisoning
> WASHINGTON — President Trump ordered the expulsion of 60 Russians from the United States on Monday, adding to a growing cascade of similar actions taken by western allies in response to Russia’s alleged poisoning of a former Russian spy in Britain.
> [...]
> On March 15, the Trump administration imposed sanctions on a series of Russian organizations and individuals for interference in the 2016 presidential election and other “malicious cyberattacks,” its most significant action against Moscow until Monday.
> [...]
> Mr. Trump has said that, despite its denials, Russia was likely behind it. “It looks like it,” he told reporters in the Oval Office on March 15, adding that he had spoken with Prime Minister Theresa May of Britain.
You have to wonder how far Trump has to go before something he does is considered hostile to Russia. Does he have to nuke Saint Petersburg?
Russian news confirmed it and later us state department.
Russian news https://amp.vesti.ru/doc.html?id=3001089&3001089&3001089=&id...
Business insider https://www.google.se/amp/s/amp.businessinsider.com/theres-a...
The CEO of Cambridge Analytica has also been recorded telling a (fake) potential client that they routinely blackmail people using prostitutes and who knows what else.
So unless you can show that Clinton's campaign was doing the same things, your claim is a false equivalence.
Using lies to convince someone to do something is going to be more effective than using truth, if that "something" is not in accordance with the truth.
The comparison with netflix really breaks down there. You're not going to be able to convince me that I liked Crash, so recommendations based off of that aren't going to be very useful to me.
But if you reinforce my false belief that Obama and Soros are gonna use the deep state to invoke Sharia Law on the 2nd amendment, then that might better convince me to vote for so and so.
I cannot find those posts for the life of me again. Not suggesting anything nefarious here, I just can't find them. Does anyone have a link to those early conversations or make copies of the papers?
I made copies earlier but deleted them before I put them into my papers archive.
https://news.ycombinator.com/item?id=14486365
https://news.ycombinator.com/item?id=14393991
https://news.ycombinator.com/item?id=14330547
https://news.ycombinator.com/item?id=14284502
https://news.ycombinator.com/item?id=13939814
And the query: https://hn.algolia.com/?query=mercer&sort=byPopularity&prefi...
I'll keep looking.
I'm not sure this is the model you want to emulate. The suggestions are terrible and continually getting worse.
But that's because they dive into other user playlists that contain the same music you play. Pretty simple. I assume Youtube do somthing similar.
If Netflix had playlists or a 'want to watch' feature I bet their recommendations would improve.
Or, just use a separate browser. I typically only use opera for YouTube & anything where I don't mind Google tracking me, but if I open a YouTube link in Firefox I'm not logged in (and w/ opera's VPN enabled it doesn't appear to affect YouTube recommendations). I'm sure Google does correlate traffic between the two to some extent, but this seems like the only useful way to use operas integrated VPN.
So, he likes horror movies and will watch any horror regardless of any signal that it's going to be poor. You can look at the viewing history and see he rarely goes beyond 5 minutes of watching any of them.
I watch Netflix regularly throughout the year, he watches in phases that last a week or two and then nothing at all for months at a time (for reasons that should seem obvious by now).
As a non-horror movie aficionado, I can say with certainty that only 2% of horror movies are ever worth watching and only 50% of these are any good. As a consequence, my personal viewing history includes almost no movies in this genre.
My favoured genre is drama and I normally watch all the way through.
You should be able to guess by now that I should rarely be recommended horror movies but, alas, Netflix thinks otherwise.
Btw. I also rate movies I watch - my friend doesn't.
I assume that Netflix’s model has the premise that profiles aren’t in fact, very easy to create. I have separate profiles for my parents, as well as a test profile to see what happens when a user only seems to like the “Human Centipede” trilogy.
And? You can't just leave us hanging on that.
It doesn't bother me since I know what I like. There are not that many good films that I'm not going to find them anyway.
Someone, or a group of people, are being paid for nothing though. I don't know anyone who subscribes to Netflix because of their recommendation algorithm.
This is far from my previous stance of "Huh, movie I've never seen before and in a different language? It's 95% so yeah, I'll give it a go."
And before someone says anything, I do vote up and down quite frequently. But can't help but notice that my suggestions were better when I had the more nuanced star system.
[1] https://i.imgur.com/MOt3XlL.png
[1.1] How I'd rate these. Cars 3: not really interested. Liked the first though. Stitches: idk, doesn't look appealing. Teeth: Classic cult film but yeah... Big Mouth: I have ZERO interest in watching this, please stop suggesting. I have downvoted this! Waterboy: I like it, but far from 95%. I'll give it like an 80.
I honestly thought the 1-5 ratings were not useful since I either like or don't like movies, I am not interested in nuances. But, it's not working out as I expected.
I do research in this area and it's fairly well established that when you go from something like five points to two points with ratings, you throw away tons of information. There's diminishing returns with numbers of points, but as you go lower you lose information.
The "ratings don't matter because what you want is implicit signals from peoples' actual behavior" is also disingenous because the rating behavior is a behavior that's directly tied to the stimulus in question. Not saying that indirect behavioral correlates aren't useful, only that the rating is a very powerful, direct correlate that tends to be very specific. Going back to the topic of the thread, sure, all those Facebook likes are going to be useful in predicting how much you like a candidate, but you're sure as hell going to get a lot of information by just asking them "on a scale of 1 to 5, how much do you approve of X?"
Or are you saying your netflix suggestions are getting worse?
Mine are pretty good and have held steady for a while, at least in terms of my own preferences, though I also am probably using it a little less as I've got Prime and Hulu now as well, so there's probably less times I'm randomly searching through Netflix and finding nothing.
I'll also add that I am more frequently searching for 15 minutes then switching to another service. I used to find a movie to watch in 5 minutes. I am a big movie person too, and will watch most things. But I am also more aware that if I watch one show that is just "meh", then I am going to be bombarded with shows of similar quality for the next few weeks.
Also, there is a fairly obvious pattern that shows up from the movies in "my list". Those do not seem to be weighted more heavily.
First, they got rid of those wonderful ranked lists that made us love Netflix in the first place, replacing them with the much more opaque cover art carousel view. Then they started mixing in lower-ranked items into the carousel. Finally, they switched from the five-star rating to the thumbs up and down buttons, which can't possibly give them as much information about your opinion.
Note, the comedy special was similarly panned across the press and social media as being repetitive of her previous work, extremely predictable punchlines, and when seemingly good later exposed to be highly derivative of other comedians work.
But regardless of the reality/honesty of her rating it was apparently not good for Netflix's business when they let users destroy content they produced destroy it on their own web site via user generated content.
So the suits (heavily swayes by their production studio and Hollywood) were able to convince the product team to hurt the UX for 90% in order to protect the popularity of the 10% of content they own.
The truth may be good for consumers in almost all situations. But sadly the interests of executives dealing closely with high value B2B partners and investors tend to have a way to out-valuing the interests of the average user (not to mention far out valuing power users).
So I guess we have to rely on 3rd party IMDb web extensions which inject into Netflix in order to get honest ratings.
http://www.breitbart.com/big-hollywood/2017/03/18/netflix-sc...
Notice that they're careful to say that they made the switch "amid" the special, not because of it. Also as far as I can tell, they have no actual data on the fact, and they're the only "newspaper" reporting it.
I'm curious why you brought up Breitbart? I just googled it and found a ton of other (non-political/right-wing) sites which drew the same exact conclusion between Schumer and the ending of 5 star reviews. Are you trying to say her thousands of awful reviews were somehow political? Or that it's all just some "alt-right" right-wing conspiracy?
- https://movieweb.com/netflix-cancels-5-star-rating-system/
- http://ew.com/tv/2017/03/16/netflix-star-ratings/
- http://collider.com/netflix-rating-system-thumbs-up/
- http://screenertv.com/television/goodbye-stars-hello-thumbs-...
- https://www.washingtontimes.com/news/2017/mar/17/netflix-cha...
I saw Amy's show live, which she later filmed, and it was just awful. Those reviews were highly justified. And I used to be a big fan of hers before anyone knew who she was.
If Amy Schumer wasn't the original reason why they started to ditch 5-stars then it was most certainly a motivating factor to get it pushed out (conveniently right after her special was famously destroyed - which got widespread press before the switch over to thumbs). As it most certainly also faced internal resistance as they were basically abandoning a decade of 5-star reviews via user-generated-content, in favour of a vague Thumbs Up/Down.
Amy provided a perfect example of the inconvenient and conflicting goals which Netflix has as a producer and content platform. The timing of the release would have been HIGHLY coincidental if it wasn't related in some fashion.
Does facebook provide an option to show a particular given ad to a particular given user? Or is it possible to select a group of people with a given set of likes? How fine-grained is facebook's audience selection mechanism for ads?
Or was the targeting performed by creating fake groups, befriending people?
http://www.stat.columbia.edu/~gelman/research/unpublished/sw...
https://www.politico.com/magazine/story/2014/01/independent-...
https://www.thenation.com/article/what-everyone-gets-wrong-a...
The last piece has a short summary of the salient point:
> In fact, according to an analysis of voting patterns conducted by Michigan State University political scientist Corwin Smidt, those who identify as independents today are more stable in their support for one or the other party than were “strong partisans” back in the 1970s. According to Dan Hopkins, a professor of government at the University of Pennsylvania, “independents who lean toward the Democrats are less likely to back GOP candidates than are weak Democrats.”
> While most independents vote like partisans, on average they’re slightly more likely to just stay home in November. “Typically independents are less active and less engaged in politics than are strong partisans,” says Smidt.
> [...]
> The conventional wisdom holds that the parties need independents to win general elections, but the reality is that they’re increasingly devoting their resources to getting their own voters—including their “closet partisans”—out to the polls rather than trying to sway the dwindling number of genuine swing voters. “We’ve seen a huge increase in technology and the ability to turn out the vote,” says Smidt. “So in terms of a cost-benefit analysis, the parties and candidates see that it’s much easier to turn out people who agree with them than it is to change someone’s mind. And then there’s also the question of how many of us are even open to changing our minds.”
It's a pretty big gap between using and abusing social media and as far as I know Obama's campaign did not 'narrowly tailor messages'. They did target broad groups using generic messages and they did quite effectively use social media presence to build support.
But they did not - as far as I know, so please correct me if I'm wrong - go so far as to single out individuals or really small groups with the express intent of flipping their votes or targeting them with disinformation in order to try to stop them from voting.
And Cambridge Analytica seems to have been doing just that if the currently available information is to be believed.
> The former president also hired Facebook co-founder Chris Hughes to help in developing his social media strategy. Obama furthered the use of Facebook for his 2012 re-election bid, utilizing it to encourage young people to cast their votes. His team developed a Facebook app that looked into supporters’ friends list to find younger voters. The team then asked supporters to share online content with these voters. More than 600,000 supporters responded to the call, sending content to over 5 million contacts.
> During his presidency, Obama continued to use Facebook to reach out to the public. In 2016, he became the first president to go live on the site, just before his final State of the Union Address.
Please read the article and compare what we know about Cambridge Analytica vs what the Obama campaign did, it is comparing snipers with someone setting off fireworks.
ie, moving beyond "soccer moms" or "defense dads" but "soccer moms with one kid and expensive tastes".
The two parties already have a list of registered party members (and they can see who on Facebook explicitly states their party preference), for those members the main goal is higher turnout (they are the training data). The other voters they're interested in are unregistered (e.g. independent) voters that are likely to be on their side ideologically.
The core idea is very simple, they believe that if someone says they're independent, but their preferences/features (age, gender, location, likes, posts) are predicting moderate or high likelihood of $PARTY affiliation, then showing this person political ads may move them from 'maybe vote for $PARTY' category and get them in the 'definitely vote for $PARTY' category.
If you have continuous access to new Facebook data as you're serving ads, you can verify your ads are working on an individual basis by checking the predicted 'score' for $PARTY affiliation predicted by your model before and after an ad (I want to stress that this can be done on an _individual basis_). The likely sequence of events is that they did AB testing on different kinds of ads and found that fake inflammatory ads were most effective at achieving this goal in a very measurable way ($PARTY score), the resulting media/political atmosphere is collateral damage (hopefully unintended).
Source: I am a data scientist / machine learning scientist and this is how I would do it and how it seems to me others. I don't work on political data but I have worked on personalized recommendations which are similar.
Plus the approach you outlined would require the user to like/dislike things based on the add they saw, so CA can observe a change in the predicted affiliation (they didn't have access to posts as far as I know). I don't think it would have that effect (even if the add influences you, I doubt that it would make you go unlike Obama's page for example). Not to mention that by any likelihood you shouldn't be able to verify that a particular add was shown to a given individual.
I suspect it was a simpler use case - they would group users into segments, and then craft different add strategies for each one (maybe based on other research or just expert opinion).
It is in this last process that it is individual-based. It is in this last process that AB tests are done individually as a function of the specific strategy applied to him/her
I assume they wrote/created different ads for different sets of users... but how many segments did they have? Did their graphic designer build 500 different ads, or was text/images dynamically inserted based on these variables? How did they figure out which message would resonate with each segment? How did they test something like this, with so many potential variables? Was this knowledge used only on facebook, or across all digital channels? Was it implemented in non-digital channels as well?
I'd kill to have access to their campaign set ups.
Unfortunately all EU institutions are terrible at marketing. If they did an ad campaign, it would probably be a TV commercial showing Jean-Claude Juncker giving a speech with subtitles in 15 languages.