So they had more than age, race, and gender, but it doesn't really say how things were weighted.
Doesn't have to be that way. It could be age, age+1, age+2 ....
The only relevant 55% that I can find is:
> Allport and Kramer (1946) randomly presented 20 yearbook photographs of Jews and Non-Jews to 223 undergraduate students for 15 s each and asked them to categorize the person in each photograph as Jewish or non-Jewish, or to pass on the trial by indicating a lack of knowledge. The reported median identification for the sample was slightly above chance (55.5%; Allport & Kramer, 1946). Moreover, they found that highly prejudiced people were more accurate at distinguishing Jews from non-Jews
So it's not the same set of photos and not the same question.
1. The data:
"We used a sample of 1,085,795 participants from three countries (the U.S., the UK, and Canada; see Table 1) and their self-reported political orientation, age, and gender. Their facial images (one per person) were obtained from their profiles on Facebook or a popular dating website... Facial images were processed using Face++37 to detect faces. Images were cropped around the face-box provided by Face++ (red frame on Fig. 1) and resized to 224 × 224 pixels."
2. The benchmarks:
"For example, when asked to distinguish between two faces—one conservative and one liberal—people are correct about 55% of the time."
3. The controls:
"What would an algorithm’s accuracy be when distinguishing between faces of people of the same age, gender, and ethnicity? To answer this question, classification accuracies were recomputed using only face pairs of the same age, gender, and ethnicity."
A. A complaint:
Geography and income are two powerful conditioners. These can leak in so many ways: uncropped background (geography), image color and quality (income), eyeglass shape (geography and income). This study really needs more controls. Geography and income would be a nice start.
But then the data wouldn't represent the natural world: nature as it is.
Raw data is the correct thing to use, because it's what a hypothetical other person would also use if you ran the same experiment yourself.
That way we can see how much of the performance is from magic AI pixie dust, and how much is from basic 19th century statistics.
Every time I read a paper like this, I have this Margaret Mitchell talk [1] in the back of my mind.
The problem was, that with the ML, they ended up building a ruler classifier, because most of the pictures with skin cancer happened to also have a ruler in them to measure the size.
> Their facial images (one per person) were obtained from their profiles on Facebook or a popular dating website
so of course the first thing to comes to mind is "how good of a predictor is just knowing which of those two sites the image came from?"
You think teachers are underpaid? Oh obviously you must be a pro-abortion, $15 minimum wage supporting, transgender-rights activist.
What's that you say, Christian bakers should be allowed to refuse to bake a cake with a pro-gay message on it? Oh, you must be a gun-toting, pro-life, anti-immigratnt Trump fanatic.
This kind of sorting people into simple binary categories, and giving them a "shopping bag" full of opinions they're supposed to hold helps nobody.
I'm not sure how this was relevant in anyway to your comment, but I just kinda jumped on my soapbox there.
A face descriptor is obtained from the learned networks as follows: the centre 224 × 224 crop of the face image is used. The shorter side is resized to 256, and the CNNs descriptor is computed for this region by extracting the deep features from the layer adjacent to the classifier layer. This leads to a 2048 dimensional descriptor, which is then L2 normalised.
People under 30: 60-36.
White men: 38-61
Black women: 90-9
So there are definitely some strong predictors there.
Source: https://www.businessinsider.com/2016-2020-electoral-maps-exi...
The abstract states:
>Accuracy remained high (69%) even when controlling for age, gender, and ethnicity.
To give some context, chance is 50%, human guess is 55% and a 100-question questionnaire is 66%.
Personally, I am surprised that the accuracy remained that high when controlling for the three variables I would have considered most telling in the determination (age, gender and race).
I'd be very curious to know what exactly the algorithm is determining from the face photos outside of those obvious variables. I know with a ML algorithm it's practically impossible to determine why the classification was made, but does anyone human here have any thoughts?
Could it be a version of this: https://hackernoon.com/dogs-wolves-data-science-and-why-mach...
> Both in real life and in our sample, the classification of political orientation is to some extent enabled by demographic traits clearly displayed on participants’ faces. For example ... white people, older people, and males are more likely to be conservatives. What would an algorithm’s accuracy be when distinguishing between faces of people of the same age, gender, and ethnicity? To answer this question, classification accuracies were recomputed using only face pairs of the same age, gender, and ethnicity ... The accuracy dropped by only 3.5% on average
Though cropping can only do so much.
I think the questions about age/sex/ethnicity are sensible in that it's a valid question to ask whether it's just doing the naive/obvious thing or something more. But if you keep on removing the less obvious things then of course you'll reach a point where it's no better than a coin flip because it's basically comparing blank pictures.