Google Photos AI still can't label gorillas after racist errors
theregister.com
theregister.com
At the same time, I'm surprised they can't get this classifier to work. It doesn't seem like a very challenging problem (I'd consider myself a computer vision expert). I wonder if they're over-reacting and just deciding to zero any residual risk by not allowing that classification anymore.
Identity politics aside, it would be an interesting study to try and break a man/gorilla classifier. Like take a picture of a man, say in a jungle setting and showing teeth or with a furry hat on, and see what the actual failure modes are. Regularly occurring misclassifications are a useful window into how a model operates.
Why are you assuming that they can't get the classifier to work? There's no evidence to that effect in the article.
> I wonder if they're over-reacting and just deciding to zero any residual risk by not allowing that classification anymore.
Why do you describe that as over-reacting? It seems like the appropriate level of reaction given there is no upside for getting it right and a massive downside for getting it wrong.
It doesn't even need to be a top-down edict, it's just the way the incentives will work out for everyone involved.
Let's say that you're an engineer working on that team, and find that there's a bunch of terms added to a blocklist a decade ago and still there. Are you going to just remove them, on the assumption that the classifications are correct? Of course not. How much validation work are you willing to do to convince yourself and others that the results are 100% correct? Is that really going to be the most interesting or impactful work you can do this quarter?
Similarly you'd find that whoever needs to approve such a removal would have incentives very strongly biased toward not approving the removal of terms from a blocklist. If everything goes well, they get no credit. If something goes wrong, they get the blame. It's possible to push through that bias toward inaction, but it requires it to be a change that somebody really wants to make happen.
It might be useful to understand why it's offensive... Likening black people to apes/gorillas/monkeys is a racist trope with a long history. It's part of casting black people as less than human, which has been used to justify their enslavement and denial of human rights.
I think it's very likely the classifier was not constructed to be racist, but once you see that it is producing racist tropes, it is racist to continue to let it do so.
It's like hiring someone to do plumbing. A few customers invite them into their home and report that the new plumber said racist stuff. As a business owner you do what after this?
And this seems to fall into the sad, overly pervasive racism of people not caring vs the more explicit “a human calling someone a gorilla is racist” category.
But also, it’s not like there’s many people who need to classify gorillas in photos so it likely doesn’t get brought up that much.
I understand that there's nothing to fix. It doesn't label anything as gorillas anymore.
Why waste effort in bringing it back, when another mislabeled photo will only serve for people to scream "racism" on social media? Better to not poke that particular beehive, I reckon.
If they cared about this problem rather than just thinking it’s a “beehive” to be avoided rather than a real problem, then they would fix it. But they don’t actually care about accuracy in classifying people of some races, so they just avoid it.
Not super racist, but definitely shows they don’t care about this type of bug in that the affected populations aren’t important to them. I understand not caring about gorillas, but I would expect they should be very interested in accurately classifying black people.
That is a very weird interpretation.
The classifier is likely accurate enough to be useful, with many edge cases where things are mislabeled. Edge cases might be difficult to fix without causing other issues, or perhaps the technology to be more accurate is just not there yet.
One of those edge cases causes a massive shitstorm on social media, where many people would love to scream "racism" at it.
So we have two possibilities:
1) Put an indeterminate amount of effort to make things more accurate. It won't bring any benefit (as the current classifier is accurate enough), and if you keep mislabeling black people as monkeys even in rare cases, people will still yell on social media. Beehive indeed.
2) Hardcode gorillas out of the classifier. This solves the issue with minimal effort and causes no shitstorm, besides some annoying people grumbling that it is still racist, but with no substance to their claims.
The choice, to me, is obvious.
You can actually find things the Google probably thinks are gorillas and monkeys by searching for "zoo".
I think intent is important for labeling something racist and a function doesn’t have intent. And it doesn’t seem like the programmer had intent.
So I agree that a human seeing a person and labeling it as a gorilla is racist, it’s because the human is making an inappropriate value judgement.
But this is AI -- so the stupid rule is not pre-programmed, but rather curve-fit to the data (uh, "learned").
So ultimately it's a matter of (1) failing to find the right training data (or procedure) and (2) more fundamentally, choosing not to correct the problem after 8 years.
I think this gets fixed by better training data and more pictures of really dark skinned people. So with more supervised labels of dark skinned people to people, properly so the matching doesn’t think people are closer to gorillas.
Comically/sadly, we’ll know we get closer to fixing the training sets to be more inclusive when google starts labeling gorillas as people.
I think there are some systemic reasons why there aren’t more diverse populations in training data. And those are more society issues than AI issues (ie, rich people are more represented, rich people are certain races, therefore races are more represented).
And finally, I’ve worked in software that people just test what they are and know so I’ve seen so many test plans that are too simple and only test the programmers dob and address. This doesn’t mean racist because all the programmers are Asian males. It just means the quality review wasn’t thorough enough to include proper test conditions.
I might be inappropriately conflating software bugs from different areas but this is what makes me think “stupidity or weakness more likely than racism.”
Is this the case? Do we even know for a fact that only non-white people were mislabeled as anything else?
Or are we just, you know, throwing out baseless speculation as fact?
Like that they don't appreciate the gravity of the problem, for example.
What would be racist from this outcome is if it kept doing this and no one did anything. Clearly it hurts people's feelings and that is a very valid issue. Googles option to just nuke it is a great start until they can hammer out the kinks.
Racism can also be measured by its effect, regardless of intent.
What would be racist from this outcome is if it kept doing this and no one did anything.
After 8 years, that's seems to be precisely what's happening.
Given the recent history of equating black people with non-human primates, and using that to deny them rights & full participation in society, making this error is going to be experienced as racist. It's not a matter of individual malice or taxonomic classification, but of history and social relations.
But it seems like if nobody is working on this, how will we ever fix this gaping hole in image classifiers? And don't we want to fix it? And to fix it, research will continue to get it wrong until they get it less wrong and more right, but can only iterate without a massive backlash. It seems like being stuck between a rock and a hard place.
I am rhetorically asking, wouldn't we have to allow researchers to iterate on this problem to fix it? That simply won't happen until we are able to allow them leeway understanding that this is an incrementally improving model. Otherwise what we have is just a sledgehammer solution (just banning all primate classifications) which actually never addressed the problem, that these models do have a race-based bias (probably in their input datasets.)
The thing you're calling racism is actually hate speech as it's typically defined in law.
The fact is that training sets usually contain many more white men than black women, especially if they're just scraped off the web. People who guided the training may have just used datasets that reflect their own culture and demographics of their own country, and didn't see a problem with that. The opposite would have been be seen as "pandering to diversity" in their country, so they've ended up with a biased dataset and a biased algorithm.
To quote from the article:
>> The company was criticized when a software developer, Jacky Alciné, found the image recognition system deployed in its Photos app in 2015 had mistakenly labelled a photo of him and his friend as gorillas.
People who are black have every right to complain about Black people being classified as gorillas, particularly as that comparison is a very common act of discrimination - most infamously in soccer, where e.g. Vinícius Júnior is routinely subjected to monkey sounds and chants [1] and the President of the Spanish Agents association literally told him to "stop playing the monkey" as a goal celebration [2].
[1] https://www.cbssports.com/soccer/news/vinicius-junior-faces-...
[2] https://www.dailymail.co.uk/sport/sportsnews/article-1121907...
https://www.mdpi.com/2227-7390/9/2/195
> "Studies have shown that due to the distribution of ethnicity/race in training datasets, biometric algorithms suffer from “cross race effect”—their performance is better on subjects closer to the “country of origin” of the algorithm."
However, these classifiers are also problematic as racial identification is a bit thorny politically. It's comparable to gender identification classifiers in the context of trans politics (as, regardless of personal identification, there are some classifier-accessible differences between the vast majority of female and male faces, although I suppose there's an androgenous middle ground).
To go out on a limb a little, imagine an 'Aryan vs. Jewish' classifier - who wants to be associated with that? Lots of downsides for any commercial outfit, certainly, and few upsides.
Because of the entanglement of people who harbor individual racial animus and the systemic factors that lead to disparate outcomes for people of different races we are in a bad situation vis-a-vie language, but I think this isn't quite right.
I agree there's no reason to believe the engineers involved had any prejudice (i.e. individualized animus against non-white people). It's the result of under-testing or corner cases or whatever. However, the system is definitely racist.
It's racist in the sense that it has an adverse impact on people of different races (white people will not be mis-detected as gorillas) and its racist in the sense that there must have been racial imbalances in development (one struggles to imagine this being launched if all humans were being detected as gorillas for instance). This doesn't require any prejudice on the part of the people making it - it's just a fact that we can observe that it has disparate racial impacts and that choices in its development have led to a product that is unevenly useful to people of different races. It's also clear because facial recognition systems perform worse across all measures on people with darker skin (not just on this metric) - it's not like the accuracy rate is the same but the mistakes are different.
Also, of course, yes - people will point out that different skin tones interact differently with photographic technology and algorithm designers are fighting an up hill battle. It's true! We have historically been good at solving the challenges of capturing images of light skinned people better than dark skinned ones. That's racist too! It's hard to not build a racist system when your tools already have racial bias. Again - pointing out the racial bias in this system is not an indictment of the people who built it. It's an indictment of the world it emerged out of and a bookmark to come back to when you see people using the outputs of the system.
This is sort of like telling residents of a historically black neighborhood which just so happens to have been exposed to drinking water with dangerously high levels of lead since time immemorial (despite decades of complaints) that it's just "shitty pipes", and that it's "idiotic" for anyone to think this state of affairs could possibly have something to do with racism.
I agree that the matter is quite nuanced, and there was no racist intent. However it's far from "idiotic" to address these issues directly. If anything it seems a bit obtuse not to address them.
Where there's smoke, after all, there's usually some form of fire.
The classifier is not the result of any racist action on behalf of the programmer.
It’s more like blaming a hospital administrator when people of a discriminated race have bad outcomes at their hospital, when the outcome is the systemic rate of all the other hospitals. The hospital outcome isn’t racist, because the systemic issues that affect the outcomes of the hospital patients have nothing to do with the decisions of the hospital.
It's the fact of this problem not being remediated in certain neighborhoods that cause people to ask, very legitimately -- how this came to pass.
There doesn't seem to be any benefit to solving this problem, so how is it an overreaction given the negative press?
But really it is history that is to blame. Gorillas are great. But if you called anyone a Lion or Lioness it would be perceived as a compliment. Gorillas don't deserve such connotation.
That is to say: not so challenging that they couldn't have got wrapped by now, if they cared enough to do so. But not necessarily trivial, either.
The lack of contrast can make things hard the evaluate, especially if you're dealing with SDR content in a format that struggles to accurately encode darker colours.
It's not racism, it's a lack of information to the evaluator and a failure of communication by the company.
If they actually put out a statement regarding the technical issues around the problem, rather than trying to bury it in order to minimise the backlash, we could be having some far more reasonable and level headed discussions.
It's unfortunate. And I can see why it worries people. Because on the surface it looks like a systemic problem. But is it really?
I think most here would understand systemic bias if reflected in a distributed system. Imagine a classifier where, through no explicit intent of the designer of the classifier, always said that some user login was malicious or some binary was malware or the like, even when it was not. And all those users or binaries had something in common that was apparent, and were all being misclassified. That would be systemic bias, and nobody would blink an eye about a call for it to be corrected.
Here we see systemic racial bias, probably not because of individual racial animus on the part of some Google engineer, but because society has systemic racial bias and the classifier is just reflecting that because it was trained on public data.
We should look at incidents like this differently: not as yet another chance to argue about what Google should or shouldn't do, but as a chance to see that a socially-neutral AI that does not have bias built into it is learning racial bias from society.
Come up with a different word. Maybe “non-uniform variance”.
Are you suggesting that the cause here is that there are a lot of racist images on the internet visually depicting black people as gorillas?
I assume that is not the understanding of most people here who are questioning what makes this racist, because if that is indeed the cause then I agree it is obviously racist.
Is it the cause though? I would be interested to see some evidence or rationale, because I don’t think I have ever seen an image like this outside of perhaps a handful of very old hand drawn propaganda images (which presumably would not affect the classifier much due to being overwhelmed by the vast number of images of real people and real gorillas).