Federal study confirms racial bias of many facial-recognition systems
washingtonpost.com
washingtonpost.com
So while some ethnic minorities might be more likely to match incorrectly to a known person of interest, they're also more likely to be let through if they are indeed that person. I think that the first case is definitely more damaging on the whole, but I still find it misleading to not mention the specific scope of the "racial bias."
The fact that it produces errors in both directions just makes the system even worse in total.
The reason it's relevant is that false positives are always a trade off against false negatives, so the "easiest" way to reduce false positives is to increase false negatives. But false negatives are also very bad, so we can't really do that or we're just trading one problem for another instead of actually solving anything.
The only real solution is to improve the overall accuracy of the system, but that's easier said than done. Some of the main sources of the inaccuracy are intrinsic. The population in question is a smaller proportion of the general population so there is less training data available. People with darker skin absorb rather than reflect light, which makes it harder for the system to identify them.
So the problem isn't some racist schmuck who just needs to be fired and the system will get more accurate, it's a consequence of demographics and physics. There may be a solution somebody can find, but there isn't guaranteed to be, and in the meantime there's not a whole lot you can do other than trying to find one.
Or maybe discontinuing the system entirely.
The principles of the law makes it clear that it's far, far worse to have an innocent person in jail that a guilty person go free. Many legal principles revolve around this.
Yes, absolutely. But there are two issues in this context.
The first is that facial recognition isn't a conviction. Presumably if it identifies the wrong person then they'll have the opportunity to demonstrate who they are, e.g. by showing that the name on their ID is not the name of the suspect, before being convicted of the crime.
And the second is that ordinary application of that principle is incompatible with eliminating the disparity, because if you shift in favor of false negatives in general then it reduces the false positive rate for all races but the proportions stay the same. But if you shift it only for one race then you're only making the existing racial disparity in false negatives even worse, which is hardly fair to minority communities because it implies they'll have a higher proportion of criminals on the streets victimizing their communities.
That being said, in this context having a preference for false negatives over false positives in general probably is a good idea, and it does reduce the absolute (rather than proportional) racial disparity in false positives -- which is not nothing. But it's the proportional disparity that people are always complaining about.
Are they? Isn't our entire justice system premised on the idea that false negatives are far more acceptable than false positives? There's a problem when your false negatives are systematically racially biased, but beyond that, false negatives seem like what you'd expect and desire from such a system.
Some might object to the decreased accuracy this gives whites — but if we’re going to have any hope of preventing existing systemic inequalities from being encoded in AI systems, we’ll have to socially engineer the data sets.
There are many problems with socially engineering data sets. For one we would need to achieve consensus on what constitutes an appropriately representative (i.e. socially just) data set. And then we need a feedback loop to make sure that our dynamically changing inputs don’t knock our meticulously tuned output out of whack. And we’d need to continually adjust our model based on our changing definition of social justice. Finally we need to find social engineers with credentials in building perfect societies (I most certainly don’t trust Joe PM) to perform all this work.
Sociologists have been arguing forever that humans are inherently racist. Maybe these “AI systems” are just proving their point. Perhaps this bias is more a reflection of our inability to maintain global context that is also maximally relevant within our local tribes. Maybe we should focus on acknowledging biases and working through them rather than trying to transcend into some unicorn world where everybody has all the context all the time and bias is impossible.
Anyway, I see your census and raise you an FBI: white people commit way more crimes in the US [1]. So if your goal is to reduce total harm then social engineering is not your solution. Perhaps more accurately identifying the most frequent perpetrators (in the case of the US, white people) is exactly how you’d end up tuning your system...
1: https://ucr.fbi.gov/crime-in-the-u.s/2010/crime-in-the-u.s.-...
The difficulty with this is that you commonly don't have any control over the training data. Someone else collected it and you can use it but you can't choose what's in it.
And in some cases you already have census-level data, which means you couldn't even increase the number of people with lower representation in the data set by collecting more data, because there are no such people not already in the data set. The only way to have equal representation would be to throw out data you do have on the groups with higher representation. Which is obviously objectionable because we need to be improving accuracy for the group with lower accuracy (if we can), not reducing accuracy for the group with greater accuracy just to make the numbers the same.
So to fix systemic inequalities for some groups, we're going to increase systemic inequality for certain other groups?
Definitely a step in the wrong direction as it's just going to swing hatred back the other way. And N-levels later of this "fiddling" you'll end up in a spot where neither side can reconcile with the other because each side has legitimate reasons to feel slighted and no one can look past that "hatred". Just look at the mess that is currently in the middle-east after people "fiddled", and then "fiddled" again to try fix it.
No; in order to accurately identify people through appearances, we'll try to sample evenly throughout the range of appearances and features instead of sampling based on proportions within the population of the people setting up the system.
You're can't be seriously making the case that white people are such an overwhelming portion of the world's population that not concentrating on them in your training data constitutes racial discrimination. Not only racial discrimination even, but is a attack that will inspire a "legitimate" counter-retaliation.
This has wider implications and tacit-meanings that I think I picked up on and definitely implies doing more than just having an equal sampling of each racial group. How that would be enacted, I don't know. The point is to not swing the pendulum the other way in order to "fix things" as it'll cause problems down the line, but to rather just "fix things" in the most non-intrusive and fair way.
Yes, it's kinda unfair on some level that the algorithms used don't work so well on women/blacks/children/elderly as well as "white middle-aged men", but it's not unfair in the "government should do something about it" kind of way just yet. If someone gets misidentified, that doesn't mean police should just go out and "arrest" them. Rather we add a person into the mix and have them talk to that person, make a final call. That is actionable counter-case data that we can use to resolve the bias in the training sets.
A lot of facial-recognition systems include human components in order to resolve errors. E.g. I've worked with facial-recognition software here in South Africa that was developed by Asian companies where the majority of their data-sets and tests were using Asian individuals, so one has to manage error risk using humans in the loop.
One of the things you have to look at is "how" that bias may manifest and if bias can even be applied in any negative sense in the realm of control you give to the user. If you give someone a task where they are to compare two faces, and they don't like the demographic of that individual, I suppose they might pick incorrectly (because you only gave them true/false type options). But you can build in review processes, reporting, performance feedback, etc to mitigate that. You need not treat it as bias at all, merely bad performance/mistakes that need correcting.
I'm not sure adding more humans to the mix is a solution to solving bias problems because humans are unconsciously biased and may be perversely incentivized.
I don’t mean to imply that’s the right metric at all, of course, because it specially privileges minorities. The proper measure should equally weight the algorithm’s performance on all people in the population, or at least weight them irrespective of race. That’s just not a measure of fairness between races.
It's almost like the purpose of the article isn't to show the tech is overall empirically bad.
If they're disproportionatly committing crimes, skin color becomes a strong prior. It would be rational, iff the goal is to lower crime rates. But I don't know that it is morally OK.
Also I'm not sure if over representation is the source of bias in this particular case, haven't read the article yet.
Edit: yeah, this article is just a smear, wapo is salivating over anything they think confirms that police are racially biased. It doesn't mention anything technical about the training data or typical nets being used. I'm not advocating for facial recognition tech, I'm personally against it, but racial bias coming from objective data is the least of your potential worries.
Briefly, bias is the wrong mean, and inaccuracy is high variance.
So, by your statistical definition, there's a bias.
And even if you get enough photons to, for example, recognize both black and white equally well in daylight, you might not at night or in a deep shadow indoors, or up close equally well but there will be some distance at which the advantage of more photons begins to matter, or some speed, or some combination of factors.
Any technological improvements will help, but there won't be any technological solution that will always work equally well with fewer photons as with more.
Plus human facial recognition is massively overpowered for normal cases. I think if you tested it in challenge cases (through blurry video, or snow, or in the dark, or at a distance), you might see darker skin tones being recognized less easily. It'd be an interesting experiment.
You can't display all of this range to people because of the limitations of display technology, but you can feed the full range to a facial recognition engine.
Security budgets might not stretch to top-end HDR equipment but the price keeps on coming down. The performance of a modern flagship phone is remarkable compared to a few years ago - and fixed surveillance cameras can have much bigger glass and sensors, making it cheaper to get super-human performance.
One new but related issue is that body-worn cameras can capture more low-light detail than the human eye. Police unions have argued against deploying these sensors, because, they want the evidential record to show what the officer could see - not what a cat could see.
While there isn't evidence that dark-skinned people have harder-to-recognize faces, there is also no evidence to the contrary.
It seems like the null hypothesis in this case, given no additional evidence, is to assume that darker images are indeed harder to recognize.
Citation?
Dark cars have more accidents- https://www.telegraph.co.uk/motoring/news/7845366/Black-cars...
There clearly will be a difference.
Whole careers are around making things more or less recognizable depending on color.
What study makes you think it's insignificant for human faces?
The linked study on cars is analogous at best, but doesn't prove anything for human facial recognition. Cars on highways and human beings in various settings are extremely different circumstances.
In the early days of color photography there was an issue with some films and reference images being tuned for the most common subject (light-skinned humans) and as a result if you tried to capture a mix of races you'd get bad results: https://petapixel.com/2015/09/19/heres-a-look-at-how-color-f... This makes sense if you consider how light is (generalizing here) a broad spectrum of hues and a given material is most reflective for specific parts of the spectrum, so if you don't capture much there you'll get a low-contrast image, like stripping the R channel out of an RGB bitmap.
It's of course possible to solve the problem for a wider set of skin tones, and it has been solved, but it takes more work. It's a subject of ongoing discussion/experimentation in film to this day: https://www.konbini.com/en/cinema/insecure-cinematographer-h...
Basically, your team is composed of white dudes who don't see the problem with a ML training set consisting largely of pictures of white dudes.
To prevent this they'd have needed to A) Employ a black person, and B) Listen to said employee's feedback, in order to recognize the problem.
Edit: Also worth pointing out, just using a representative population sampling would still show racial bias, essentially weighting accuracy with respect to population percent. You'd probably need to have equal samplings of pictures of people from all races/genders/disabilities if you wanted equal accuracy across the board. That also includes picture quality and range of picture quality. Doubling up images, or using corporate headshot white dudes and grainy selfie People of Color could still cause issues.
Same logic applies to labelling. That minimum wage contracting firm used to decide who's who in the photos may exhibit racial bias, by virtue of the fact that most people do. If their accuracy in labelling is racially biased then so too will the algorithms that it's based on.
In short: Racist garbage in, racist garbage out.
But more to your point, if these oh-so-competent ML practitioners were doing their jobs right, we wouldn't be having this discussion.
The whole reason diversity is championed in hiring is precisely because a single individual's perspective can only see so far. And if you have a monoculture team who has experienced very similar life circumstances, you end up with the kind of narrow perspective that leads to more racist soap dispensers.
And not to say this happened in your case, but even with that considerable effort, it's still very easy to end up with blind spots in your product that a more diverse team would have caught.
It's the same as hiring for any other level of experience for more routine technical skills. If your team has no experience in this area, they'd need to expend a much greater degree of effort to answer questions that someone who is experienced would already have known the answer to.
In other words, ideally the racial and gender distribution of a team would be as inconsequential and unbiased as blood type or handedness, in that the aggregate demographic ratios on your teams would at least match that of the residential population in your area, and ideally that of your broader geographic location.
I'm not doing a good job explaining this clearly, but the simple answer is: more than one. No one wants to be the token hire.
Ok, I don't know about race, but for gender look up the "gender equality paradox". In countries with greater equality rights for women they show less of an interest in STEM subjects.
https://en.wikipedia.org/wiki/Gender-equality_paradox
Like I say I don't know of any similar studies done for race, but it would indicate that you shouldn't necessarily expect outcomes that "would match that of the residential population in your area, and ideally that of your broader geographic location".
In my opinion we should be pushing for equality of opportunity, not equality of outcome (you appear to want the latter).
Realistically, the most that hiring managers (save for huge FAANG institutions) can do is thoroughly ensure that their team isn't inadvertently (or blatantly) racist/sexist in their hiring process and on the job, and to post the job in enough places that a diverse applicant pool will see the posting.
With that said, hand-waving away that there are few to no women or African Americans/Latinos/Native Americans/etc. on the team with an overzealous application of the Equality Paradox is a pretty dangerous mindset to get into. It's essentially passing the buck, and is eerily reminiscent of the claims made by 1950's Southern US Politicians that Blacks were the ones who were self-segregating because they wanted to, not the other way around.
What I'm saying is, the ideal 50/50 gender ratio/representative race may be unrealistic for a myriad of reasons, but if you're a 50-person start-up with 2 women, one of whom is HR, and no black people, I'd take a good, hard look at the company culture that's being fostered, and particularly whether turnover for women and People of Color at your company is worse than average.
I have been involved in hiring people before. There was absolutely nothing racist or sexist in the way we hire. Fact was we go two applicants. Neither were women or minority status. Fact is that the industry is full of white men (even here in Europe).
We should be hiring on ability to do the job and nothing else.
> We should be hiring on ability to do the job and nothing else.
This is exactly my point! Yet there is quite a lot of inadvertent, or even blatant, racism and sexism that happens during the hiring process and on the job.
Forgive my ignorance, but this seems to lend credence to the popular idpol claim that the white guys programming AI are ignorant of the inherent bias their models might have.
These are two bold assumptions. Would you really have us believe that this is the case? I think you will find that some important information has been left off. The article suggests that asian created algorithms are better at recognising asians. It implies that they all have problems with darker skin tones. Why does the article not explore the reasons why? Perhaps they are wanting to say that algorithms can be racist, rather than the truth of the matter, which is there are certain technical issues that are difficult to overcome with darker skin tones.
The first sentence of the article backs up its usage of the word bias instead of inaccuracy:
> "...casting new doubts on a rapidly expanding investigative technique widely used by law enforcement across the United States." emphasis mine
Can you point to any references supporting this claim?
It's absolutely irrelevant whether the heuristic is unfair to individuals of a particular race, or unfair to individuals of a particular nose length. Prioritizing the former over the latter is a totally arbitrary value system that deprioritizes the most important metric for maximizing fairness, which is the overall accuracy rate.
Ultimately any inaccuracy of the heuristic is an instance of the heuristic treating someone unfairly. The objective should be to minimize the inaccuracy rate overall, not the inaccuracy rate in relation to politically prioritized groupings like race.
'relationship “between an algorithm’s performance and the data used to train it,”'
To solve this problem they should look at the relationship between the number of pictures for a given race, and the accuracy in recognizing members of that race. It probably does best at identifying European descendant faces because it was developed in a country where European descendants are the largest ethnic group. I'd bet dollars to donuts that if you can increase the number of faces for each race this disparity will largely disappear.
Quick, somebody offer the guy who realized this a job.
Not blaming this group or anyone in particular – we all just need the be cognizant of the fact that we have our blind spots (racial, gender, and otherwise).
Actually, as far as I can tell, it's dominated by asians - specifically Indians.
https://www.revealnews.org/article/heres-the-clearest-pictur...
The ACLU rep that is quoted is IMO relying on hyperbole to make his point (“One false match can lead to missed flights, lengthy interrogations, tense police encounters, false arrests, or worse”). If the match is close enough that a human would confirm it as well, and conduct an “interrogation”, then presumably facial recognition technology did not incrementally add to the problem of mismatching since a human would also make the same false match.
So if it makes policing more efficient and enables more criminals to be nabbed, and it is used as a first pass filter (so the false positive rate doesn’t matter), I am all for it.
it seems likely to me that cops might use it as justification for a fishing expedition, or TSA agents may just go along with it as a cover-your-ass measure
“Sir, this dog detected something on your person”
“Sir, the scanned detected an object under your shirt please step over here”
“Sir, you’ve been flagged by this facial recognition software”
... agreed. I think this is better to head off before that last one can be abused.