Amazon’s Face Recognition Falsely Matched 28 Members of Congress with Mugshots
aclu.org
aclu.org
It also possibly highlights the fact that algorithms are not immune to bias when they are designed by humans. Obviously I don’t think this bias is intentional, but so much of it isn’t and happens anyway.
I would wager "absolutely not" for the exact same issue as firearms or cars. The person who mis-applied the technology or a middle man vendor though? Unless they fall under qualified immunity, there's your fall guy.
Almost every engineering discipline has a code of ethics [0][1][2][3]. It's time software "engineering" grew up and did the same.
I rarely see ethics mentioned on HN, and granted, people's view differ. But it's weird we're not having that conversation at all.
[0] https://www.raeng.org.uk/policy/engineering-ethics/ethics
[1] https://www.ieee.org/about/corporate/governance/p7-8.html
So. Let's talk about ethics. I, personally, subscribe to the ACM code of ethics.
I think Rekognition, as built and presented, falls fully within that strict ethical code. It can be put to uses that are unethical, but that does not fall upon the people who made it. Certainly, an engineer creating a system such as the one the ACLU created would be acting unethically.
Will that slow things down and make them more expensive? Absolutely, but that’s the cost of safety.
In this case though what is probably most lacking is public knowledge about the shortcomings of such systems. The public needs to understand that 90% accuracy, while it sounds high, mean it’s wrong a lot, and the chances of it being wrong if it was continuously run all the time are actually quite high.
Car manufacturers are sued for incompetent use of their products.
- Show us the side-by-side images of the false-positives. Are the matches plausible?
- What is the demographic distribution of the mugshot database? If the data is disproportionately biased, then that bias would be reflected in the false-positives. A casual skimming of some mughot websites shows a potentially significant racial bias.
At the same time, they do still have the correct final conclusion. Dragnet facial recognition is a bad idea that will produce too many false positives to be of use. Or at the very least, dragnet facial recognition used by people who put too much faith in it is a bad idea. It's at most useful to produce flags for slightly more attention; on its own, I would not consider it anywhere close to enough to arrest someone on.
And it can serve as a demonstration of how a clumsily-put-together system can be used to produce bad results, so, if we correctly assume that there may be people bodging together a system in much the same way the ACLU did, well, hey, it's a valid result then!
Perhaps lawmakers should consider implementing rules to ban facial recognition software until the developers can prove some like 99.999% accuracy. Otherwise we end up harassing a large number of people, and the credibility of the technology will stuff. As it stand any lawyer will be able to argue that facial recognition isn't admisible as evidence, do to the high margin of error.
You're right that systems have previously learned to infer black -> criminal. It's absolutley possible that could have happened here! Given that not every one flagged is a person of color, there's clearly more at work.
Systems like Rekognition let a user specify what level of certainty the system has that it has found a match. In this instance, the system was told to use 80%. Rekognition can be told to use any number from 0 to 100. Your number of 99.999% could be used without any difficulty whatsoever! Since the ACLU doesn't seem to be sharing enough to reproduce the test, we can't tell what would happen.
Regulation based on accuracy is a wonderful idea! One well worth pursuing. With that said, "accuracy" is perhaps difficult to define here. How would you determine that? How accuracy is measured is highly sensitive to what you use as a dataset, and different test datasets can easily produce very different impressions of how common false positives and false negatives are.
Yet, might there not be other factors to consider? Bear in mind that this isn't pushbutton easy off-the-shelf. Some actual work is required here. There's probably a batch of Python scripts sitting around that did the work. The most charitable assumption is that whoever did this at the ACLU didn't look at how the tool under examination should be used at all and missed everything about confidence levels in whatever few docs they did read. Under the ACM code of ethics, I am comfortable calling that an unethical degree of incompetence and carelessness.
A less charitable interpretation is that they knew and didn't care because it produced the result they wanted. This is similarly unethical, in my mind on a level with the sort of p-hacking prevalent in social sciences.
You're completely right! They just used the default. It's just perhaps possible that there might be other factors worth considering.
Same thing re data setup. Not the point again. Even if law enforcement hires the world's best computer vision experts (fat chance, it'll go to cheapest contractor) this is a note that there can be mistakes that get people killed.
Regarding sketches, they're actually interesting because they force a human to exercise judgement, being ambiguous by nature. This is fine - any law enforcement officer stopping someone who looks like a sketch is going to give them the benefit of the doubt. On the other hand, if a computer tells an officer a person ahead is an 89% match with a known killer, the conversation will start with guns drawn, and that will greatly increase the likelyhood of things going south.
You really don't want to be scanned by a trigger happy cop having a bad day. God knows who you happen to look like from that angle.
That is where the two sides of this argument diverge. First: That things "will start with guns drawn". And Second: That things are more likely to "go south" for an innocent individual incorrectly blamed by this.
If police engage a target with a weapon already drawn, there is a long list of things that person can do that will get them shot, and a very short (and often unclear) list of actions that will keep them safe.
Source: 10 years PMO in the Marine Corps, 1 month SFPD (then quit because their training was so shitty)
For the most part, this sort of article makes things seem like there will be some sort of racial-based, dystopian police future where automated algorithms will target people and have them put into jail. I don't see that happening in the near future at all. The police departments' personnel currently using this probably LOOK at the matches and only then make a further call. Just like I would assume AFIS fingerprint hit results are double-checked by skilled operators after matching to reduce false-positives.
>Reached by The Verge, an Amazon spokesperson attributed the results to poor calibration. The ACLU’s tests were performed using Rekognition’s default confidence threshold of 80 percent — but Amazon says it recommends at least a 95 percent threshold for law enforcement applications where a false ID might have more significant consequences.
Presented without real comment on my part.
> A confidence score is a number between 0 and 100 that indicates the probability that a given prediction is correct. In the tropical beach example, if the object and scene detection process returns a confidence score of 99 for the label ‘Water’ and 35 for the label ‘Palm Tree’, then it is more likely that the image contains water but not a palm tree.
> Applications that are very sensitive to detection errors (false positives) should discard results associated with confidence scores below a certain threshold. The optimum threshold depends on the application. In many cases, you will get the best user experience by setting minimum confidence values higher than the default value.
Which I am interpreting as similar to your question, but perhaps should not be understood as a statistical probability.
They should rerun this at 95% confidence and see how many mismatches they get, I suspect it will be much lower than the 6.4% they have at 80% confidence, thought it might cost more than the $12.33 they were willing to invest in this hitpiece against rekognition.
That's not in https://aws.amazon.com/rekognition/faqs/ - while that threshold may minimize false positives, I'm curious if it was in any documentation that the ACLU saw. Or, to the bigger point, is it in any public documentation?
They were slow to let go of phrenology as well.
It can say exactly what the Amazon response to the Verge dismissing the ACLU test should, and, in fact, it should say that if Amazon really does have such critical, specific domain-specific recommended settings, rather than it being something made up off the cuff to mitigate bad press.
With that said, it's possible that no amount of clear guidance would have mattered here. The ACLU isn't exactly asking for better docs.
If you're a researcher making a deliberate choice to use the defaults to simulate a random police officer with no real understanding of their tools, then say as much. And compare with more cautious results.
That’s not how these numbers fit together. Using rekognition, you would create a collection of faces based off your image corpus and then ask the service if a provided image has a match against anything in your collection. Lowering your confidence threshold, invariant of your sample size, is by definition going to increase your false positive rate while lowering your false negative rate.
I think it would be worthwhile to compare a mixed test set and see false positives/negatives at various confidence thresholds, because the fear (IMO) is that law enforcement will minimize false negatives at the expense of a huge increase in false positives.
Falsely arrest 28 members of Congress due to poor face recognition, and the problem of facial recognition in law enforcement is resolved the next day.
Congress is a battleground. Not a club or secret society. You send representitives from your state to try and get a slice of the pie, and I send my reps from my area to try and get a slice for me.
They are trying to make this racially charged without giving enough information to verify their claims. If you use a dataset of mugshots, that's statistically going to have more data on people of color. If you have more data on people of color, it is more likely to match people of color. Claiming the algorithm is racist because your data is racist is inflammatory bullshit.
The whole racism angle of the article seems way too loaded for me.
Moreover, if a algorithm enforces bias to the detriment of inoccents, it's a bad algorithm
I actually did not know this. Thank you. I had conflated incarceration rate with inmate population. If anyone is curious for a source, see here: https://www.bop.gov/about/statistics/statistics_inmate_race....
> Moreover, if a algorithm enforces bias to the detriment of inoccents, it's a bad algorithm
I agree with this 100%. But they haven't provided enough information about their dataset to make this conclusion. If they provided enough information for someone to independently analyze the data and reproduce the experiment, I would flip immediately.
Personally, I would assume the mugshots are a random sample of a public corpus, and thus likely 35-40% people of color. Which is definitely skewed from the background population. I'm assuming the similarity of 40% is a coincidence.
> Moreover, if a algorithm enforces bias to the detriment of inoccents, it's a bad algorithm
Is it? It sounds like it might be working exactly as it was set up and specified to work. I would call that the fault of the designers and specifiers, rather than the algorithm. It's a good algorithm. It fits its purpose well. Its purposes just happen to be completely evil.
Unless the designers of the system are aware of and account for biases in the training data, the system will be biased and will be biased due to the actions of its designers.
I don't see what's so hard to accept about that. It's literally the oldest problem in computing (going back to the "Pray Mr. Babbage, if we put into your machine wrong figures, will the right answers come out" days).
The problem, of course, is how woefully underprepared the average tech person is for realizing the existence of even extremely obvious biases, let alone more subtle and pernicious ones, and a tech culture in which people are lauded for being knee-jerk contrarians on topics like "does our society have a history of systemic racism".
This whole topic is racially charged because incarceration rates are a race issue and you are being disingenuous with your comment pretending the problem with facial recognition can somehow exist and be fixed in isolation.
I'm not saying the topic isn't a race issue, I'm saying the ACLU isn't providing the data for someone to verify their results and is possibly making false claims that no one can verify.
Oh, wait: https://www.usatoday.com/story/tech/talkingtech/2018/06/29/c...
The technology is not perfect. And should never be used as evidence of a crime, just as an indication that two images might be of the same person and to have humans look at them.
And then those humans will also make mistakes. People do look alike. And there are twins.
I think a jury should require more evidence than just similar appearance. But such a match is a strong indication of where to look.
This is especially important as automation enters the equation.
How many people have been falsely convicted because "DNA"? 1B-1 odds of a match, "and sir, yet you claim you were not even in the area?".
There should be a way of recognising the parts of the science that are basically correct and the parts that are either less reliable, open to bias or could simply be broken due to incorrect process/mistake in the lab.
https://www.newyorker.com/magazine/2009/09/07/trial-by-fire
> Many arson investigators, it turned out, had only a high-school education. In most states, in order to be certified, investigators had to take a forty-hour course on fire investigation, and pass a written exam. Often, the bulk of an investigator’s training came on the job, learning from “old-timers” in the field, who passed down a body of wisdom about the telltale signs of arson, even though a study in 1977 warned that there was nothing in “the scientific literature to substantiate their validity.”
> In 1992, the National Fire Protection Association, which promotes fire prevention and safety, published its first scientifically based guidelines to arson investigation. Still, many arson investigators believed that what they did was more an art than a science—a blend of experience and intuition. In 1997, the International Association of Arson Investigators filed a legal brief arguing that arson sleuths should not be bound by a 1993 Supreme Court decision requiring experts who testified at trials to adhere to the scientific method. What arson sleuths did, the brief claimed, was “less scientific.” By 2000, after the courts had rejected such claims, arson investigators increasingly recognized the scientific method, but there remained great variance in the field, with many practitioners still relying on the unverified techniques that had been used for generations. “People investigated fire largely with a flat-earth approach,” Hurst told me. “It looks like arson—therefore, it’s arson.” He went on, “My view is you have to have a scientific basis. Otherwise, it’s no different than witch-hunting.”
If the algorithm is, say, 99.99% accurate on faces that are the same race, gender, and roughly age, and there's 100 faces in there that match those, it'll correctly say no match .9999^100 = 99% of the time. If there's 5000, it's only going to find no match .9999^5000 = 60% of the time.
? That doesn't add up to the rest of your comment.
Presumption of Guilt is the new normal, isn't it?
When the gov steers opinion, we call it manufactured consent, when public advocacy organizations engage in sloppy methodology to further a cause, I propose calling it manufactured outrage.
One of the perennial problems with machine learning is that it has a tendency to intensify pre-existing biases in the system. It's not just that the algorithm tends to reflect biases in the source data, it's that that reflection tends to encourage the people using the system, who generally have some role in creating that source data in the first place, to become even more biased.
I haven't made much of a study of primary sources, but I wouldn't be surprised if a common thing that happened every one of those times was that people would start trying to figure out how to turn "appeal to acrimony" into a convincing debate tactic.
Bone marrow transplants for cancer treatment used to have a 20% failure rate. I didn't think for a second that we should cease using bone marrow transplants.
And yet why does the ACLU think we should cease using a technology to deliver potential location hits on wanted criminals because its not 100% perfect?
The cost associated with being just 5% wrong is huge, and that's not including the emotional damage done to falsely accused citizens, or the decrease level of trust in law enforcement.
And you've highlighted the _exact_ problem here. That is not at all what has happened, they have been identified as "possibly having a similar appearance to one or more particular images of a suspected criminal."
> And an officer can review the evidence presented by the report and make a human determination to follow up with a physical arrest.
Which already presumes that the system will only have false positives. When the system produces false negatives, then this mode of investigation is entirely flawed.