Thoughts On Machine Learning Accuracy
aws.amazon.com
aws.amazon.com
What'd they think Amazon was going to do, roll over and be like "turns our our facial recognition software is racist, woops!"? Now instead of a meaningful dialog, we've got a line in the sand where on one side, the ACLU, champion of the people's rights, doesn't understand technology, and on the other side, facial recognition software is Bad and Evil. Shits so polarizing these days it seems there's no room for negotiation, as much as I'd like there to be.
Did they? This blog post describes some implementation choices which could make their false positive rate lower.
It reveals no information about how Rekognition is being deployed by LEAs, and there’s no meaningful regulation or oversight about how that happens.
We can’t tell what the assumptions and implementation choices of local police departments are, because that’s top secret information. Most people acting on the predictions have no idea what they are, let alone the implications.
It’s fine to ding the ACLU study on methods —- you can, because they published some.
No-one would argue that a toy model based on members of Congress is representative of the public at large. But it’s more representative that a press release from a company that are trying to normalize and monetize mass surveillance via facial recognition.
It puts directly the blame on Amazon and its technology and doesn't mention any configuration they used.
They used the default settings, and a training set which seems extremely likely to be used in ‘LEA-production’
Any user, including LEAs, is free to use which ever configuration options they like.
The criticism with the most statistical implications is the confidence level used. The ACLU were clear they they used the defaults. I can take at face value that moving from 80% to 99% confidence on a sample of 500 faces could produce 0 false positives. However, on the faces of the tens/hundreds of thousands of people that might move through the center of town on any given day, the implications for causing ordinary people serious harm are large.
I am personally shocked that Amazon are willing to “recommend” law enforcement actions at any level of confidence, from an unintrospectable machine learning system.
I would imagine the confidence level would be tweaked by Law Enforcement or anyone else based on the results. If you're monitoring cameras for missing persons or shooting suspects and getting 0 hits, you might lower it because it's better to dismiss false positives than the alternative. Conversely you might increase it if you're overwhelmed by the number of false positives.
I think there should be a discussion about laws/regulations about facial recognition use. Even if Amazon removes their service there will always be companies that don't care about the ethics as much. Strict rules instead of public shaming seems like the better approach
If you screw up with facial recognition a human double-checks and nothing happens.
Scissors are dangerous but we don't regulate those either.
First of it is not a given if you "screw up with a gun someone dies" that is not the only result
Further it is laughable to believe that with facial recognition a "human double-checks" we have seen time and time again that new technology entering the realm of criminal justice can be used to wrongfully convict people at an alarming rate. See Hair and Bite mark controversy or the number of Labs that have been compromised in recent years
The fact is that a "humans" that are suppose to be "double checking" do not in reality, if the computer spits out "Well the computer says it is person X" that is what the Jury will hear, and the Jury will not question the technology and the person will end up in prison until such time as someone like the Innocent project comes along to invalidate either the technology on the whole or the application of the technology
You have too much faith in humanity: a lot can be, and will be done "because the system says so". Nevermind integrators who will create processes that assume recognition is 100% correct "Yes, I know you're not the wanted criminal, but now that I have tagged you, I have to take you to the station and process you out...Monday morning because that's when Jim is in"
Only if <gun maker> was specifically marketing their hunting rifle to be aimed at members of the public, and for the trigger to be frequently pulled. Happily, gun-makers are far more responsible than that.
> It's a tool, you're responsible for how it is used, not the manufacturer
Not so. Private businesses can refuse to offer services largely as they wish, and AWS frequently do, based on their terms of service.
For example, you are not permitted to run a web-scraper or open mail relay on EC2. You are not permitted to host bestiality on S3. None of these things are illegal. But Amazon have taken the stance that they just don't want them to happen on AWS.
I am of the view that mass facial recognition poses more problems for society than running web-scrapers.
Many civil liberties organizations have serious concerns about the use of facial recognition for law enforcement, including pro-industry organizations such as the EFF[0]. This is not some new reaction; they have had Know Your Customer guidelines about terms for ML for states since 2011[1].
Even competitors who are incentivized to sell facial recognition tech to LEAs are warning that it's a bad idea for society, with serious and likely lethal implications. Here's the CEO of facial recognition startup Kairos[2].
> ...the confidence level would be tweaked by Law Enforcement
This is a large part of the problem. If you don't think so, then we simply have very different visions of the future that we want to live in, and must agree to disagree.
[0] https://www.eff.org/wp/law-enforcement-use-face-recognition
[1] https://www.eff.org/deeplinks/2011/10/it%E2%80%99s-time-know...
[2] https://techcrunch.com/2018/06/25/facial-recognition-softwar...
I mean, accurate, but then again they are making guns, so ... low bar i guess.
To me, this is the issue. We make all sorts of statements when it’s congress, but when it’s real life it’s not enough to make the news.
After all, how many members of Congress have been shot because a police officer thought their cell phone was a gun?
THAT is what the ACLU is trying to point out here.
However, this article states clearly that there are recommendations to LE on how to use Rekognition in a meaningful and correct manner for LE's use case. I am highly confident that Amazon will provide technical documentation and manuals for a contract of that size; of course Amazon will not be able to conduct oversight on LE operations. LE may missuse it, the way they've missused Stingrays, stoplight cameras, clipper chips, and other devices.
You state that having 99% accuracy for identifying an individual will do large scale harm, and I agree; what about the 1 person out of every 100 that is falsly identified? The other side to this coin is the number of people law enforcement would miss recognizing without this software... This argument comes down to human ability vs. machine learning ability, and human discretion in interpreting results. I think with proper use it'll be a benefit to identying culprits quickly.
Regardless, it is still pretty evil giving LE any tool that can autonomously monitor people who have not committed any crimes. But who's going to stop them?
BTW . Why is the default 80% Does it deliver 4/5 TP in the wild? Has anyone done an evaluation? What about vs. the kind of corruptions that are common in the wild - how does it handle odd light, reflections, shadows?
For those of us without domain knowledge, however, what the ACLU published is wildly misleading, to the point of being actively disingenuous. What are folks without baseline STEM education, much less AI education, supposed to take from the ACLU article other than "Amazon AI is racist and dangerously ineffective"? Do you think they are opening up the Rekognition docs to evaluate if the default config is appropriate for this test?
(By the way, I wouldn't be surprised if Rekognition is actually racially biased. I also feel a bit icky criticizing the ACLU here — they do great work, and I encourage everyone to donate.)
Hopefully, and I think this scandal will help, the engineering departments/contractors for the government will develop software that is appropriately configured for each case.
What I imagine is this: 1. The cop uploads a picture of the suspect taken from camera surveillance at crime scene 2. Software displays the mugshots (and names) of PROBABLE matches from the database with confidence %, color coding according to confidence level, and notice of proper use 3. Cop compares man/woman's picture with what's displayed 4. Investigation continues if necessary
The former ACLU, yes, but not the current ACLU: they have repeatedly said they dont want to protect Free Speech anymore.
https://reason.com/archives/2018/07/26/liberty-makes-us-unfr...
Far worse if you're being polarized. Also they ACLU hasn't released their methodology, are you jumping to conclusions with the ACLU "doesn't understand technology"?
If this is used to misidentify someone and then that someone gets shot how will the folks at Amazon feel about that? To me failing to consider that and failing to act appropriately is reckless. It's as reckless as Uber turning off the Lidar and as reckless as Tesla using the road markings as a dominant steering method.
When you have to ask "how many innocent dead will it take?" do we not all think that something might just be going wrong!
Why on Earth would they put something like that in their ToS? If anything, the place for that is in documentation ("this product may not be suitable for use in law enforcement"), but then again, since when Amazon does have a clue about the specific requirements of government services, to issue blanket statements like that?
> It's as reckless as Uber turning off the Lidar and as reckless as Tesla using the road markings as a dominant steering method.
Yes, you don't blame the Lidar company for equipping their products with an off switch, nor do you blame whoever designed the road marking detection for Tesla's decision to base their self-driving tech on it.
For law enforcement, I assume they wouldn't be hitting the Amazon API directly, but rather using software someone sold them which as part of its service uses Amazon's API.
I have stock in Amazon, but also caution about law enforcement (and a desire for technology to increase accountability of government to citizens rather than vice-versa). For me, the fear of going even one tiny step closer to China outweighs everything else.
For those who are arguing on behalf of Amazon, what's your emotional calculus here?
You, as a bystander, can be caused great inconvenience (where the line of injustice gets crossed is up for debate) through someone else's expectation of safety.
I guess the way I look at it are the odds of my child being abducted are one in a 300,000. On the other hand, a significant portion of the world (30%?) lives under oppressive regimes. So there's that.
That's such a huge margin that I have trouble understanding any other way of looking at it.
Secondly, I think allowing this technology in America makes it more tolerated in other countries. Perhaps facial recognition in public should be banned by international treaty.
The best practice is to show a ROC curve, which shows the tradeoffs of selecting a given threshold: https://en.wikipedia.org/wiki/Receiver_operating_characteris...
[0] https://en.wikipedia.org/wiki/Receiver_operating_characteris...
[1] https://en.wikipedia.org/wiki/SCR-270
[2] https://archive.org/stream/Vol1PlansAndEarlyOperations#page/...
Which it never is.
So cut it out with your justifications about ovens and other nonsense and acknowledge that you are just a corporation releasing a product that tries to make a buck and/or not get left behind some other competitor of yours.
Seems like more productive conversation could be had over how to use, than blanket dismissals. The proverbial 400-lbs man in his basement can code this up relatively easily with existing tools, to near state-of-the-art results. So the guy that works for the FBI will too. Let's come up with appropriate, targeted safety mechanisms so the oven doesn't light the whole building on fire.
As long as the results are not taken as judgements, but only auxiliary tool, it would not be a concern.
The article indeed argues that any positive match should be vetted by a human.
I have said before and I will say again: after the disaster that was the Amazon Fire Phone I have zero trust in any Amazon.com employee to pick up that any time phone to sound the alarm or if they do, I don't expect Amazon to "do the right thing". The reason I mention this is because I don't think this PhD fellow speaks for the the sales "engineers" at Amazon.com. The sales people will say what they have to say to make a sale.
Lets face it: we programmers just work at a company, the business actually runs the show. I have zero confidence that the customer will understand that "any positive match should be vetted by a human". What does vetting mean anyway?
No police has ever stopped me for matching the profile of a suspect, ever. However, I know people who say it happens (or at least used to happen) to them at least twice a year.
Is a blanket ban on law enforcement from using facial recognition possible? Or at least restrict it to federal agencies (a ban on all law enforcement that is not a federal agency)?
Or, don't use the oven on setting 3 because if you do your house will burn down?
This is my biggest gripe with how the ACLU has conducted this. I find it hard to distinguish their "test" from clickbait.
He attempts to refute one unverifiable Rekognition configuration with another.
>In addition to setting the confidence threshold far too low, the Rekognition results can be significantly skewed by using a facial database that is not appropriately representative and therefore is itself skewed. In this case, ACLU used facial database of mugshots that that may have had a material impact on the accuracy of Rekognition findings.
Although, it's important to note that machine learning algorithms generally have a more difficult time with darker skin because of lower contrast (see scandals on Google identifying black people as gorillas and cameras not auto-focusing on black people). It is a legitimate concern that black people may be misidentified more often with any machine learning algorithm.
But we have no objective proof that every one is as recognizable as another.
https://webcache.googleusercontent.com/search?q=cache:wzZHRC...
This blog shares some brief thoughts on machine learning accuracy and bias.
Let’s start with some comments about a recent ALCU blog in which they run a facial recognition trial. Using Rekognition the ACLU built a face database using 25,000 publicly available arrest photos, and then performed facial similarity searches of that database using public photos of all current members of Congress. They found 28 incorrect matches out of 535, using an 80% confidence level; this is a 5% misidentification (sometimes called ‘false positive’) rate, and a 95% accuracy rate. The ACLU has not published its data set, methodology, or results in detail, so we can only go on what they’ve publicly said. But, here are some thoughts on their claims:
1. The default confidence threshold for Rekognition is 80%, which is good for a broad set of general use cases (such as identifying objects, or celebrities on social media), but it’s not the right one for public safety use cases. The 80% confidence threshold used by the ACLU is far too low to ensure the accurate identification of individuals; we would expect to see false positives at this level of confidence. We recommend 99% for use cases where highly accurate face similarity matches are important (as indicated in our public documentation).
To illustrate the impact of confidence threshold on false positives, we ran a test where we created a face collection using dataset commonly used in academia, of over 850,000 faces. We then used public photos from US Congress (the Senate and House) to search against this collection in a similar way to the ACLU blog.
When we set the confidence threshold at 99% (as we recommend in our documentation), our misidentification rate dropped to 0% despite the fact that we are comparing against a larger corpus of faces (30x larger than ACLU’s tests). This illustrates our point that developers should pick the appropriate confidence threshold best suited for their application and their tolerance for false positives.
2. In real world public safety and law enforcement scenarios, Amazon Rekognition is almost exclusively used to help narrow the field and allow humans to expeditiously review and consider options using their judgement (and not to make fully autonomous decisions), where it can help find lost children, fight against human trafficking, or prevent crimes. Rekognition is generally only the first step in identifying an individual. In other use cases (such as social media), there isn’t the same need to double check, so confidence thresholds can be lower.
3. In addition to setting the confidence threshold far too low, the Rekognition results can be significantly skewed by using a facial database that is not appropriately representative and therefore is itself skewed. In this case, ACLU used facial database of mugshots that may have had a material impact on the accuracy of Rekognition findings.
4. The beauty of a cloud-based machine learning application like Rekognition is that it is constantly improving as we continue to improve the algorithm with more data. Our customers immediately get the benefit of those improvements. We continue to focus on our mission of making Rekognition the most accurate and powerful tool for identifying people, objects, and scenes – and that certainly includes ensuring that the results are free of any bias that impacts accuracy. We’ve been able to add a lot of value for customers and the world at large already with Rekognition in the fight against human trafficking, reuniting lost children with their families, reducing fraud for mobile payments, and improving security, and we’re excited about continuing to help our customers and society at large with Rekognition in the future.
5. There is a general misconception that people can match faces to photos better than machines. In fact, the National Institute for Standards and Technology (“NIST”) recently shared a study of facial recognition technologies that are at least two years behind the state of the art used in Rekognition and concluded that even those older technologies can outperform human facial recognition abilities.
A final word about the misinterpreted ACLU results. When there are new technological advances, we all have to be careful to be calm, thoughtful, and reasoned about what’s real and what’s not. There’s a difference between using machine learning to identify a food object and whether a face match should warrant considering any law enforcement action. The latter is serious business and requires much higher confidence levels. We continue to recommend that customers not use less than 99% confidence levels for law enforcement matches, and then to only use the matches as one input across others that make sense for each agency. But, machine learning is a very valuable tool to help law enforcement agencies, and while being concerned it’s applied correctly, we should not throw away the oven because the temperature could be set wrong and burn the pizza.
1) When the ALCU tested mugshot photos VS congress member photos, the ALCU set it to use an 80% confidence level threshold, which should result in a 5% false positive rate. There are 535 members of congress, which from this setting should have resulted in 26.75 misidentifications. The ALCU got 28 misidentifications, which is pretty darn close.
2) The ALCU use a comparatively small dataset of photos. Using a different, 30x bigger dataset, and a 99% confidence level resulted in no misidentifications.
3) The police are only supposed to be using this system to narrow down choices, and then have a human sorting out possible matches, which this should do very well.
---
My thoughts: The system appears to be functioning exactly according to specifications. The eternal problem of systems is that the specifications do not match what users are actually expecting/needing the system to do. If the ALCU is confused about how it works, the police probably will be as well.
Potentially, although big volume customers like public bodies are almost certainly in contact with eg. AWS solution architects, who would hopefully guide them around pitfalls like this.
The technical staff interfacing with AWS is probably many layers removed from the day-to-day employees that use this sort of software in law enforcement. Federal agencies like the CIA & FBI have trained analysts to do this work, but your local law enforcement agency isn't going to have this level of specialization or training. You can look at the history of any new tool of forensic science and see unfortunate abuse when it's first deployed (like DNA evidence).
Now the question is: Is amazon responsible and should it get the bad press? Or should we blame the local cop and his training? Or the police software implementing the technology?
Unfortunately a lot of innocent people will be victims of bad policy and training until then.
> the ALCU set it to use an 80% confidence level threshold, which should result in a 5% false positive rate.
I'm not sure where the "should" in this sentence is coming from. As far as I know, Amazon's documentation doesn't claim any particular relationship between confidence and false positive rate. The 5% number in the article was simply talking about the actual false positive rate that the ACLU observed.
Not to say the conclusions of the article are wrong, though. It certainly seems unreasonable for someone to assume that "80% confidence" would give an FPR of less than 5%.
The point is that all of these decisions are unregulated, and complicated systems are being put into the hands of people who have very different professional backgrounds (and incentives) to Applied Machine Learning Engineers.
- Just because it’s possible to use a better or more representative data set, it doesn’t mean LE agencies can or will
- Just because a particular confidence interval is more suitable, doesn’t mean it will be used. What if there’s pressure to hit targets? What if a senior cop sits on the sys admin’s desk and tells the them that “We’ve really gotta nail this guy. He’s armed, dangerous, and on the run”.
- What if cops are radioed and told that a man who HQ are 95% sure is the armed killer is in the car/room/office ahead? How will the LEO react and behave when approaching the suspect, based on that information? What if 10k faces from CCTV were put through the system to get that match? What if there are 10 other officers drawing their pistols approaching 95% matches for the same subject at the same time?
These are all difficult questions for statisticians and ethicists to answer. We should not be forcing these decisions on LEOs to deal with in real time.
And I, personally, would prefer to live in a free country without facial recognition cameras pointing at me.
There is no democratic consent for any of this.
But should you regulate the tech itself? Like ban the development of facial recognition all together? Or for identifying human faces?
I feel if you regulate, it should be in its proper usage. Which would mean the law enforcement is to blame. Ot maybe you'd need to get an explicit license to be allowed to use facial recognition and Amazon could make sure that only licensed individuals get access to their tech flr example.
Point being, nothing here indicates that Reckognition has inherent bias in its design.
I'd rephrase that Amazon recommends (Rekomends?) using a CI of 99% for law enforcement, and that it only be one factor in policing.
---
We're getting to a brave new world here, but there is a lot of parallels with the current one. A common trick for reducing speeding tickets is to request things like when the last time the radar was calibrated, and the date of certification of it's operator.
Soon, your lawyer can challenge ML based convictions by requesting the coefficients of the model.
Also, the police will use every trick in the book to get an arrest, so unless it's clearly illegal (and in some cases that's not enough to stop the police), you can't say the "police should..." do anything.
My understanding about Amazon is that they don't have a culture of shooting from the hip. They certainly encourage risk taking, but not without putting together a solid plan with review.
Sadly, these days I see too many ML practitioners and "data scientists" without the necessary prob/stats foundation. Misinterpreting data happens even to experts, so not surprisingly it's going to happen even more so to amateurs. Suggesting more foundational knowledge is considered elitist. Why shouldn't a month-long bootcamp be as good as a MS/PhD? In this case, the original ACLU results fits a certain narrative and hits all the hot topics so it was bound to be picked up.
Facial recognition is likely to help rapidly match known photos of a suspect with security footage and imagery gathered near by to help pin down movements of a suspect prior to and after a crime, in order to focus the search for further evidence.
This is a tool to catch stupid criminals; it seems easily fooled by those who approach crime in a more systematic manner.
There should be laws that prevent using facial recognition to track the locations of people who aren’t the subject of investigation. I am sure that technology would be very desirable to advertisers, and I don’t want them to have it.
I’ve said this many times. What it means is that even if you think it’s great to outsource your machine learning model to a third party like Amazon, and consume it like a service, you really need in-house expertise in machine learning anyway, to help you interpret the accuracy, diagnostics, and some “integration test” notion of model performance on a sample that is agreed upon as truly representative of the stakeholder runtime conditions.
So either way, you still have to pony up the dough to employ adequate in-house expertise in modeling and domain machine learning.
As a result, there actually are not that many use cases when it would make sense to not build your own tailored model in-house. If you already must hire most of the staff with necessary skills to evaluate a third-party, the incremental effort to train a good-enough fine-tuning of some imagenet-based model is really not that high. And you can customize the training and diagnostics.
This is especially critical when there are also stakeholder-specific latency or throughput constraints.
For face detection in particular, I know a great deal about this directly related to AWS Rekognition, because my team extensively evaluated it for the possibility of outsourcing face detection calls in an in-browser image editing application.
Not only could Rekognition not meet our runtime requirements, but separately it had poor accuracy and coverage (often capping out at detecting a small number of faces per image) and the Rekognition-reported confidence score aligned horribly with our own ground truth bounding box data that allowed us to quantify overlap of the detection with intersection-over-union scores.
We tested a handful of fine-tuned prebuilt models in Keras: Resnet, Inception, and our own port of MTCNN, and had drastically better face detection (on a very large training corpus) than Rekognition in about a week.
And our cost per request, even deploying to AWS, was much lower than Rekognition even at our high volume of traffic and even inclusive of the cost to spin up GPU instances for training.
What’s more is that we tightly controlled how to write the web service layer wrapping it, and how to optimize the image preprocessing steps, etc., that we have no visibility into with a third party.
I think there is a conceptual gap with this stuff where people just think that because you could commoditize a web service around a prediction algorithm, it must mean it would be economically valuable.
But that part is the absolute least meaningful part. The whole enchilada is diagnosing how the model performs on the exact stakeholder use case, inclusive of performance constraints.
For this reason, I think the only market for “ML tools for people who don’t know ML” is going to be just like vapid corporate IT consulting.
Don’t get me wrong: big tech companies will profit from this. But not because it solves any actual prediction problems or saves anyone money compared with in-house ML development. It’ll just be the standard Dilbert-y politics that has always yielded profits to IT consulting services.