A picture with a white college kid in the frame would get the output of "human". Put an African borrower in the frame and you got at best a failure to recognize a human, and at worst a reference to an animal.
I would hope the situation is much better now, but the bias (and just sheer inaccuracy) of the tools was readily apparent and we gave up on image recognition for the time being.
Computerized facial recognition has been progressing for something like 50 years, with each paper building on previous work. I suspect a parallel world where the dominant social group hard dark skin would explore different technical avenues, focus on ways of recognizing faces that lean less on tone, etc. Obviously it's hard to say with what we know now.
Humans have difficulty recognizing cross-ethnic faces, however I don't think dark skinned people are worse at recognizing same-ethnic faces than light skinned people. To me that suggests our social history is more to blame than some fundimental quality of skin tone.
This can extend beyond the film or image sensor to biasing lighting equipment, defaults, standards, and thresholds.
Looking through the literature you'll see things like, "the design decisions in the NTSC’s test images effectively biased the format toward rendering white people as more lifelike than other races"
https://sterneworks.org/Mulvin-Sterne-Scenes.pdf
I'm mostly familiar with the history in the US. I'm earnestly curious about how other countries accommodated these types of things as they adopted the technologies.
1. https://petapixel.com/2015/09/19/heres-a-look-at-how-color-f...
https://www.npr.org/2014/11/13/363517842/for-decades-kodak-s...
but, yes, dark skin reflects less light than light skin. by definition.
The camera settings and background lighting make a big difference.
Some notes:
- the eye can handle the dynamic range, but there's no correct exposure for that photographic combination, so they play with lighting, exposure bracketing and dodging when printing
- it's recommended that brides wear off-white (like cream) or other colors of fabric if they're concerned about the photos
Mind you, the hobbyist uncle may produce great photos an may be a wizard with lightroom and photo shop. But a lot of hobbyists try to compensate with gear for skill, or more charitable may not be even aware of related skills an possibilities. So in the end the snapshots your friend Andy took with his 2 year old iPhone might look better (have better automatic processing) than those photos Herbert took with his expensive prosumer or even pro kit.
Seeing how many great photos come out of the event in sometimes very strange lighting conditions (as you might expect from an event that includes everything from brightly lit outdoors events to indoor conference spaces and theaters to rock show-lit concerts, and everything in between), I'd imagine that little metering trick is doing wonders.
https://www.quora.com/Why-are-all-bodybuilders-tan-Do-they-u...
Moreover capturing Black/Dark skin and features requires more accurate light metering & lighting because dark skin absorb more light. There's a lot variance in cheekbones, nose and lips.
Humans' features in general, are more complicated then you realize.
I'm curious how many white people there are on earth vs how many black people there are, and other races. A couple google searches didn't give me any easy finds
The US is the only country I am aware of who still uses this term, everywhere else was using some thinking like ethnicity to indicate different culture, origin...
https://www.tandfonline.com/doi/abs/10.1080/0031322050034783...
This meme dates back to a loose claim made by R. Lewontin back in the 70s. In fact, you can very precisely and reliably recreate the "intuitive" human racial categorization using unsupervised algorithms, like doing multi-dimensional clustering over fixation indices. (It does not work using single-dimensional clustering, which is what Lewontin was talking about.)
Modern biologists usually talk in terms of clines rather than races, but this is just using the first derivative instead of the zeroth - you'll get the same result either way.
> Africa actually has more human genetic variation than anywhere else.
SNP diversity has ~nothing to do with phenotypic variance.
Not really - almost everyone can agree on "middle eastern/north african", "east asian", "south asian", "black african", "white", etc. If you force people to pick a single-digit number of major categories, they're probably going to come up with the same categories that k-means in fixation space would.
That almost everyone can agree on these categories is also contrary to reality. For example, many of who you describe as East Asians consider themselves racially distinct both within their societies and from their nearby neighbors. Also, what major categories do mixed race people fall inside?
> many of who you describe as East Asians consider themselves racially distinct
That's why I specifically mentioned the number of racial categories involved. Obviously as the number increases you can have different clustering results.
> Also, what major categories do mixed race people fall inside?
Obviously not into any of them, if we're talking about a simple mechanical classifier with high separation.
It may not be “your job” to provide such evidence, but you’ve made a series of specific claims about things like the rate of phenotypic variance among different racial groups. If you don’t want to defend them, that’s your prerogative, but you also can’t expect them to be received as authoritative or remain free of challenge.
Even the man who coined "Caucasian" as a racial category recognized that there was more physical variance among African populations and individuals than compared with Europeans.
> some who have naturally blonde hair and some who have blue eyes
I know they exist, but we are talking about statistical properties of entire populations, and these people are very rare.
> Humans' features in general, are more complicated then you realize.
It's not about what I realize - it's about what can be mechanically detected.
You missed this part: "Moreover capturing Black/Dark skin and features require more accurate light metering & lighting because dark skin absorbs more light."
I have to see the images used to train the ML model, to be certain, but based on my experience working in photography and programming, I believe it is more likely than not that they used essentially poor quality images for the training.
Moreover, after the model has been trained, to use the system effectively the facial recognition camera has to be set up to capture both light and dark skin, in the case of dark skin, it typically means not relying on available light alone indoors, an additional camera light must be provided.
The reality is if you want a facial recognition system that accurately detects dark skin it will cost a bit more to do it right.
But if it were so, then it sort of invalidates your argument because the bias complaints are that the facial recognition algorithms often misidentifies other races (not generally as white but as perhaps not human)
Specifically the scandal some years ago when black people where being identified as gorillas, it seems obvious to me that if non-white people have less facial variance it should be easier to identify black people as black instead of more difficult.
I thought it was a basic understanding that unfamiliar people, events, places, etc _do_ look alike, because before you get sufficient experience and exposure, you don't have enough skill to know what features are important to focus on and which ones carry no information.
It's also worth considering what saying that implies to the people you're saying it to. Especially when said dismissively. I'm terrible at names, faces, voices, pretty much any way of telling people apart. But that's my problem, and I need to be careful not to push the burden of that problem onto the people around me.
The first one CAN also be troubling if you have had ample exposure to become familiar, and are indirectly admitting that you did not believe it to be an important skill to learn.
I'd say it's an ignorant thing to say in the way you say it. Like you said in this post, you just don't know the features to look out for when you're not familiar with them
You might be right in terms of hair or eye color (not too many gingers outside Europe), but Africa has a vast amount of genetic diversity, so I'd naively guess that variance along other feature dimensions would be correspondingly higher too.
This is one bias I think benefits folks with dark skin. I'd love to be in the group that doesn't get recognised by this terrifying tech.
And the truth is, the danger zone is when the algo is somewhat ok, but a few percents worse on the discriminated group. If it were obviously stupidly bad (like for example any two black people are a positive match according to it) then people will disregard it. But if its just slightly bad thats harder to notice and suddenly your loan/car rental is rejected because the security system says you look too much like a specific known fraudster. And of course they wont even tell you that thats the reason, because security.
The indications in question were extremely rare for black people on the other hand. Another use case was the detection of Vitiligo, which is significantly more easy on black people. There was a great demand for therapies because patients are stigmatized for different reasons, especially in Africa because the symptoms can look similar to some deadly diseases.
There is a lot of discrimination in imaging, but it doesn't have to do with race and more with the creation of binary images.
As for facial recognition. I see some use cases for convenience, but it would be far more effective to disallow widespread application because of the numerous unsolved problems. So I see the move from IBM positively, even if they continue their work. In the end I believe face reg to be possible without bias, but I don't have these propped up security needs.
If your a member of Congress who is also a person of color, it will do that at twice the rate of your white colleagues. [0]
[0]: https://www.aclu.org/blog/privacy-technology/surveillance-te...
[0] https://www.nist.gov/programs-projects/face-recognition-vend... [1] https://www.nist.gov/news-events/news/2019/12/nist-study-eva...
At best, you have AI's that can easily recognize pathologies that an average rad could recognize, but are useless when it comes to trying to recognize a pathology that 99.999% of rads would miss. Why? Probably because the system is biased towards pathologies that rads recognize. Why? Probably because that's what's in the learning set. I understand all that, but that makes the AI almost useless in a production setting simply because of the way many healthcare delivery networks are structured with respect to radiology.
At worst, you have AI's that seem to recognize pathologies that an average rad could recognize, but then inexplicably have a horrible miss on an obvious study that any first year radiologist would have caught while half asleep. It's the sort of miss that makes the astute observer wonder if the other vendor's AI, the one that didn't have any misses on your test set, was simply not fed an example study that triggered its blindspot? You start to wonder where the blindspot on that one was? You start to wonder does it even have a blindspot? How can you work around blindspots like this in general? Etc.
But here's the thing, it's never good when you're thinking about mitigations while you're still testing the AI. My first RSNA was over 20 years ago. To this day, we're still hearing the same promises, and the production testing, (it never fails), is still uncovering the same issues.
Now I would have thought that recognizing a human would be easier than trying to ascertain, say, calcification in a DX, but apparently these same issues crop up in a number of different applications of ML based technologies.
I understand that a training set, by the very definition of the word "set", does not contain everything. Obviously, there will be bias towards whatever is in the training set. Essentially, most techniques today, simply train AI's to tell us whatever we tell the AI's to tell us. But for this reason it should be unsurprising that these sorts of AI's work best in settings with well defined domain spaces and are challenged when the domain space is less well defined. (And in either case, these AI's will have an inescapable bias towards whatever was in the training set.)
What does a tumor look like? Is probably too broad a question to give these AI's. What does a stop sign look like? Is probably a question that these AI's could answer relatively reliably. (I hope?) You would have thought that "what does a human look like?" would be closer to "what does a stop sign look like?" But I guess it's not surprising to hear that it seems to trend towards "what does a tumor look like?" in practice.
It is true that ML algorithms are almost always trained on radiologist labels on the same modality, and thus take in the reader biases. I also agree that some radiologists are better than others as you imply.
As a patient, one does not know who will read their film. IMHO we as an industry should aim not at beating 99.999% of radiologists. We should merely make products which consistently perform not worse than an average radiologist at a particular institution. It is always thrilling to outperform humans with your software, but at the end patient outcomes are what matters. Those are about consistent performance over a long period of time.
Demonstrating this consistent performance is the challenging part, but it is possible to prove it through sufficiently careful and lengthy prospective trials. That’s what we are focusing on, and I would love to see the other players in the industry do the same.
I believe that ML should not be taken in lieu of human opinion. The consensus, be it medical or legal, has to be explicitly human with all the responsibility attached.
Shifting the responsibility for the misses onto a faceless ML is only eroding trust in the professional opinions and cementing the biases.
The problem with this argument is it glosses over that fact that in the tail, where the ML is making a wrong decision, sometimes a catastrophic one, the behaviour of the ML algorithm is not well understood. How can we deploy something in such safety critical applications that we do not fully understand?
Please also note that there are several important differences as compared to the automotive industry. First, one could argue that the task at hand is trivial as compared to the self driving car. We are operating in a heavily constrained setting with much better understood data inputs and a hundred-year history of medical professionals trying to classify and systematize them. Moreover, our task is not time-critical. It sometimes takes more than a week for such an image to be reported on.
That said, a non-diverse facial dataset in a diverse society like the US is simply useless. It doesn't help saying the AI is suffering from a human bias, and dropping these projects entirely, unless they are being used for a malign purpose like what China does.
That happens uncomfortably often - and not even anywhere where AI is even remotely involved, like soap dispensers: https://reporter.rit.edu/tech/bigotry-encoded-racial-bias-te...
I would hope that on HN one would take the time to explore the technical issues before even suggesting bias (which implies human racial discrimination of some sort).
I am going to venture a guess that there's a large audience in the image processing/AI/ML world that lacks a fundamental understanding of image sensor and lens technology. I have never seen mention of concepts such as well capacity, quantum efficiency, noise floor, dynamic range, thermal noise, gamma encoding, compression induced errors, etc. in most work I have reviewed.
Sensors used in the general class of imagers found in these experiments are nowhere near adequate to capture the full dynamics of a lot of real life images. The lowlights (referring to the lower portion of the dynamic range of a camera, encoding, compression and image processing system) can be some of the most challenging portions of the dynamic range to get quality data.
The old idea applies: Garbage-in, garbage-out.
It should come as no surprise that algorithms trained on (likely) bad images with bad lowlight detail will fail to deal with people of darker skin. It's almost a given. One can't assume cheap cameras and the data sets produced with these cameras will see the world the way our eyes are able to. Not to mention the fact that we have something called "understanding" while classifier systems have no clue whatsoever what they are looking at, all they can do is put things in buckets and that's that. In other words, there is no inherent comprehension of what a human being might be versus a bear or a teapot. That's a major problem.
The answer isn't to give up. The answer is to understand and then go back and do it right. This isn't going to be cheap and it will likely require rethinking how we build and train these systems.
As a tangentially related data point, I have three German Shepherd dogs. Two are the traditional black and brown coloring. The third is 100% black. It is virtually impossible to take a good picture of him. In anything but the right lighting he shows up as a dark amorphous blob. For all the prowess of the mighty camera in an iPhone 10, you'd be hard pressed to use those images to recognize him as anything other than a blob on a dark couch.
I do have access to high performance images with far greater well capacity as well as 100% uncompressed data output. In that case there's usable data in the lowlights that, through gamma and LUT manipulation can be extracted. When you do that he quickly goes from looking like a blob to looking like a happy dog.
Anyone interested in learning more, I would highly recommend looking up Jim Janesick:
https://www.google.com/search?q=jim+janesick
The popular phrase "he wrote the book" applies here. Jim's books on the subject of image sensor technology (science and design of sensors) are the reference work anyone in imaging studies. He designed so many sensors for space applications I am not sure he even remembers how many. I was fortunate enough to study CCD and CMOS sensor design under him a couple of decades ago.
ML has to start with good data. Inadequate sensors coupled with compression and other processing artifacts leads to bad data, a formula for failure.
But that's still a cop out. You can't use models with these problems. It gets you disasters like face unlock that doesn't work for black people. This happens even if people aren't white. Samsung famously released blink detection that didn't work on a lot of East Asian faces. The products are broken because the white bias is the default and pervades everyone's thinking.
Vision researchers should take a class or two in cinematography and photography, it would serve them well. Even the quality of the lens makes a difference. Most work I've seen out there uses cameras that barely pass as security cameras or webcams.
That said, yes, ML needs to work with crappy images and every single camera out there. My argument is that you are not going to be able to train using crap data. And the images in a data set would be crap if the data --the images-- were not acquired using cameras and techniques that provide enough data across various segments of the dynamic range.
Again, I gave the example of my black GSD for a reason. You are not going to be able to recognize him as anything other than a blob on a couch without a camera that can capture enough data at the low end of the dynamic range and a system trained with that data.
The fact that Samsung (or anyone else) failed means nothing. In order for that data point to be meaningful you'd have to have intimate knowledge of what they were doing and what capabilities they had, both in terms of science and engineering as well as the consumer hardware they developed.
I have competed against multi-billion dollar multinational corporations who, despite their financial prowess and scale, could not design their way out of a paper bag. They don't understand the problem, lack creativity and, most importantly, absolutely lack the passion necessary to solve it. Ten 9-to-5 engineers can't compete with a single engineer passionate enough to devote every waking hour to solving difficult problems. It doesn't matter how much money you throw at them, they just can't perform.
This doesn’t help when the variance in feature space of white people’s faces is objectively a lot wider than the variance of black people’s faces. You can have the best camera in the world, but it’s not going to change the fact that someone from Ireland looks way more different versus someone from Italy (in any mechanical basis) than someone from Cameroon looks versus someone from Malawi.
To add to this considering that genetic variations between peoples on the African continent are larger than the variations to anywhere else, this is also very unlikely. I think it's a reasonable assumption that genetic variation and variation in physical features are somewhat correlated.
You are misinterpreting this, as does almost everyone else. Phenotypic variance is almost completely orthogonal to number of unique SNPs (which is what this generally refers to).
> genetic variation and variation in physical features are somewhat correlated.
This is not the case. You can have narrow population bottlenecks (reducing SNP diversity) followed by high variance in selective pressure. This is exactly what happened to early Europeans. You had a small initial population spread out to inhabit a large variety of ecosystems.
The second issue is believing some technical arcana are a valid excuse for selling products with such errors. It isn't! If your product reinforces entrenched discrimination, you either fix it or stop selling it.
In my experience, they rarely will even consider it.
Therefore, the training data should have the same technical shortcomings as what will be used in production.
I think training and the application of the trained system are separable. Nothing is accomplished by training with data sets that lack data or detail. Inference or classification is impossible or deeply impaired by the lack of data.
As a hypothesis, the solution is likely to involve training one network with good data and a separate network to be the interface between the first network and the imperfect perceptual data in real-world applications.
At the end of the day AI/ML need to leave the world of classification behind and move on to the concept of understanding. This is not an easy task, yet it is necessary in order for these amazing technologies to truly become generally useful.
We don't show a human child ten thousand images of a thousand different cups in ten different orientations in order for that child to recognize and understand cups. The reason for this is that our brains evolved to give us the ability to understand, not simply classify. This means we need far fewer samples in order to have effective interactions with our physical world.
The focus on using massive data sets to train classification engines is a neat parlor trick, yet it will never result in understanding and is unlikely to develop into useful general artificial intelligence. The problem quickly becomes exponential and low quality data becomes noise. We need to develop paradigms for encoding understanding rather than massive classification networks that can't even match the performance of a dog in some applications. As I said before, this is a very difficult problem. I don't think we know how to do this yet. Not even sure we have any idea how to do it. I certainly don't.
And for those who don't know the reference: https://www.imdb.com/title/tt1346402/
This was exactly my translation of the statement.
When analyzing large business decisions, I find the most success when viewing through cynicism. Self-interest is after all the primary driver of most businesses.
If every person who breaks into a car or sells something in the street is caught, this would be an insane surge in the prison population.
Isn't this a good thing? Breaking into peoples cars is wrong, we want to prevent people from doing this.
Would you rather lose your car window or would you rather sacrifice someone's life?
You're exaggerating the argument. Stop trying to justify social degeneration as acceptable.
A few months back, covid-19 was the perfect scapegoat for a lot of companies to drop projects (or go bankrupt, cough OneWeb).
I'm generally of the opinion that tools are intrinsically amoral; neither moral nor immoral. A hammer could be used to build an orphanage, or be used to hit orphans; it's an amoral tool that could be used for good or evil. But I think when tools become sufficiently powerful, that sort of perspective starts to break down. Weapons of mass destruction are good examples of this from the past century. Could there conceivably exist digital systems too powerful to consider morally neutral? Personally I think so, and I think facial recognition technology might qualify.
This entire “tools are amoral” argument seems to deny the fact that in most cases we don’t need to judge the intrinsic morality of a particular tool, because we can just look at how the tool is actually being used.
Personally, in cases where there isn’t a specific explanation for why a particular decision has secretive motivations, I do find it a touch too cynical.
Not saying that this is what happened at IBM though.
So far nobody has dug up all the little nuggets I've left in public records. Sucks