> but the bias (and just sheer inaccuracy)
I would hope that on HN one would take the time to explore the technical issues before even suggesting bias (which implies human racial discrimination of some sort).
I am going to venture a guess that there's a large audience in the image processing/AI/ML world that lacks a fundamental understanding of image sensor and lens technology. I have never seen mention of concepts such as well capacity, quantum efficiency, noise floor, dynamic range, thermal noise, gamma encoding, compression induced errors, etc. in most work I have reviewed.
Sensors used in the general class of imagers found in these experiments are nowhere near adequate to capture the full dynamics of a lot of real life images. The lowlights (referring to the lower portion of the dynamic range of a camera, encoding, compression and image processing system) can be some of the most challenging portions of the dynamic range to get quality data.
The old idea applies: Garbage-in, garbage-out.
It should come as no surprise that algorithms trained on (likely) bad images with bad lowlight detail will fail to deal with people of darker skin. It's almost a given. One can't assume cheap cameras and the data sets produced with these cameras will see the world the way our eyes are able to. Not to mention the fact that we have something called "understanding" while classifier systems have no clue whatsoever what they are looking at, all they can do is put things in buckets and that's that. In other words, there is no inherent comprehension of what a human being might be versus a bear or a teapot. That's a major problem.
The answer isn't to give up. The answer is to understand and then go back and do it right. This isn't going to be cheap and it will likely require rethinking how we build and train these systems.
As a tangentially related data point, I have three German Shepherd dogs. Two are the traditional black and brown coloring. The third is 100% black. It is virtually impossible to take a good picture of him. In anything but the right lighting he shows up as a dark amorphous blob. For all the prowess of the mighty camera in an iPhone 10, you'd be hard pressed to use those images to recognize him as anything other than a blob on a dark couch.
I do have access to high performance images with far greater well capacity as well as 100% uncompressed data output. In that case there's usable data in the lowlights that, through gamma and LUT manipulation can be extracted. When you do that he quickly goes from looking like a blob to looking like a happy dog.
Anyone interested in learning more, I would highly recommend looking up Jim Janesick:
https://www.google.com/search?q=jim+janesick
The popular phrase "he wrote the book" applies here. Jim's books on the subject of image sensor technology (science and design of sensors) are the reference work anyone in imaging studies. He designed so many sensors for space applications I am not sure he even remembers how many. I was fortunate enough to study CCD and CMOS sensor design under him a couple of decades ago.
ML has to start with good data. Inadequate sensors coupled with compression and other processing artifacts leads to bad data, a formula for failure.