Not really surprising. The first thing they teach in data science is that bias is everywhere. One of the first things taught in programming is garbage in garbage out and that computers do exactly what we tell them. Once you start making decisions with biased data you will start to prejudice some group.
The quest for non-biases systems is a little like a perpetual motion machine. If we all have biases and these machines learn from the same data we do, using systems we write, how could one expect a different outcome?
The
To respond to some sibling comments: Yup, this is prejudice. I'll try to analogize the thereom with an example: Without prejudice, you can't recognize a leaf in a figure, because alternate hypotheses (there are an arbitrary number of things in this universe that look like leaves but in fact are not) are equally likely.
My advisor one told me that machine learning is the study of biases.
"Without the aid of prejudice and custom, I should not be able to find my way across the room." - William Hazlitt
Seems like a biased premise.
Which is... most of the thing people know? Including, ironically, this very definition, which I learned about from a HN comment that quoted a Google search result...
This is a misrepresentation of the parent comment and the article.
So a face corpus with only white faces doesn't reflect the diversity of faces one encounters in the world.
With that said, unbiasing data is extremely difficult because the true distribution of things is unknown and sometimes subjective. The visual images you would encounter as a human from birth to death growing up in a first world country would be very different from that of a drone's video camera. Are we really sure that imagenet should be K% animals and not K/2% animals? And if you train a machine learning algorithm on every possible image with every possible pixel, it will just learn noise.