For example, imagine that you wanted to train an algorithm to distinguish photos of dogs from photos humans. So you collect a bunch of photos of both dogs and humans and use them to train a classifier. You do all the proper cross-validation, bootstrapping, etc. to ensure that you are not overfitting, and you get really good results. Then, looking at the mis-classifications, you notice something: all the photos that are taken looking at an angle down toward the ground are classified as dog photos, and all the photos taken looking straight ahead are classified as human photos. It turns out that in your training set, most of the dog photos are taken at a downward angle while must of the human photos are taken facing straight ahead, because humans are taller than dogs, and your machine learning algorithm identified this feature as the most reliable way to distinguish the two groups of photos in your training set.
In this hypothetical example, no overfitting occurred. The difference in photo angles is a real difference in the training sets that you provided to the algorithm, and the algorithm did its job and correctly identified this difference between the two groups of photos as a reliable predictor. The problem is that your training set has a variable (photo angle) that is highly correlated with what you want to classify (species). This is considered an unwanted bias (and not a reliable indicator) because the correlation is caused by the means of data collection (most photos are taken from human head height) and has nothing to do with the subject of the photos.
(Though maybe the term as used in industry is less strict.)
When AI makes a decision, right now, people only uses the probability output. Hiring A has .6 probability while hiring B has .4. then we will hire A instead of B. However, if we consider the confidence intervals, the decision might not be that clear. Say +/- .5 to hire A but .2 to hire B. If exploration is considered too, very likely that we will give B a chance.
AI is in the realm of probabilistic decision making, while normal people don't follow. The bias is not from the training side. It's the decision making process incorporating AI should change.
It would be like if your car was driving in circles and you called a mechanic to fix your steering, and they told you that the actual problem was that both right wheels were missing. That's not a steering problem, and no repair to the steering system will fix it. The only fix is to put new wheels on.
...which means that whether a model is "biased" depends on where and how it's applied. This is an important point that is missing from most discussions, articles and even research papers on the so-called "AI ethics".
If the biases are consistent with other real-world data, then it's not overfitting.
If the results are odious to us, it should be impetus to critically analyze not only the AI/ML systems, but also the underlying assumptions that they're built on. Instead, developers become defensive and cage-y about their processes.
If you don't want systems to have disparate impact, you have to be adamant about it in your design. If you think society is better off with systems that reflect preexisting biases, then fine, but be ready for the backlash.
At the end of the day, it really is up to what humans want to do with themselves. It's an opportunity to be truer to our intent, not a bug to be covered up.
When you say unconscious bias, you are kind of implying that the model learns something that is false. But more often it's the case that the model learns something true that we don't want it to learn. That's what makes the problem so hard, you are trying to hide the truth from a system you only half-understand processing data you only half-understand. There is a big risk the truth slips through the cracks if you aren't careful.
It's more that the model learns something that is undesirable. It could be the case, for example, that the true thing that the AI learns is that your resume screening process tends to exclude women. This is true, sure, but it could lead to the undesirable outcome where the presence of a female name on a resume might be weighted heavily against the candidate.
I agree that this process is one of resolving blind spots, but I disagree that the blind spots are simply areas devoid of light. AI/ML systems are frequently employed to augment or stand in for human perception, which is known to be necessarily incomplete with respect to reality. In other words, they can learn things that seem true to us but that are false from another perspective, or undesirable once exposed. What's exciting about them is that they provide an opportunity to interrogate the flaws in our individual perception with a systematized observation and analysis, in a much more sophisticated manner than in the past. But fulfilling that potential requires humility.
https://www.wired.com/story/best-algorithms-struggle-recogni...
Their failure was not just in lacking diverse training sets, but diverse QA, or at least QA looking for those blindspots which eventually became evident.
So, correct, it's not as simple as having "sufficient" data.
Your expectation was the same one they had, and it was wrong, which is the crux of the issue.
Do you have a source for data set mis-labelings being a problem?
If you asked the developers of the facial recognition library, "does your software have problems with very low contrast conditions" they'd surely have answered yes. Fully conscious of the issue but, that's software. It's hard to get everything right 100% of the time.
However, ML is often sold as a solution for generating outcomes, not for finding truths, wether they be true or false.
The distinction is huge.
No matter what humans do, they will reap what they sow. Consequences and outcome matter more than "truth" (which may be in the eye and competence of the beholder).