Andrew Gelman on Pedro Domingos' claim that algorithms are incapable of bias
statmodeling.stat.columbia.edu
statmodeling.stat.columbia.edu
Consider the gender imbalance in tech jobs. Let's say that is because more men pursue tech degrees than women and this is because the field is already male dominated and women don't have female role models.
A hiring algorithm will optimise for the best candidate for the job. If we assume men are as good as women, the gender ratio will be skewed simply because the underlying population has a skew. However, we, as a society, may want to overweight women candidates even at a cost of choosing a slightly less fitting candidate because, in the long term, this will provide female role models and convince more women to join tech (there are reasons why this may be desirable, but it does have to be argued for).
Under this view, the algorithm is biased because it's optimising for the wrong thing - it is assigning no extra value to choosing a female candidate even if we, as a society, think there is value.
This is a thorny issue, though, as I've made some assumptions above and plenty of people will disagree with them, on various grounds. Still, I think this is a useful way if thinking about biases.
Also, viewed like a problem of figuring out the correct function to optimise, this suggests fixing this problem via slow and heavy handed law making (e.g. a blanket rule enforcing hiring gender ratios) is probably the wrong way to do this.
By the way, hiring committees can also be thought of as algorithms and at least part of their biases can be explained when viewed as them optimising for the wrong thing.
"And here’s the half that I think Domingos gets wrong: He’s too sanguine about existing algorithms being unbiased. I don’t know why he’s so confident that existing algorithms for credit-card scoring, parole consultation, shopping and media recommendations, etc., are unbiased and not capable of outside improvement. I respect his concern about political involvement in these processes—but the existing algorithms are human products and are already the result of political processes. Again, his concern that “progressives will blithely assign prejudices even to algorithms that transparently can’t have any,” is missing the point that the structure and inputs of these algorithms are the result of existing human choices."
It would be very hard to convince me that this was not motivated by cheap point-scoring for political reasons, perhaps as defense in anticipation of an attack upon himself.
Domingos’s research is specific to naive algorithms learning from data in experimental realms. Gelman’s work often relates to expert-designed simulations which are actively used in economic policy, banking, insurance, and legal systems. From this perspective, it’s somewhat clear that Domingos’s approach has far less capacity for bias, and thus invites criticism of the older method. Learning from data versus designed simulation. This reads somewhat like a reply suggesting we should trust hand-engineered bias from the right hands rather than seek to eliminate it.
But in this case, the hand’s argument is an ad-hominem, so we see that it’s a religious argument about social values rather than a theoretical argument about statistical bias. Maybe we should have that conversation. Yes, we can use algorithms that have no statistical bias. Yes, we can insert any bias we want into an algorithm. Now that that’s established, what should we do?
Don’t get me wrong - You have a duty to get the best models possible, and you should try variables that generalize at higher levels - income instead of race) for example and always be testing.
I heard a story about a particular parole board trying to use an algorithm to predict recidivism rates. It often predicted one race more likely to do so. There is some argument that each case should be looked at individually. I agree. But should we avoid addressing the issue if the algorithm is explainable and a good predictor? Instead we should use it to help lower reoffender rates NOT by denying parole but perhaps assigning additional resources to those more likely to offend once we have already granted parole.
Having inmates trying to guess at what some algorithm values so they can game it? Kafkaesque
A bit more leeway can be given to algorithms that companies use but justice system is very sensitive.
if the algorithm is public, it can and will be gamed. The solution is to make the algorithm be based on the things that you want people to game, but that becomes not much of an algorithm at some point...
In the business world, the answer is reached by answering the question "How would our reputation change if the public had access to this information?" The end result is that AI ethicists are hired for optics and subsequently have no real capacity to publish this information.
Is this similar to how we interact with food? Is the 1990 Nutrition Labeling and Education Act a good or bad analogy to use? How similar is this to a food company hiring a food inspector and then firing them when they discover that peanut might have contaminated some other products? I guess if people don't know they won't complain too much.
I believe in the wisdom of the crowds. I like democracy, direct democracy even more so.
We should stop pretending we know better than what our data is telling us, even if it hurts our feelings. Not everything true is going to make us feel good, more often the inverse is true.