How we use big data can reinforce our worst biases, or help fix them
nautil.us
nautil.us
A potent analogy.
The cycle is truly terrifying:
1. deep institutional racism compromises training data
2. data classified through that lens
3. ML insights and patterns inform institutional policy
4. goto 1
https://99percentinvisible.org/episode/the-age-of-the-algori...
http://www.econtalk.org/archives/2016/10/cathy_oneil_on_1.ht...
Are Algorithms Building the New Infrastructure of Racism? Possibly. But what I find even more worrying: One day algorithms may take so many factors into account that they can discriminate against minorities so small that they neither have a name nor a voice. If an algorithm like that affects you your life could turn quite Kafkaesque.
If you want to use algorithms in these areas you will eventually have to prove that they don't illegally discriminate.
So this narrows to these basic outcomes if we assume any given algorithm isn't purpose built to always have a discriminatory output:
1. The input data is junk data, so the output is junk data.
2. The input data is fine and humans don't like the output for human reasons.
The problem I'm hitting on here is that if we assume any sort of "discrimination" based on a data set is "bad", then we don't really care what the outcome of a given algo is, and we're just trying to massage data sets into an "acceptable" output. Seems like the opposite of good science to just throw away results that we don't like, because reasons.
It's quite possible for an algorithm that doesn't know what race is to be biased despite ignoring race if the real world is full of racism, and that racism manifests in dependent variables that the algorithm bases its decision on.
This concept isn't even a little novel. Google the phrase "previous servitude" for an example of just such a dependent variable in action.
My point is that when you get a good input dataset and there is still some sort of bias in the output, you have to start asking questions on why the results are biased towards one group by examining the variables pertaining to that group.
The position you're starting with seems to be that the output should never be biased, which means you're not really approaching this scientifically and won't stop messing with input sets until you get the output you agree with; not the output which is necessarily true but you disagree with for personal reasons.
You can't just claim "bad data" every time you get a result that doesn't jive with your politics.
I'm just pointing out that the machine doesn't know what anything at all is. All it can do is work with what humans are telling it. If the real world is biased it can then itself be biased.
There are many feedback loops in societies. The rich get richer and the poor get poorer and those subjected to discrimination often have objectively worse outcomes. An algorithm can very easily be susceptible to just recursively documenting actual bias that exists in the real world and giving it the veneer of objectivity.
It's an actual hard problem that is the topic of this thread, hand waving isn't going to fix it.
> If the real world is biased it can then itself be biased.
And if said algorithm corroborates that bias, then is reality somehow "bad" as a result, or do you just not like the outcome because it "feels bad"?
There is a difference between truth and an idealized outcome that you desperately want to be true.
Also for historical reasons all the green people live in certain neighborhoods.
Now imagine creating bureau that would review driving and accident data and attempt to answer the question "How good at driving is this person?" The idea being that we are going to try to keep the worst drivers off the road.
People scream hey wait a second that's not fair at all. So you take steps to address this whole system instead and solve the real problem. You try to crack down on all the rock throwing, and not prejudge someone without knowing the circumstances.
Until one day an algorithm comes along to answer the question. But wait, we thought of that, we make it so the algorithm can't take into account what color the driver is. It can only rely on other things.
See the problem yet?
The algorithm won't corroborate the bias—it will include the bias, because the bias will be built into its data. It's not an independent variable.
An algorithm is a set of steps given to it by a human. If the human can discriminate then so can the algorithm.
In the case of ML, let's say you are making lending decisions and your data set consists entirely of the race of the applicants. Obviously that ML algorithm would illegally discriminate.
Now, extrapolate from that extreme case and say that you feed your ML algorithm a whole bunch of random data points about your potential borrowers.
How do you prove that your ML algorithm is not making its decision based on data points that are proxies for race?
If you don't have an answer to that question good luck with the swarm of lawyers who will be suing you into oblivion.
The error they make is always the same. I highly recommend this piece: AI 'Bias' Doesn't Mean What Journalists Say it Means https://jacobitemag.com/2017/08/29/a-i-bias-doesnt-mean-what...
> However, there is a significant cost to forcing algorithm outputs to reflect wishful thinking. If we are issuing loans, we will issue more loans to people who do not pay them back. If we are making parole/sentencing decisions for convicted criminals, we will release more dangerous criminals who go on to commit new crimes.
> This may not be a concern for the journalist, but it should be a concern for the rest of us.
I agree that this is the heart of the issue but by leaving this to the conclusion, I think the author didn't end up arguing for why we should optimize for objective cost instead of for our ideal reality. It might be financially profitable for a bank to use a racist lending policy but as a society we think that it is wrong so we passed regulations forbidding the banks to do so. Similarly, the criminal justice system is supposed to serve society and follow our ideal sense of justice and shouldnt be blindly optimizing recidivism rate. As the Nautilus article said, "What gets chosen is usually whatever is easiest to quantify, rather than the fairest".
This article is trying real hard to find victim's anywhere and everywhere. Not that long ago there were many people talking about how wonderful the future was going to be with algorithms and big data to make unbiased decisions. Now we are basically there and they still find a problem with it because of "calibration issues" which were mentioned to be due to mathematical incompatibility. You'll also note the wording about the improved algorithm being "independent" of race. Not that the results were more fair or accurate, just "independent".
Then there's the ever popular gender pronouns that were mentioned. Guess what. In nearly every country in the world doctors are majority male and nurses are majority female. Funny that we never hear these "gender biases" mentioned with things like "___ is a construction worker" or "___ is a bus driver", isn't it? Those are male dominated fields as well yet no one ever talks about gender imbalance or discrimination. It's only ever with the higher status/income jobs that you see this talk of imbalance. Many of these aren't "biases" but rather statistical probabilities.
The problem with modelling real world systems is there are too many variables and we have no way of knowing just how many their are and what effects they all actually have on a system. We can never make a complete model yet we act and make decisions as a society as though our models are complete and infallible.
Its gotten to the point where I do not click media links where race is included in the title. Wild accusations of injustice get click$.
Also, saying that there is one system of inequality doesn't mean there aren't others. Class is certainly another, related, system. But one doesn't reduce to another.
Yes it does. Every day I see people far more affected by class based inequality than race based inequality. Right now not having money would affect you far more than being a different race than you are.If you suddenly became a different race overnight chances are your life really wouldn't change that much, if you lost your job and all your money you'd be totally fucked. That means by definition one has a higher impact on the world.
It sounds like you're saying that one is more important. I honestly don't know whether class or race-based inequality is more of a problem; I think it depends on context. Both of them have caused great problems. I will point out that many people have been as "fucked" as imaginable for being labelled as a particular race: murdered, enslaved, forcibly displaced from their homelands, etc. And this sort of thing continues to this day, e.g. in Myanmar right now the Rohingya are arguably being genocided. If one of those people could change their race, they would be decidedly unfucked in that context.
When is violence just violence. When is it racism?
You see humans are terribly biased. Racial bias isn't even anywhere near the strongest bias we have. Unattractive people get sentences twice as long as attractive people. Judges give far harsher sentences before lunch, when they are hungriest. Socially awkward people seem to be pretty strongly discriminated against. Studies have found people discriminate by politics even more than race. Job interviews have actually been shown to degrade performance over just judging resumes. Before statistics and credit ratings, getting a good loan required being an old friend of the banker.
It's not just that humans are unfair. We are objectively terrible. Very simple statistical algorithms beat human "experts" in almost every domain they get tried on. Way back in the 1920s, a statistician came up with a formula that was better at predicting recidivism than a group of 3 prison psychologists.
Simple linear regression has predicted the success of medical treatments better than doctors, diagnosed psychoticism better than trained psychiatrists, predicted academic success much better than admissions officers, predicted loan risk better than bank officers, etc, etc. To say nothing of modern machine learning methods on modern computers. It's insane we allow humans to continue doing these tasks at all.
But there has been huge resistance to algorithms in every domain. From people who stand to lose their jobs and be put to shame by them of course. But also even outsiders tend to reject algorithms. And overly trust humans. Psychologists have actually studied this. They call it a bias labelled "Algorithm Aversion". (http://opim.wharton.upenn.edu/risk/library/WPAF201410-Algort...) The study showed that humans were willing to forgive the mistakes of humans far more than those of algorithms, even when the algorithm made far fewer of them.
This is why they aren't everywhere already. The last thing we need is fear mongering articles like this. As shown, humans are far worse. If an algorithm shouldn't do it, then a human certainly should not.
A big part of the case this article makes is a reference to a propublica study that found a slight bias against race in an algorithm once. Yet that study wasn't peer reviewed. The findings weren't statistically significant, which is a pretty low bar to begin with. And yet it always comes up as the main reference in these discussions. It's the only piece of evidence they can find of this nonissue ever happening. Stop. You are just making the problems you claim to care about worse.
Arbitrary clusterings such as zip codes, while certainly statistically significant, do not translate to intent by algorithms or algorithm makers.
Fairness is somewhat subjective, and I think it's fair for companies to pursue profits such as "same day shipping", even if that cannot be simultaneously rolled out to all markets. If it is profitable, eventually, it will be.
[0] Invoking Betteridge's Law of Headlines ( https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline... )
Sounds like the answer is yes.
What are the chances this won't be co-opted by the other side, to enhance divisions?
That's particularly ironic, since a lot of ML can treat it that way.
intention is orthogonal to structural bias
from the last section of the article: "infrastructure beats intent"
Why don’t you engage with the content of the article: But when COMPAS’ prediction was inaccurate (either predicting re-arrest when none happened, or not predicting an actual re-arrest) it routinely underestimated the probability of white recidivism and over-estimated the probability of black recidivism. In other words, it contained a bias hidden from the perspective of one set of statistics, but plainly visible in another..
Edit, because you noticed how unhelpful “ha ha, Betteridge” by itself is: intent may not be necessary for discrimination, nor can you deny intent when you continue to use methods shown to produce discriminatory outcomes.
A profit motive is similarly useless as a defense of discrimination. Maybe young, white, heterosexual female employees are better for sales, because you operate in a xeno-, homo-, gymno-, gerontophobic neigbourhood. But it’s no excuse to limit your job ads to that demographic.
One reason is that such logic perpetuates its own causes: if the police is more likely to stop&frisk non-whites, you will have fewer whites in your arrest statistics, even if they commit the same number of crimes. To then turn around and use those statistics to justify your racist police procedures is both morally and scientifically bankrupt.
Civil law isn’t about morality. It’s about righting a wrong. If you die because a mechanic negligently forgot to reconnect your car’s brake lines, you (and your family) will be harmed. The question for the courts then becomes shifting that harm to the party responsible, as far as that is possible, and your family will likely be awarded damages.
That assumes companies are completely logical, singular entities which all follow the same motivation without deviation. As companies are the sum of the humans working there with all their personal problems, faults, opinions and attitudes (plus the organizational framework itself) that's rather unlikely.
It also assumes companies always know beforehand what the best strategy is to make more profit (just look at VW for a recent example where that didn't pan out).
Yet historically, a number of businesses in various parts chose not to serve certain sections of the population, and sometimes still do (off the top of my head, wedding cake companies choosing not to sell to gay couples).