You want it fair based on race ? Ok, don't input race. Everything else stays the same. Why this complexity ?
You want it fair based on race ? Ok, don't input race. Everything else stays the same. Why this complexity ?
There are several reasons for this from both technical and legal perspective.
It is incredibly easy to find statistically significant correlations given just a few (more than 7) different views of the data. In general these ml models are not working with less than hundreds or thousands.
If the model learned this suppose racial bias, once, you deleting this column is not going to stop it from learning it again, and I believe some research showed that it actually can make the unfairness more severe.
from a legal standpoint a company that may or may not be infringing on rights could just say, oh we can't be because we don't have these fields in our data: which makes it harder to monitor and audit wrong doing.
most of the methods that I am familiar try to ease the effects of the learned biases as a post-processing step for the model.
I would be interested in seeing examples. So far in this thread the arguments were along the lines of: "but then the algorithms punishes poor neighborhood instead of race" but you shouldn't have address in the data either as (I hope) nobody is ok on punishing people for living in bad neighborhood. We should only include data we would like to see in the explanation of the sentencing.
"You are not getting parole because you live in a poor neighborhood" is unfair while most people would be ok with:
"You are not getting parole because you willingly associated with people who committed crime".
Why ? Because crimes aren't committed by all groups equally ... so you wouldn't expect the risk to be equal.
Even for completely practical reasons there are differences. For obvious reasons a 1.4 meter individual is limited risk for armed robbery, to give an extreme example. But there's hundreds to thousands of things like that, that all randomly affect the distribution. The result is that when all put together, things are seriously out of whack.
Also I fear the reverse effect. Communities work, to some extent, because of lack of crime. Doing this will necessarily make those neighborhoods with the "victim" genes more violent, more criminal. The result will be improvement of those neighborhoods ? I have to say, I'm VERY skeptical.
here is a good resource on that: https://arxiv.org/pdf/1606.08813.pdf European Union regulations on algorithmic decision-making and a “right to explanation”
What you should do is to keep the ratio of people classified as dangerous same among different groups. But this comes at a price as reduced classification accuracy.
Of course you don't include location. Why would anyone be ok with "you got longer sentence because you live in poor neighberhood"?
>>What you should do is to keep the ratio of people classified as dangerous same among different groups.
Sounds like a terrible idea. What if there is a community which just doesn't contain many dangerous people at all? If someone from this community commits a crime once in a blue moon they will be unfairly punished.
Now imagine if people were unfairly punished based on something they were born with and not where they lived.
No, the algorithm just gives negative weight to locations where dangerous groups tend to live. You infer race from that measure.
That said, it'd be interesting to see how other "algorithmic bias" might play out in more or less racially homogenous, if not entirely ethnically homogenous populations --for example Rwanda, Finland, S Korea. Things like class, neighborhood, education etc.
[1]Given they themselves don't officially collect race information.
Or let's say the algorithm is deciding on insurance of some sort's premiums and someone is living next to a SuperFund site.
I don't want to punish people for living in polluted areas therefore I don't include that data in the training set.
>>Or let's say the algorithm is deciding on insurance of some sort's premiums and someone is living next to a SuperFund site.
I am not ok with doing that therefore I don't include address in the data.
Am I ok with punishing people for living in poor/dangerous neighborhood? No? Then don't include that in the data. Am I ok with people paying more for car insurance if they drive in a city with more cars/accidents/worse driving culture? Yes? Then I include it in the data there.
I'm not sure I can reconcile that.
1. in the first option, we're not applying the consequences random people, we'd be affecting those already found to have committed a crime.
2. in the second option, if we think we'd be affecting innocent people in option 1. why do we think we would not be affecting innocent people (good drivers who happen to live among poor drivers) in options 2?
That's to say, either both are right or both are wrong. I see the same effects in both.
Here's a very simple refutation:
Is past arrest history a valid factor in the model?
What if past arrest history is biased due to biased policing (Hint: it is, well proven)
How about zip code?
If you don't believe in using probability to determine risk. How about you hire a pedophile to babysit your kids? How about throwing down your retirement on lotto tickets? ETC.
Do you think biased policing causes people to commit crimes? How about not committing crimes to begin with? How about not hanging with troublemakers? You think it is impossible for people to avoid being arrested and convicted of crimes because of "biased policing"? Please think harder.