> But look how that data came to be -- by police making arrests.
Yes. And rape statics come from people getting rapped. I'm going to assume you are talking about convictions here, as arrests and convictions are different.
> And there is also plenty of data out there showing that black men get arrested more often than white men for the same crimes
Ok? Then add that as another data set to your model. You will also need to add in crime relative to the population. Ethnicity of that area. et cetera. You can't just make this statement, say it affects your conclusion without adding in everything else.
This fact doesn't change the initial statement that "Black men are more likely to be criminals than white men." It exposes issues in the society we live in, not the data that is being reflected. Any data scientist worth his salt would be able to take this into account.
White Population of the US is 62%. Black Population in the US is 12.6%. White arrests in the US were 70%, Black arrests were 27%. Whites are over represented by 8%. Blacks by 14.4%. Asians are under represented by 4%. FYI.
https://en.wikipedia.org/wiki/Demography_of_the_United_State...
https://ucr.fbi.gov/crime-in-the-u.s/2016/crime-in-the-u.s.-...
> and get longer sentences for the same crimes than white men.
This is irreverent to your point. Ignoring.
> Now imagine building an AI (or actuarial table) based on that data. It would necessarily identify black men as "more likely to be a criminal" simply based on the fact that they are black.
A Machine Learning model, not AI (as AI is incorrectly used everywhere), would come to that conclusion- yes. The problem with your argument is that you want to bring your bias into a system that needs to be based on facts. Do black men get shafted? Yes. But approaching the problem at the end of the pipeline doesn't solve anything.
Going back to the initial discussion point on insurance rates. Men are more likely to die behind the wheel. Ok. We don't fix the problem by saying the data is bias, wrong, whatever. We fix it the source. You're entire argument boils down to "I'm not comfortable with what the data is telling me and I want the data to say something else." Which is emotional- I get that.
> So now you've magnified the bias in the data.
You've built a system that reflects the realities of the world around you. There was an article somewhere about a robbery on a BART train. The police wouldn't release the ethnicity of the suspects for some stupid reason about not wanting to feed into stereotypes. That does nothing to fix the problem, it just tries to cover it up. But, your reaction is the right one you just are not focusing it correctly. Instead of asking why our models look like this, and trying to manipulate them to make you feel better you need to ask why the data is happening the way it is.
My point is this. The conclusion of "Black men are more likely to be criminals than white men" should make you want to fix the reason why, not try and manipulate the data or system that produces that result. If you have a Machine Learning Model that draws that conclusion, you have two options:
1. Fix the source of the data. Isolate why Blacks in the US are over represented in the criminal population by 15%, and solve that.
2. Add in other data sources to further refine your model.
The fix is not run around screaming bias, as that's an incorrect characterization of the situation and actively hurts your cause.