Even if being an immigrant is correlated with higher rates of some crime, an algorithm cannot take that into account without bigotry. Imagine being an innocent immigrant. If the algorithm tags you as guilty, no amount of explanation about averages or neural nets are going to make you feel better about your erroneous charge. This is the reason “stop and frisk” was such a huge deal: it let the cops “racial profile” for drugs, meaning that at any time all black and brown people were more likely to experience 4th amendment violations by the cops.
Our language around probability doesn’t help either. We say things like “being an immigrant makes it /more likely/ that you will take some action”. There might be a correlation in aggregate data but in any individual case we can only look at the relevant facts. Think again about being an innocent immigrant. Would you accept someone saying that /you personally/ were more likely to have committed a crime? Of course not!
The moral of the story is that stats describing a cohort cannot be projected onto individuals in that cohort. The data flows the other way; individuals comprise an average, no individual /conforms/ to an average. They may happen to coincidentally be average, but no causation has occurred.
Our justice system is supposed to be “innocent until proven guilty”, not “innocent until Bayesian inference suggests high enough probability.”
"What would be wrong though is to blindly take that probability for the fact."
And it's right: your algorithms can use all sorts of indicators to flag suspicious cases, but in the end the evaluation needs to be done by a person on a case by case basis. As long as the investigation on the single case is performed "silently" (i.e. without causing any loss of time or money or any annoyance to the person who is investigated) I don't see the issue. Otherwise it's the same as saying you cannot sort rows in an excel sheet according to some criterion and start looking at the data from where it looks more promising.
As a concrete example of where this can go wrong, let's say that (hypothetically) 1% of people commit fraud, and that this prevalence is not correlated with race. But, due to past discriminatory practices, more people of one race than another are convicted of fraud. Then, this algorithm would perpetuate this discrimination, since, say, it would identify 80% of fraud in the discriminated group but only 50% of the non-discriminated group.
Also, dual citizenship is not a protected class last time I checked.
A model that does not incorporate all the relevant information will have more false positives. The practical consequence of doing what you suggest is, cetaris-paribus, that you will incorrectly penalize more people.
> There might be a correlation in aggregate data but in any individual case we can only look at the relevant facts.
Definitely. Thats where courts and human review come in. The mistake the Dutch made here was not the variables in the model, but that they had limited human review of a model with (I'm assuming) low predictive power.
> Our justice system is supposed to be “innocent until proven guilty”, not “innocent until Bayesian inference suggests high enough probability.”
"Guilty beyond a reasonable doubt" (or whatever the Dutch equivalent is) implies exactly the Bayesian calculation you are rejecting. That said, I agree that a machine learning model with limited inputs is not sufficient to determine guilt.
A lot immigrants from Marocco and Turkey own real estate in their country of origin. Their IRS doesn't readily share information with ours.
See: https://www.uwv.nl/zakelijk/images/handreiking-inkomen-en-ve...
That's insane. Men commit 99% of violent rape, imagine not being able to say that. And I say that as a man, innocent of rape.
Similarly, if people with dual-citizenship smuggle more you're wasting time by treating everyone equally when looking for illegal imports.
> If the algorithm tags you as guilty, no amount of explanation about averages or neural nets are going to make you feel better about your erroneous charge.
Yes, any black-box trial is crap. But a transparent algorithm based on real data that a certain subset of the population is more likely to commit a crime, which is used for proper planning and investigation, is not.
> Our language around probability doesn’t help either. We say things like “being an immigrant makes it /more likely/ that you will take some action”. There might be a correlation in aggregate data but in any individual case we can only look at the relevant facts. Think again about being an innocent immigrant. Would you accept someone saying that /you personally/ were more likely to have committed a crime? Of course not!
You have one proper point about terminology and then an appeal to rage.
You are right that nothing my demographic does (men, raping) makes me more likely to rape. But you're wrong later where you say "[claimed] /you personally/ were more likely to have" because, yes I (by the nature of being physically capable of rape) am more likely to have committed it than other people. To me I'm not a statistic, to you I literally am a population sampling.
> Would you accept someone saying ...
If they're right, yes. Otherwise you're just saying that my outrage should trump the truth.
If a woman wants to organize a women's only bus, or hotel room, because she fears the harm I as a man could commit I shouldn't have the right to force myself on her.
> This is the reason “stop and frisk” was such a huge deal: it let the cops “racial profile” for drugs, meaning that at any time all black and brown people were more likely to experience 4th amendment violations by the cops.
No, the problem with stop and frisk is that it was (allegedly, I'm not from there) used by racists to target black people. If it was used as commanded by an algorithm then it would only happen where data showed an actual correlation.
What if somehow we as a society magically get equality overnight and there's no more connection between race and wage? But you use historical data where they were connected, still excluding wage data, and so even now your model connects race and fraud.
In this hypothetical example, I'd argue it would be inethical to include race without wage. Hell, wage is only going to correlate with fraud via yet other variables. Maybe an ethical system needs to go out and gather those measurements too.
(side note, in an ML system with regularization, even if you do include wage in your data, the regularization might pin some fraud on race anyways.)
The solution is to acknowledge that there are always going to be omitted variables and either a) be extremely careful and rigorous in your data gathering, design, and roll out or b) don't try to automate this thing, leaving it the slow expensive way where people can gather facts as needed, case by case.
In this sense, dual citizenship is basically code for being Moroccan.