Is the goal of the AI model to predict crime rates in a hypothetical world where everyone has equal rates of lead exposure? Or is the goal of the model to predict crime rates in the real world?
Is the goal of the AI model to predict crime rates in a hypothetical world where everyone has equal rates of lead exposure? Or is the goal of the model to predict crime rates in the real world?
> Is the goal of the AI model to predict crime rates in a hypothetical world where everyone has equal rates of lead exposure? Or is the goal of the model to predict crime rates in the real world?
The goal is to use the results of a model for something (pet peeve, I hate the use of "AI" to describe what are usually pretty standard statistical or ML models). The model you create, and how you apply/interpret it, depend entirely on what you're actually trying to accomplish or change with the results.
Depending on what that is, the kind of "bias reflection" we're discussing is hugely problematic.
For example, crime rates are not equal between men and women. If we force our AI to assign equal risk of crime to men and women then we will have introduced a bias that either under predicts the rate of male crime, or over prdicts the rate of female crime.
What matters is: What are the outcomes and consequences of active systems, AI or not.
For instance: How do the algo cope with derivatives of its own output being fed into itself as input at a later stage?
The reality is, truth is relevant and sometimes the truth is inconvenient. Tech workers may want to build an AI that measures risk of recidivism that produces uniform risks across race and gender. But the truth is, rates of recidivism is not the same across all groups. If we produce the desired outcome of equal reporting of risk, then the consequence is that men have their risk underreported to put them on parity with women, or vice versa.
[1] See the 1999 documentary Blast from the Past.
But I think we're diverging considerably from the original point: that forcing an AI to produce equal outcomes despite unequal behavior in the real world is not the elimination of bias, it's the deliberate introduction of bias. If we have an AI that predicts recidivism rates, and we engineer it to produce equal predicted rates across all groups despite different between rates groups in the real world then we are deliberate introducing bias. The truth, regrettable thought may be, is that a magical AI that operates with 100% accuracy - the only people who it flags would have re-offended - is going to produce disparities because recidivism rates are not equal.
The blindspots end up being extremely problematic.
The justice system should absolutely not use this system for any purposes, since justice is based on the circumstances of the individual case in front of it, not the societal statistics which apply in the aggregate but may not apply to that specific case.
And if an activist engineer deliberately biased the model to avoid indicating disparities in crime, then we will have sabotaged police's abilities to allocated resources. Hence, why this assumption that disparate outcomes are indicative of a biased model is a problem
We already know where crime is occurring. We don't need an AI model for that.
People aren't arguing about not using biased data, they're arguing that the model needs to be designed and trained so that the bias in the data doesn't affect the predictions in the model. And yes, that means deliberately de-engineering bias out of the model, which may involve introducing a counter-bias.
For example, you and others kept bringing up race earlier as a legitimate bias for criminal profiling. But socioeconomic status is far more correlated with propensity to commit criminal acts than race. A model of crime in LA based on race, for example, would assume that people in Ladera Heights are just as likely to commit crime as people in South LA because they have the same race...but Ladera Heights has a fraction of the crime as South LA (and several times the average income). Similarly, you would expect South LA to have less crime than the largely Caucasian Joshua Tree or Fontana...but both cities have higher crime rates than South LA, and for a period were some of the most dangerous cities in California. (Joshua Tree was the inspiration for, and original setting of, Breaking Bad. Fontana used to be known as the Felony Flats.)
What I am saying is that pointing to disparities in the outcome of these models to claim that the models are biased is not, on it's own, a reasonable conclusion. As you point out, people in Joshua Tree and South LA have higher rates of crime than average. So if our model flags people from these areas as high risk more frequently than other places, is our algorithm biased? If we deliberately make the model to produce uniform results across different locations because an engineer feels that it's problematic to have a model that produces different results between different geographic cohorts, then have we mitigated bias? No, that engineer intentionally introduced our own bias to make the model adhere to his or her worldview.
Statisticians refer to this process as "controlling for confounding factors." What really matters is what questions you're asking. Data is too often abused, not always intentionally, by people with vague questions.
If you're going to use a model trained on simple, biased data, to get, say, insurance estimates for a for-profit company, the model will probably successfully increase profits, so it was a good model.
On the other hand, if you're going to use the same model to help with sentencing, where your goal is to see equality and justice, then the model will do very badly, since it will punish many people for the community/skin they happened to be born in.
For instance, men commit more crimes than women. If we are building an AI that predicts risk of committing crime (say, estimating rates of recidivism) and we forcibly make it report equal rates between men and women then we will be creating a discriminatory system because it will either under report the risk of men or over report the risk of women in order to achieve parity. Engineering parity of outcome in the model when the real world outcomes have disparities necessarily results in bias.
What a terrible example. There are many other features being omitted that would predict crime rates. If anything, this is an example of enforcing your own bias on the model by not including all relevant features.