Obviously that is absurd, but where does it go wrong?
Defense means 'what are you going to say when the feds come knocking and want proof you are not discriminating against a protected class?'
As a practical matter it isn't too hard to come up with a legal defense for that, of course. Lots of plausible deniability here. IMO most of this effort is by people trying to help prevent computer nerds from becoming (or staying, as we may already be there) complicit in discrimination by race while hiding behind algorithms as if they are somehow infallible.
want proof you are not discriminating against a protected class
If black people are poorer on average, then even if the system takes only takes personal history in account, black people will inevitably get worse credit rating on average. It just follows from the premise, unless banks participate in some form of race-based redistribution which I find distasteful. How's it supposed to work in your opinion?
It might well be the case that a totally 'colour-blind' scoring algorithm still ends up sorting people by one of those mentioned, unrelated data points. This is not racism, it is just the consequence of the mentioned relationships. Calling it racism is a sign of ignorance and does a disservice to the effort to get rid of true racism.
Each has a plausible justification for inclusion. But the more data points you have, the better you're able to predict race from them (or anything else you're not supposed to).
So how do you decide if an algorithm is discriminatory? Maybe a whitelist or blacklist of what data it can use?
If you use time in bar to predict cirrhosis it's ok.
etc
We may say that now, but what if later some algorithm finds there's a hidden correlation between driving performance and race/gender? will it still be an ok metric then? where do we draw the line?
The line is between inferences of bad behavior -> bad result such as "you drive risky so you must be at risk of an accident" and pure correlation "you drive risky so you must be a man so you must be at increased risk of prostate cancer".
And conclusions should be assumed to be of the latter kind unless shown otherwise.
The solution is pretty much that simple. I would argue that all correlative models dealing with human beings will always be discriminatory.
Will loans become more expensive for many people? Sure. But that is the true solution. Sometimes the right decision is simple but unconfortable. In my opinion, this is one of those times.
Your financial history includes your rent or mortgage payments, which are for living at your zip code. Restriction to financial history is not all that much of a restriction.
Of course, things are not so clear cut. Your financial situation can also be a consequence of discrimination, e.g., if you were denied a job or unfairly arrested. At the end of the day, some human has to sit down and decide what is OK and what is not. Personally, I believe the goal of research on algorithmic fairness should be to give the people who will ultimately make these decisions (e.g., judges, politicians, etc.) the tools to understand both broadly, the kinds of things that can go wrong when using algorithms, and also to understand what might have gone wrong in a specific situation.
But zip code? I understand there is empirical data showing how race is correlated with zip code, but you still choose your zipcode at the end of the day. If you have data saying that applicants from "Zipcode X" are less likely to repay loans, why is one's choice of zipcode out-of-bounds but their choice to have a hard credit check several months ago not? What if a poor credit history is correlated with race (fwiw I don't know if it is or not)? Is it then discriminatory to use credit history as a covariate?
The truth is that machine learning to predict people's behavior is inherently discriminatory and involves generalizations by definition. Once you start disallowing training on covariates that exist due to voluntary behavior, it's not really clear whether you should be allowed to use machine learning to predict individual behavior at all (I'd probably listen to that argument actually).
There's clearly a tradeoff to be made in terms of how much discrimination is allowable in exchange for a given unit of predictive performance, but that sounds incredibly difficult to regulate for every business's ML problem. I don't know what the answer is, but I think it's hasty to argue that anything that's correlated with race is out-of-bounds.
Unless you live in a city that was redlined[1], which is to say, unless you live in virtually any city. In which case, you really did not and in many cases still do not actually have much of a choice where you live. And that's without even looking at the fact that the racial wealth gap means that even if there weren't still strong barriers to integrated neighborhoods, lots of members of groups that were and are discriminated against would be unable to move into the neighborhoods that would theoretically let them escape that zip code bias.
[1]: https://www.chicagomag.com/city-life/August-2017/How-Redlini...
https://www.kff.org/other/state-indicator/poverty-rate-by-ra...
As an example, some 'professional' sports are dominated by people of certain skin colours. Basketball is predominantly 'black', ice hockey predominantly 'white' while soccer is more representative of the population at large. If you're looking for people who have the potential to become good basketball players you might look at data on their athletic achievements, on their length and other such things. I don't think there is a need to have a dark skin to be able to play basketball at a high level even though the majority of players seem to have such. I assume scouts for the NBA, NHL and MLS don't profile people based on their skin colour but on the aforementioned things and others related to a person's ability to play the sport.
[1] http://www.chineseorjapanese.com/wp-content/uploads/2009/04/...