Would giving out more loans than is rational by excluding stuff like zip codes be a good thing? Wouldn't that lead to more defualts among those groups of people zip codes can discriminate against?
Would giving out more loans than is rational by excluding stuff like zip codes be a good thing? Wouldn't that lead to more defualts among those groups of people zip codes can discriminate against?
There's no Hippocratic oath for our profession, and in many ways that's important, because we create systems whose impact may very well outlive us and out-scale anything that a single medical professional could do. But that also doesn't mean we should operate in a utilitarian environment without constraints.
Sadly there is no profession for our profession. I often think the model for any software profession (if we can create that - something I doubt) is railway engineer - where the professional signs off on the safety / completeness of work done on the railway - and that no train can travel without it. It leads to plenty of uncompetitive practises - but also to ... y'know ... people not dying in crashes.
How we start that is hard (probably something to do with safety critical software systems) because we aren't too sure what is the right way to build software.
And then we have the fun problem of the members of the profession trying to decide the answer to your trolley problem. Sorry scratch that. The various legislatures proscribing the answer and the profession trying to implement the conflicting results !
Two ways where these systems may give out more loans than is strictly profitable, and they are both investments:
- Fairness. If you have a variable race and a zip code, you could account for discrimination via redundant encodings, while still using the feature for the optimal trade-off between a fairness criteria and model performance.
- Exploration. Concept drift (the correlational and causal meaning of variables shifts over time) can introduce wrong predictions. If all you have is few samples from a zip code, the model will always be uncertain. You can counter this by exploration and active learning: gather samples, not because this makes the model max-profit, but gather samples, to better learn how to predict these samples in the future.
But yes, giving out too many loans to minorities, may very well lead to further crisis and defaults, and tainting the credit scores of people will low access to finances even further. A bit like how well-meaning people donate money and food to Africa during Christmas time, then a few months later when donations subside, there are increases in famines. There is such a thing as "being too good".
Another startling insight from that talk was how normal services (i.e. home utilities, insurance, etc) will do credit inquiries and set a customer's rate based on their score. Which results in people with low credit scores pay more for services that are traditionally unrelated to borrowing. I'm worried that it could lead to a positive feedback loop that heavily affects those with poor credit scores in the long term. For this, it seems like a limitation on the kind of business relationships allowed to perform credit inquiries should be implemented.
The problem is, that's every parameter. The outcome you're trying to predict (will borrower repay debt) correlates with race, so other things that correlate with that outcome also correlate with race.
What you really need to do is to include as many non-race factors as you can, to give people with a couple of negative factors (who are often black or hispanic) more chances to redeem themselves with the other ones, to get the false positive rate down as much as you can.
Because the false positives (and false negatives) are really the problem. The true positives are... well, true.
It's an epiphenomenon of the interaction between this information and the normal workings of society on so many levels, and that feedback loop has been in place for a LONG time and is only getting stronger, now with real intentionality.