It used to be that if you tried to get a loan there was a loan officer in your town who had the final say. Someone who knew the locals and could override the "model" based on human knowledge. The Kennedys are good for it, their business is still solid enough, the Johnsons are really unreliable and I wouldn't do that.
I'm much more concerned with the "false positives" here than the "false negatives". The loan officer can give a few bad loans based on his gut feeling... and then he's going to lose his job. There's a direct feedback loop on them. Someone who doesn't get the loan to expand their business is going to suffer much more immediately and much more deeply.
That's what I'm fundamentally against - the abolition of the "loan officer" in this situation, a human who can countermand the models when they're obviously wrong. At the end of the day these are truly just classifier models and there's no guarantee that any given output is valid for a given input - someone has to maintain the feedback loop and keep training the model back.
And not only that but these aren't meaningless "training runs", each one can potentially screw up someone's life. So again, big consequences for error here.
And indeed the socially-just answer may not even be the mathematically correct one. Is there actually a check that your training model doesn't discriminate against black people? If you weight that to zero, are you sure it isn't going to start picking up the addresses where they live instead? Or names?
The problem with black-box models is they are designed to identify arbitrary or hidden features. Even if you forbid one feature they often will just find another proxy. That's what the models are supposed to do, actually. That's super problematic when there's nobody around to tell the model "no", and it can ruin someone's life.
I'm picking on black people as an example here because redlining is a blatantly obvious case of a rational individual decision with massive social consequences. But you can substitute in "high risk financial transactions" like handling lots of cash if you like. Those are pretty obviously prone to false positives just like redlining.
Frankly there's a lot of things about yourself that have recently become "public knowledge" that you probably don't want a government agent to analyze with a blunt instrument. For example, the USPS saves "mail covers" i.e. addressing information for all letters/parcels in the US. Or, based on commercial information that can be gathered without a warrant (reddit/HN datasets are on BigQuery let alone actual ISP- or forum-level data which can be subpoena'd - similarly courier services like UPS/Fedex can be subpoena'd without a warrant) they could analyze users to see what topics they post about on the internet. Post a lot about drug legalization? Our model says that's suspicious. Don't mind the dogs, they're just sniffing.