One, if datasets are biased --- if you build your system to only work on white males --- then it may have suboptimal results for other groups. This is a common problem: you use your company's faces, or college students enrolling in data gathering exercises, etc, who are not representative of the population at large. We can fix this by being careful about dataset bias.
But the second issue hits right at the heart of a major societal problem/debate. When we use AI to make decisions about people, will the system become racist--- even with representative datasets? If you train something to predict, say, odds to default on a loan-- will it figure out things that correlate to race and be making mostly racial decisions? Different races do have different default rates, but we've decided as a society that it is unfair to use race to determine an individual's probability of default. But if we choose things that are correlates of both default rates and race, when is that fair and when is it just veiled racism (redlining)? What things are measuring a causal relationship and what things are just racism in disguise?
This second problem is much worse with ML, because we have the ability to accept a whole bunch more things into our models and explaining the rationale of why decision are made is much harder.
And of course, the first problem-- both bias in datasets during use and biases during research and development -- makes the second problem worse. It can't even necessarily be addressed by broadening the dataset and retraining: if, in this case, you do your research and training with just white faces, and report positive results, it may not generalize to work as well for everyone with a broader dataset.
E.g. for a non-ML example -- redlining is theoretically "blind" to race, but makes extremely racist decisions.
ML will figure out both but can't explain the hidden variables: it'll figure out behaviors that directly indicate higher risk, and behaviors that indicate you're a minority that tends to be a higher risk.
This is a hard problem.
Many people would say that identifying people who spend a lot of money at bars or casinos as a credit risk isn't racist, even if it happens to pick up more minorities. The mechanism for the credit risk and behavior seems tightly correlated.
Many people would also say that identifying people who spend some money at clothing retailers that market to minorities as a credit risk is racist. Here, the relationship seems like a hidden way to spot minorities, who happen to be more of a credit risk.
When banks drew red lines around all the minority neighborhoods and didn't lend to people there, because they couldn't consider race anymore, most thought this behavior unacceptable, even though it did genuinely reduce credit risk for banks. ML can't explain its rationale, and very often does the exact same thing inside an opaque box.
My tendency is to think that since it is optimized to maximize profit, the more features about a person we get in the data, the less "racist" the system becomes under this definition. Yes it's more possible to "hide" racism in a complex method using tons of features, but increasingly less likely it would happen. If we can use income and debt and whatever other things to make either a good classifier that ignores race, or an inferior one that sneakily identifies race and bases the decision on that, mathematical optimization of the model should result in the former. ML can be trusted more than humans in this situation, not less.
Of course there's still the issue of the data being biased, which is where it all started.
I'm guessing it's illegal to discriminate based on race, gender or ability no matter how much it extra it costs a business, whether its deliberate or not.
This makes "buys fashion for minorities" quite an accurate predictor of high credit risk, but also a profoundly racist one.
Strawman. The question is whether predictions are improved by using race (or direct proxies for race), not whether it'd be wise to use proxies for race alone.
Loan default rates are higher for disadvantaged minorities, even after controlling for many, many other variables (income, neighborhood, level of education, etc). Therefore, using race (or inferring race) improves prediction quality, but is ethically dubious.
Can we skip this kind of snippy arguing please? Anyway I take it you agree the statement is true but don't want to.
So now you're saying those great minority loan customers are impossible to identify? I think you just need to figure out what information is still missing. What's the effect size at this point anyway?
> Can we skip this kind of snippy arguing please?
If you don't start off by willfully mischaracterizing your opponent's argument in order to be able to more easily refute it, I think you'll find that they accuse you of this less.
I think you're just willfully missing the point.
Ideally, you'd just lend money to the people who will pay you back. Unfortunately, we can't predict this perfectly. Adding race, or proxies for race, to the things you consider improves your predictions somewhat.
And what have I claimed to "easily refute" exactly? More like I ran with the definition, and considered how to address the problem as stated. I said more features were needed, I didn't say enough features were currently used. You keep pointing to a dichotomy between perfect and flawed, while I was talking about relative improvements. There isn't even necessarily a disagreement there.
It is mainly deep learning for which this is difficult. Which I'm sure is the method in the twitter discussion, but that is a different kind of problem where interpretability doesn't have a clear use.
They may be less inscrutable than deep learning, but can easily still draw a box around most of the black people in a clever, non-obvious way.
https://consumerist.com/2008/12/22/amex-lowers-your-credit-l...
Capital One has admitted doing the same, as well as a number of other lenders.
There's been analyses done since, that show that when controlled for all other variables in credit reports, credit issuers issue less credit to blacks.. e.g. Cohen-Cole, 2011, Credit Card Redlining, Review of Economics and Statistics
You asked for a source.
I provided a source of the same. Yes, things like purchase histories are used as proxies for race and used to deny credit to mostly people of color.
Then you go here:
> Credit Card applications do not ask for race.
which was never asserted.
> I suspect it’s more likely based on zip code or similar.
Yes, zip code can be another proxy for race.
As for the latter, no, it’s not a proxy for race. It might be highly correlated with race, but that is very different. I don’t know why everyone does these gymnastics to find a racist angle to everything, but it’s nonsense and it’s very harmful to spread that narrative.
Nobody is searching for black zipcodes and using that as an input to deny credit. The goal is to increase revenue by giving as much credit as possible with the least risk possible, so it wouldn’t even make sense.
Scoring based on where you shop is still behavioral analysis. I think it’s a bad policy, but it isn’t targeting people of color. Your zip code affects insurance rates too, and that’s based on claims. It doesn’t cost more to insure your car in south side chicago than beverly hills because there are black people there, it costs more because there are more claims. The same effect is seen in “white” zip codes that are in areas with severe winters.
Black people default on loans more, adjusted for income, education status, etc. But we've decided it's socially un-okay to ask people if they're black and adjust how much we're willing to loan based on the answer.
But instead, we measure all kinds of ancillary things, many of which are highly correlated with being black and not obviously correlated to ability to repay, and dump them into a model, where we come up with weights that make basically the same decisions as if we'd asked if they were black. Is this is fundamentally better somehow?
Mind you, I don't have a wonderful answer as to what to do: it's important to accurately price credit risk. But it's also important for society to treat minorities equitably, especially when inequitable treatment might reinforce the very problems/reasons why some minorities are worse credit risks.
It might be better to explicitly have the "is African American?" question in the model, because then regulators, policymakers, the public could perhaps eventually know the exact contribution of this factor... while approaching it obliquely makes the effect far less clear.
For example, from the outset, would you object to an AI that made decisions on how harshly to sentence someone based on age & number of prior crimes?
What about hiring or admitting people to college based on the results of an IQ test?
Yeah, those end up being badly biased, this first has been studied and the second is the reason for dropping SAT/ACT scores--every mental ability test correlates with IQ.
The more interesting thing, IMHO, is that the opposite is also helpful. For example, if you help all poor people equally, you help even the playing field by disproportionately helping out all disadvantaged groups.
Education is different. IQ shouldn’t be relevant but competency and aptitude should and in addition we should recognize some people are better off going to vocational school where they may do better economically during their lifetimes.
Yes. This is an overly simplistic system for which it would be hard to justify the cost of developing.
>What about hiring or admitting people to college based on the results of an IQ test?
Yes. For the same reason.
For the second, any mental ability test correlates with IQ, so you end up with a different set of difficulties. There have been attempts to, e.g. correct for cultural bias in the tests, but these actually made the problems worse.
I'm not presently aware of anything that makes that situation better.
If you're looking at things that correlate, you've got a lot of choices. Saying that your data set is representative, your algorithms are correct, doesn't mean your choices of what correlations to use are objectively or morally correct.
Suppose that you have a system that accurately tells you men have higher car insurance claims on average than women.
The same system might also be able to tell you people with low credit scores have higher claims on average than people with high scores.
Is it right or ok to use the first correlation and ignore the second? Or vice versa? Or maybe both of them are unacceptable? There's nothing objectively inherent in the correlations that tells you it's ok to charge men who are good credit risks the same as men who are bad credit risks. Or optimal economically! Those are independent and difficult questions.
You're worse off lending in traditionally black neighborhoods; the question is, is it ethical to make a map of all of these traditionally black neighborhoods and refuse to lend there? A model that takes this into account would be biased in the same way the actual world is, but many would consider it deeply racist and unfair.
ML has problems with all of these at times. The particular case here with the Obama picture is a particularly vivid illustration of one that can be used to bring attention to the overall problems.
I feel like you've missed at least half of what I was trying to communicate, because you're still presenting arbitrary discrimination as reality-based.
Discrimination as you describe is not only unfair, but it costs the lender money, because they are not identifying the true risks of individuals within the group. Because it's suboptimal, it can't be called an unique, objective reflection of the world.
All correlations are inherently imperfect and therefore unfair; if they weren't you'd have an definite causal relationship.
When you use one correlation, or set of correlations, maybe you have completely missed something that would be much better, not just from a moral perspective, but a business perspective.
If you discriminate based on gender, and haven't considered age (hypothetically it being legal), maybe the latter would be a much better proxy for the real causal factors. Your gender model can be correct in itself, and yet from a completely amoral perspective be terrible and uncompetitive. The model is not "the way the world is" just because it's technically correct in isolation.
No model is perfect. Odds are you can better predict risk by (directly or indirectly) attaching bonuses or penalties to given races, all else being equal.
e.g. Say, irrespective of all other measured quantities, people of race A are more likely to default because of other systemic racism against them. As a lender, it'd be completely rational to consider this and make "better" choices. And if you're not allowed to measure race, it'd be completely rational to find other variables that don't have an obvious causal relationship to credit risk, but predict race and thus have some information about credit risk. This is the exact kind of thing ML does.
Then, in turn, this becomes self-reinforcing. Because other institutions discriminate against race A, the risks going forward of dealing with race A increase...
But some are definitely better than others.
It sounds like you're looking at things purely from the point of view of getting the correct average for a group. But whether or not you get the correct average doesn't tell you if you're using a good enough or the best available model completely apart from fairness or justice.
If you do some type of testing and you know 5% of the tests should come back positive, is there a difference between reporting 5% at random and actually doing the tests? Of course!
No, I'm looking at things from the point of view of making a model that fits well to the original dataset and then is verified in the actual accuracy it makes over time.
If adding race-- or inferring race-- makes the model substantially better in predicting outcomes, is it right to do so? Credit default risk is correlated to race, even controlling for other variables. Hence, using race would help you make more accurate predictions.
This entire discussion is about ML models, which can often be fairly described as very fancy curve fitting.
> And when inserting a fitted curve into a feedback system you're very likely to just perpetuate the problem you're looking to eliminate.
This is not a problem for the individual credit issuer-- they're not looking to eliminate the problem. They've avoided some credit risk by taking race into account. They've improved their expected value, even if society is stuck with the cost of the problem getting worse.
And I repeat curve fitting isn't modeling. Because in that case it's not a model it's a prescribed outcome.
That is, what's your loss function and how does it prevent race from being considered in the credit decision?
- a first network take the input data and return a representation A (like an embedding vector): let's called it the "censor network"
- a second network take this embedding A as input and is trained to predict the class that should be censored (for example the gender of a person) : the "discriminator network"
- a third network take the same embedding A as input and is trained to predict the real task of interest (for example the probability of credit default) : the "predictor network"
The idea is that, by training the censor to make the discriminator fail (predict the wrong class) while making the predictor work, it will force the censor to learn a transformation of the input data that keeps the task related information in the embedding A, but removes the information correlated to the "censored class" (and that could be used to discriminate).
Here's a reference about this kind of methods, but it's still an active domain of study in ML and there are many papers that followed this one: https://arxiv.org/pdf/1801.07593.pdf
If neural nets provide a way maintain plausible deniability while breaking the rules, it's going to get used.
Imagine you have two groups of people that have different means and distributions of “input” (could be intelligence, wealth, interests, ...), and this input is somehow correlated with some output (insurance rate, criminality, university acceptance rate). Both input and output are one-dimensional.
Then you only have two choices. (1) You “unify” the distribution and use it as joint input, resulting in one group (the one with the higher mean) getting “more” - many people think this is unfair. (2) Or you can subtract the means first and so make the distributions as equal as possible, which can result in a person from the “more” group having worse outcome than an equivalent person from the “less” group (if the means are close together and the distributions are wide, there will be many such people) - many people would consider this unfair.
There are plenty of examples, current and historical, of society picking one or the other, for different groups and topics: wealth/success/income (old vs young, Jews&Asians vs other Americans, men vs women, whites vs blacks), university acceptance (Jews used to be discriminated against (look up “Jewish problems”), now Asians are, men vs women [1], “positive” discrimination of blacks) - just a few. In many examples historically, we consider (1) fundamentally better than (2) - e.g. discrimination against Jews at universities - but it seems it takes the society a long time to reach this conclusion in each instance.
[1] https://en.wikipedia.org/wiki/Simpson%27s_paradox#UC_Berkele...
Approach #1 is to just go to maximize expected value, using all variables, including race.
Approach #2 is to use race to adjust the distributions as you suggest. Call this "affirmative action."
The inbetween approach #1.5, that we usually follow, is to make a value judgment about whether any individual item is OK to use.
Income? Of course it's reasonable to use income to determine whether to lend-- even if it's correlated with race. Race? No, that seems wrong. Location? Sure--- wait, you've drawn a box around all the black neighborhoods, nope! Buys certain products? Sure-- wait, you don't mean "buys products that blacks stereotypically like," do you?
#1.5 is already iffy, but becomes completely untenable once model complexity gets high and we use approaches that do not do well at explaining their decision: it basically becomes #1.
A major problem with #1 is it becomes self-perpetuating and reinforcing: if everyone is going to be biased against you, you're going to do worse, and hence the models are validated/more models are created with these assumptions. Everyone using #1 may improve their own expected value at the expense of a worse outcome for society as a whole.
Do you have any proof of that? The world is improving on all levels, even the income gap between blacks and whites has been narrowing [1].
We're not talking about 100% segregated populations like during slavery or patriarchy, the means between the groups are very close and there's a huge overlap between the distributions. It's obvious that the vast majority of outcome isn't determined by race (otherwise you couldn't have white homeless addicts and black presidents) or any such "group" characteristic but instead by "individual" characteristics (skill, intelligence, drive, effort, hereditary wealth, chance, ...), so the evidence for this "self-perpetuating" unfairness is very weak.
[1] https://www.pewsocialtrends.org/2018/07/12/income-inequality... but you'll have to manually calculate the percentages
I'm completely overwhelmed by your evidence of median black income moving from 59% of whites in 1970 to 65% in 2016, and appearing to not change substantively in the last 15 years of the dataset. If the overall trend continues, we'll reach parity in another 269 years.
And it hardly has been happening in a world where considering race and proxies of race is considered "OK".
This is an absolutely comical piece of evidence you advance to try and justify your view.
When more variables come up and get more complicated it seems like it would be impossible to say "You're based on race" versus "You're based on a multitude of factors, several of which correlate to race, and consequently you give more or better loans to race X over Y."
You find criteria that are not inherent to race.
Reluctantly using your twisted example, imagine the left handed are more likely to default on loans because the checks they send are illegible.
Then you could use the acceptance rate of written checks as your criteria instead of handedness.
Yeah this is tough, answering the question of the odds a person may default on a loan will probably reflect historical bias. In fact I'd argue that it should show this if the model is accurate.
The moral issue is what do you do with that information? Do you use that knowledge to improve the outcomes of the disenfranchised or do you reinforce the bias based on the information you have?
Both might be right, for their respective time scales, which is why it's so hard for either one to back down.
Many many ways to talk past one another. We need a small book on techniques to short circuit these situations and salvage conversations.
No decision is completely safe and no system gets everything right in the first pass.
All we can ask for is a serious effort to understand risks sufficiently well to reasonably believe that the system is safe.
Certainly treating PULSE as a simple engineering problem and fix-as-you-go isn't so bad, is it?
This has already led to numerous troubles in the applications for loans, job evaluation, criminal matters and so on.
Which I'm arguing about elsewhere... but demanding that university research about upscaling a face to be perfect to publish is a little much.
> . But success isn't a success when the dataset isn't representative and datasets used by academics as a rule are not representative
Yup, so, academics should never research this stuff or publish? That certainly is an ... interesting ... recipe for progress. And research progress is one of those things that could help us ultimately address some of these issues, because as the original argument makes clear: improving the quality of datasets is not enough.
This is pretty much the issue at its core.
The real question is, if Yann is so smart, why is he arguing with people on Twitter?