That's the computer science problem being worked on here.
That's the computer science problem being worked on here.
Not everyone shares your politics.
I'm perfectly happy with algorithms detecting that certain people are more likely to be safe drivers than average, and giving them lower rates, and concentrating premiums on the groups more likely to be in accidents, even if I don't understand why Armenians (in your example) get in more crashes.
The distinction is that with young male drivers, we have two supporting classes of information:
* A clear statistical observation
* A conceptual understanding of why the observation is likely to be valid
With machine learning, we might have neither of these classes of information.
<shrug> What's the null hypothesis here? From my point of view, you're the one who defined into existence as a problem something that is not a problem.
A great example is how very resentful many young white men of college age are that universities are requiring them to take sensitivity courses designed to reduce the instance of campus rape, but strictly speaking men of that age are the overwhelming majority of bad actors in that environment. Statistically and logistically speaking, it's smarter and cheaper to just require all men of college age to take courses reminding them that rape is not okay rather than dealing with the moral, legal and healthcare costs of the alternative.
In some cases, our relatively primitive algorithms pick up on correlations that should not be acted on because we're actively working to correct them. For example, it would be inappropriate to pre-reject job applications based on skin color if in a certain culture, it's less likely for that person to have a college degree.
Even if that insight is correct, it's usually part of something that society hopes to correct or that applicants should be given the benefit of the doubt about, otherwise very serious negative responses will emerge.
Acting on existing categories may reinforce. It may not. For now, it's a case-by-case basis we'll have to act on. Maybe one day, modeling techniques and data sources will become sophisticated and robust enough to make every decision for us. That day is not today.
I'm sure that I've experienced increased costs based on this.
The issue is that I don't consider that the morality of forcing my will on other people depends at all on whether their current behavior is advantageous or disadvantageous to me.
Why on earth would I believe anyone who says this? I don't think humans are capable of such abstractness. It's not a matter of wanting to, biology itself is at odds with this mindset.
But also, there are many markets which are not effectively free markets. Insurance is a good one, and health service is an especially good one. In these cases, it's very dangerous to start agreeing that insurance agencies can start to pay "pass the puck" with human life.
And I'm not sure what sexual orientation has to do with any of this.
Nice bait, I guess?
The problem being worked on here is "what if Armenians shouldn't get car loans because they don't pay them back as much as other groups?" I.e., algorithms rightly classifying people leads to results that we believe are "unfair".
This research is illustrating that you can't simultaneously have accuracy and fairness. You need to explicitly decide how much accuracy you are willing to give up to get fairness. I.e. it's computing the tradeoffs needed to evaluate the ethical question: how many Armenian deadbeats should you extend credit to in order to be "fair"?
Go play with the simulation to see. The various fairness criteria all achieve lower than maximal profits.
A key result in the paper by Hardt, Price, and Srebro shows that—given essentially any scoring system—it's possible to efficiently find thresholds that meet any of these criteria.
The crucial thing is that this "credit score" doesn't actually predict defaults quite well, because blues who are 50% likely to default have higher score than oranges who are 50% likely to default. If you base decision on a simple credit score cutoff, the blue defaulters will screw you up because of their higher score. You can defend from this by tweaking cutoffs per-group to achieve profit maximization or various "fairness" goals. In a sense, it's an attempt at producing non-garbage out from garbage in.
And of course whether real-world implementations of this idea are "good" or "bad" is another can of worms altogether and varies case by case. They probably could have came up with some better examples than credit scores and labeling people by colors.
The best predictor is one which is explicitly discriminatory based on race: it takes both FICO score and race into account. There are worse predictors which also discriminate explicitly based on race, but in some "fair" way.
Finally, the worst predictor throws away directly relevant racial information.
This is not about garbage input data - that's not the problem being addressed here at all. The input data is perfectly fine. The problem is just that the input data says "race discrimination is the best way to make money".
> The problem being worked on here is "what if Armenians shouldn't get car loans because they don't pay them back as much as other groups?" I.e., algorithms rightly classifying people leads to results that we believe are "unfair".
No. If you stop playing with simulation and look at the input data, Armenians don't pay back less. They pay exactly as much as Iranians except that Iranians consistently have higher FICO scores because Iran infiltrated FICO with their suckxnet(TM) worm which replaces R binaries with hacked versions. Or something like that :)
The general abstract idea is: you have some input "score" which is know to inaccurately predict the outcome, use the input score and measurements of its biases to produce more accurate prediction than naive threshold classifier would.
And yes, the other guys talking about women's healthcare costing more or men causing more traffic accidents got it wrong too. I somewhat arbitrarily responded to you because you said something about "the real computer science problem being worked on here" and then continued to talk about other things like everybody else.
Why don't you quote the place in the paper where they make accuracy go up, fix overfitting, or build an improved risk score? Or even just quote a place in the paper where the risk score is treated as anything other than an accurate black box?
They pay exactly as much as Iranians except that Iranians consistently have higher FICO scores because Iran infiltrated FICO with their suckxnet(TM) worm which replaces R binaries with hacked versions.
In the example provided in the paper (see Fig 7) that's explicitly NOT true. Blacks pay back their loans a lot less than asians/whites holding FICO fixed. For example, at a FICO score of 500, blacks pay back their loans 10% of the time while Asians do about 40%.
I.e., blacks have consistently lower FICO scores because they don't pay back their loans. Further, FICO score is biased in favor of blacks. If we made it more accurate we'd be actively discriminating against blacks. For example, a black person with a financial situation reflecting a FICO of 500 would have their FICO score lowered to approx 450 to reflect their higher default rate.
Did you even read the paper, or the linked article?
I think we can agree that using the same cutoff for both races gives less accuracy than using higher cutoff for Blacks and lower for Asians, for some values of "higher" and "lower". You are right that "race discrimination is the best way to make money", but at the same time I think I'm right that this happens because FICO score already is racist to begin with, which is the "garbage in" I talked about.
But this paper is NOT about correcting that bias. If it was then it would be a pretty short paper:
Abstract: Go use isotonic regression in scikit-learn, once for each group.
References: Some papers from the 1970's when isotonic regression was developed.
http://scikit-learn.org/stable/modules/calibration.html
...but at the same time I think I'm right that this happens because FICO score already is racist to begin with...
This is also incorrect. The conclusions of the paper would remain valid if FICO were calibrated identically for all groups.
Can I suggest reading the paper and working through the math?
I guess you are right, this was just the first disparity I noticed in the orange/blue example and I got fixated on it. And yes, I haven't read the whole thing and my maths may be a bit rusty nowadays :)
Now I see that the problem they attempt to solve with "equal opportunity" is FICO's (in)ability to fish reliable borrowers out of the whole population. Currently, FICO identifies a small number of reliable Black borrowers whom they give high scores, plus there are many Blacks who would pay back diluted in a sea of unreliable borrowers with low scores. While in, say, Asians, the ratio of reliable borrowers who had been given high scores is higher (fig. 8).
I think the issue is a bit more nuanced than "algorithms rightly showing that fairness is opposite to profit".
In particular
> The best predictor is one which is explicitly discriminatory based on race: it takes both FICO score and race into account. [...] the worst predictor ["race blind"] throws away directly relevant racial information.
As I noted, this is a case of reverse bias directly compensating for FICO's bias. Max profit simply grants loans to all FICO score buckets whose default risk is sufficiently small to be worth it. If FICO vs risk was race-independent, "max profit" would use the same FICO threshold for every race and hence would be equivalent to "race blind".
Analogously, "max profit" could be equivalent to "equal opportunity" if FICO was better at finding Blacks who can pay and putting them in low risk buckets (high score) so that it becomes feasible and profitable for banks to grant them loans. I believe this is what authors meant by incentivising classifiers to improve accuracy and yes, I was wrong suggesting that they found a way to improve accuracy here by using the input classifier as a black box. Their ideas only compensate for the aforementioned race-biased score inflation, which isn't the entirety of the problem, and put some financial burden for some other mispredictions on banks/FICO to pressure them into getting their shit together.
I think trying to apply "equal opportunity" in the real world may indeed turn into handouts to people who can't pay, because <handwaving> it's possible that poor people are just hard to classify correctly and if certain ethnicities are poorer than others, they will appear to be given less opportunity even though actually it's simply poor people in general who are being given less opportunity. If FICO finds ways to classify poor Blacks better, it may turn out that applying the same solutions to other groups will improve their true positive rates too and hence "Black opportunity disadvantage" will stay.</handwaving>
But otoh, it also is possible that <handwaving> for some reasons Blacks are classified with less accuracy than Asians, which contributes to the lower overall true positive rate of the race.</handwaving> Authors appear to be assuming this possibility, though I don't think the distinction can be made from data presented in the paper alone.
TL;DR: I think you were oversimplifying things, and so was I :)
Your dispute may be with the authors.
If a computer program spots an irrelevant unhelpful correlation its generalization error will noticeable go up. If it is an irrelevant "helpful" correlation it means there is a problem with the data (such as leakage), not with the algorithm. If there is a problem with the data, all bets are of, both for black and white box models.
A blackbox model will probably not find that being an Armenian alone will lead to more crashes. Being non-linear in nature it will find interactions (young male Armenians are more likely to crash than young males in general). If the importance of such a feature is not significant enough to distinguish it from noise, then regularization may automatically remove it.
Even if we observe that 99 out 100 Armenians crash their cars, and you decline someone a loan, because he/she is Armenian, you may just have discriminated against the 1 Armenian who is a safe driver. Young male drivers who drive safely have a worse time getting loans, because their group (the set of young male drivers) spoiled it for them. So their only hope of getting a loan is you adding more features (like nationality), to be able to distinguish them as safe drivers, not removing them and lumping them into the status quo.
The simulator is a great illustration of exactly what they did; the entirety of the work is generalizing that to arbitrary predictors (subject to a few conditions) and bounding the accuracy penalty.
Feel free to cite the theorem improving accuracy if you disagree.
[1] One idea I'm kicking around is the following. Banks/others are legally required to issue bad loans for fairness. I suspect there is a lot of money to be made hacking this, I just haven't figured out how yet.
But sometimes they loan to people they wouldn't normally loan to, just because that loan adds the kind of data they would otherwise not get, and they take a risk on default in exchange for more data. I'm assuming smaller loan amounts.
So I guess that is one way for lending institutions to say they are issuing "bad loans for fairness" while still benefiting from it in terms of getting more training data and making the model more robust.
To use your Armenian example, it could be that while being Armenian doesn't actually affect your driving, a "true" model could still end up being bad for Armenians if being Armenian is correlated with the things that actually do affect crash risk.
But what about this: we don't need to solve all social ills every single time. What if we let the algorithm correctly decide crash risk, and if we notice that it unduly impacts Armenians and that's not an outcome we want, we via a separate channel compensate the Armenians? That is, acknowledge the fact that Armenians may crash more, but give them a government subsidy to offset the higher premiums, and work to bring the premiums down (i.e.: fix the underlying issues that being Armenian is correlated with).
It's related to something I've been thinking about lately with regard to minimum wage. I like the idea that everyone should have a livable income, but tying the implementation to businesses that have low wage jobs seems like mixing concerns. For example, my company doesn't have any minimum wage jobs, but shouldn't my company chip into this social ideal same as any other?
What if we let the businesses pay whatever the market will bear, and if we decide that as a social concern people should get more than that, we subsidize them from the government, which is wear this concern is coming from in the first place.
I'm just incredibly impressed that you came up with a way to make this sound even worse.. can you imagine the government giving subsidies to Armenian drivers with good records because insurance companies are overcharging them due to the statistical performance of their category?
A tax benefit for being an outlier from the mean? My god. Not only does that sound like a terrible idea, but intensely complicated to organize in the general case.
And what happens when an individual is in multiple categories? What if only gay Armenians who listen to disco are the at-risk driving category? How do we aggregate the tax subsidies, multiplicatively, additively, ..? This sounds like an administrative nightmare, I congratulate you on your evil ingenuity.
Lets say that I am male, and am young. Lets say that as a young male driver, I am likely to be a cause of an accident 20% of the time i.e. 1 in 5 young male drivers will cause an accident. Lets say as a young female driver, I am likely to be the cause of an accident 5% of the time.
So on the face of it, young male drivers are riskier, and should pay appropriately.
But lets say there was another measurable factor, such as a 'recklessness' score, which describes how recklessly you behave. And lets say that if you are 'reckless', you are likely to be the cause of an accident 90% of the time. And the maths works out that if you take the 'not reckless' group from men and women, they are equally likely to be the cause of an accident, i.e. there are more reckless men than women.
If you were to take the stats at face value, then a good proportion of people are being over and under-charged, because you are using a proxy measure rather than digging deeper for a root cause/don't have the data available. I think this is exactly what people are afraid of, businesses/people being lazy and making assumptions that are correlative, based on characteristics that one can't change.
This research is about what happens when combining the recklessness score with gender-based distributions provides more accurate results than recklessness alone. Many people consider this to be "unfair". This article shows how much profit you need to sacrifice to get fairness.
Even if that was true[1], where are you getting the data from? Choosing which types of data to use is just as important as the model.
[1] as others have already pointed out, it depends on the technique/etc
However with a lot of ML classes we are told to look for the simpler explanation which would be nationality in this case, in the absence of collection of the real reasons.
Not sure how that can be gotten around. I'll read the paper.