Cathy O’Neil on Weapons of Math Destruction
econtalk.org
econtalk.org
In the business of making credit decisions (specifically, who to extend credit to, and at what rates), preventing banks from using better information only harms people who would have been the subjects of erroneous decisions in their favour, and the ability of the bank to consistently make better judgements means they can also offer lower rates and fees to people they choose to extend credit to.
This is in contrast to sentencing, where you prefer (in theory, at least) to bias judgements in favour of leniency (in the spirit of Blackstone's formulation), and you might even to prefer to make mistakes, if making the right decisions alienates people who interact with the justice system.
I disagree. For instance, in the credit-card business a key factor in the ability of a bank to be successful is the ability to assess the likelihood of a person defaulting on their payments. There are a lot of factors that go into the formula to calculate that risk today, especially things like previous payment patterns and previous borrowing history. There are some factors that do NOT go into the formula like the borrower's race or the payment patterns of family members.
Now, I am sure (although I have not actually run the numbers) that an analysis would show that race and family-member credit scores are fairly strongly correlated with default rates. That means that a credit card company chose to use factors like race as part of the scoring decision the company would do better than another bank that didn't use those factors.
But we do not WANT to be using race or family wealth to decide credit decisions. Speaking as a banker, my company does not want to be using those criteria; speaking as a citizen, my country does not want banks to be using those criteria. Restricting what criteria are permissible for making credit decisions enables the banks that would refuse to use that data for ethical reasons to remain competitive. Restricting it allows us to craft a society where any citizen has an equal chance of success.... well, that may be a stretch but at least it is CLOSER to such a society than if we did not have such restrictions.
Sure, you can categorize this as "erroneous decisions in the favor of those who come from a poor or minority background", and I suppose you are technically correct. But that pretty phrasing doesn't make it ethical.
This isn't a black-and-white choice between including peripheral information into credit lending. When banks and lenders make a deliberate decision to ignore information that could allow them to be more accurate with lending, those costs are passed down to users - and although it may not hurt the typical HN user, marginalized credit-seekers can literally have their hopes of home-ownership or education denied because of a bank's choice to be ignorant but considerate. Neither choice is completely without consequences.
We could just let the market decide, and since we know the antiracist hypothesis is true, racist companies would miss on a huge amount of qualified labor and get beaten in a free market, forgetting for a moment that free markets are a myth and don't exist, and that the markets we do have aren't even close to efficient.
But we don't. We force companies to act a certain way even though they don't wish to because we are forcing ethical standards on them. This is, akin to your point at the expense of all of the workers not protected by the law. Every time we prevent a black person from being fired because he is black, it's at the expense of the white person who would take his job. The costs are passed down to white people. Neither choice is completely without consequences, yet we've made the one against racism.
In this situation, the question is between whether ensuring privacy is more ethical than ensuring access to capital - and this is almost entirely focused at minority groups. If we assume that banks can make more efficient and competitive lending transactions given more demographic information, then denying that information raises the ceiling on financial capital for those marginalized groups.
As of now, I don't have a definitive answer. Although I think it would be beneficial to examine how certain data impacts credit-lending and move from there. A lot of these concerns may be moot if the information in question isn't even relevant to credit lending.
That's true. If we didn't give credit to more than a few outstanding black folks, then some poor whites could get credit cards at lower rates.
The most important thing to realize here is that DISCRIMINATION CAN WORK. If everyone agrees that redheads are no good, and everyone is extra careful about lending money to redheads and extra reluctant to hire redheads and extra-strict when deciding how to prosecute and sentence redheads, then network effects will make it a self-fulfilling prophecy. A redhead will be more likely to get caught up in criminal proceedings, will be more likely to get fired (or not hired in the first place), and therefore will be more likely to default on their loans.
For the most part, society has decided that this is either a moral outrage or a case of tragedy of the commons. From the moral point of view we say it's just not ethical to discriminate against people based on race, sex, family, and such. From a purely utilitarian point of view we can say that discrimination can benefit one party at a cost born by all of society. As is usual with tragedy of the commons situations, we can repair the problem with regulation. Regardless of whether you prefer the moral approach or the utilitarian one, there is a pretty strong case to be made that it is GOOD (on a society-wide basis) to give better deals to some (those marginal whites) than others (those marginal blacks) by prohibiting the use of certain information in granting credit.
You're making it sound as if lenders don't apply criteria such as race out of high-minded civic duty and a commitment to ethics. They don't because laws were passed prohibiting them from doing so.
I work for such a lender. Yes, that is precisely what I am saying. Although I might call it "basic ethics" not "high-minded civic duty".
Typically now "you" (someone raising this issue, because I don't want to put words in your mouth) point out that corporations are not ethical and simply act to seek maximal profit. Then I point out that corporations have a culture, which may have values other than just profit, and that corporations are made up of individuals who ALSO have motivations other than just profit.
The company I work for also has resisted the urge to open accounts that customers didn't ask for (as recently revealed of Wells Fargo) and many other ethical deviations. I am sure we have our own sins, and should work to improve them. I agree that laws constraining corporate behavior are ONE good tool for managing corporate behavior, but it is not true that absent such laws there would be no other constraints.
And this is simply not true. Especially by race where IQ can vary 3 standard deviations or more.
Improving your ability to predict defaults is simply moving in the direction of that perfect algorithm. The inputs you use, whatever your emotional reaction to them, are irrelevant.
By not optimizing your predictions to the fullest extent possible, some people are worse off just as some others are better off. There is always such a trade-off. There is absolutely nothing special about the status quo. The idea that one side of this trade-off is somehow more ethical than the other is completely absurd.
But the world is not immutable. And if we silently allow companies to optimize for the unjust world of today, we make it harder to build a more just world for tomorrow.
Denying loans to people of color may be price optimal today, but that's true only to the degree that the world of today sucks. We should fix that, not optimize for it.
Or to put it more specifically, why should a middle-class white family have to pay higher rates on their loan to help some family of color get a loan, while the rich white family that buys the house outright doesn't?
Shouldn't we help people of color using a system that distributes the costs in a widespread and progressive way?
For example, imagine if income perfectly explained default rates. Then in that case, the race of the person wouldn't matter at all given income.
P(default | race, income) = P(default | income)
The only time this equality would be false is if race was being used as a proxy for another variable that isn't being collected.
Imagine If incomes were lower for certain races, in that case the algorithm would be biased in favour of those discriminated people, not against them.
P(default | discriminated race, income = X) < P(default | ~discriminated race, income = X)
this article goes into more detail https://www.chrisstucchio.com/blog/2016/alien_intelligences_...
Before you choose to make a tradeoff, you should know what it is that you're trading off. If anything, much stronger forces than this book are pulling in the other direction.
"In baseball, a team can’t create bad or misleading data to game the models of other teams in order to get an edge. But in the financial markets, parties to a model can and do. ...Silver gives four examples what he considers to be failed models at the end of his first chapter, all related to economics and finance. But each example is actually a success (for the insiders) if you look at a slightly larger picture and understand the incentives inside the system. ...Silver confuses cause and effect. We didn’t have a financial crisis because of a bad model or a few bad models. We had bad models because of a corrupt and criminally fraudulent financial system.
...Silver has an unswerving assumption, which he repeats several times, that the only goal of a modeler is to produce an accurate model."
Her other examples are things like pharmaceuticals research. http://www.nakedcapitalism.com/2012/12/cathy-oneil-why-nate-...
https://www.nyu.edu/projects/nissenbaum/papers/biasincompute...
I found this recently in an appendix of Susan Leigh Star's book "Standards and their stories" which outlines a syllabus for the teaching of "infrastructure studies". The reading list also discusses the consequences of systems of categorization such as the DSM and medical notions of gender.
(1) binary propositions instead of assessing functionality/dysfunctional on a continuum (i.e. you either have 'major depressive disorder' or you don't); and
(2) discrete and distinct 'diseases', instead of the cumulative effects of dysfunction/abnormality in multiple neural 'sub-systems'.
If you're interested in the DSM case, and the likely way forward (from my understanding, the convergence of psychiatry and neurobiology and increasingly accurate and affordable neuro-imaging techniques), the textbook "Stahl's Essential Pharmacology" is well worth a read.
Although not ONLY these types of topics, I should say. It's a great nerdy finance podcast.
That being said, I took issue with the discussion at the end of this episode regarding Google's ad targeting being used for 'bad' products like payday loans or for-profit online universities. Even though they appeared to be on opposite sides of the issue, neither addressed what seemed to me to be the core point. Which is, why is it better if rich people have to see ads for payday loans too? She seemed to be suggesting that Google's targeting somehow makes this problem worse by focusing these ads on vulnerable people. And while that may be true, if the thing is harmful when sprayed across un-targeted media, why is it so much worse when it's targeted? Just because it gives these people a better ROI on their spend? It just seems like a total red herring issue to me.
I totally agree that things like sentencing or policing using machine learning algorithms will strongly tend to reinforce the status quo. But ad-targeting just doesn't fit into that mould, IMO.
You can tell she's interested in preventing knowledge from how she handled her job to predict effectiveness of homelessness services - she actively decided to not use particular variables on the grounds they might show a result she was uncomfortable with. That isn't an issue of "being aware of the limitations of machine learning", it's intentional ignorance.
http://www.slate.com/articles/technology/future_tense/2016/0...
Sometimes correlations are self-perpetuating, and it's better to not know about them when making decisions.
Race is not a valid biological category; it is a social construct and a salient social category. & as such it remains the site of many unethical uses of power. At times measuring them can reify those injustices; at times it is essential in combating them.
You can always decline to use a variable after the fact; she decided she didn't want to know the effect in the first place. Suppose the effect had been the opposite of what she suspected?