Mathematician Says Big Data Is Causing a ‘Silent Financial Crisis’
time.com
time.com
Last year I gave a talk on this at 32C3 where I tried to elucidate some of the problems that can pop up when using data analysis to automate things that were previously done by humans:
https://www.youtube.com/watch?v=iRY9IceaVig
Personally I see a lot of potential in data analysis & artificial intelligence, but like all technologies they pose significant risks as well. What we really need therefore (IMHO) is to teach ethics to people that work with data, and establish organizations and methods that ensure data analysis isn't used to do harm (deliberately or not).
> What we really need therefore (IMHO) is to teach ethics to people
The issue with almost all organized teaching (as opposed to when you don't even know you are doing it, e.g. to your children, every day) is that it only targets conscious thinking. Do you already see the problem with teaching "ethics" (and empathy)? It's the wrong brain area. People will ace the tests but how they actually act won't change. Another big reason for why teaching ethics (the way it is usually done) is that much of behavior is environment-driven. So you can teach someone all you want, if you then place them in a cut-throat competitive "result-driven" environment you can see empathy and ethics quietly sneak out the back door.Instead of learning lessons, the players continue to 'print money out of thin air'[2] and deny any fault for the last collapse. The rationale seems to be, 2007 was an extreme culmination of events that wouldn't happen again in millions of years. Problem is, in a universe of infinite variables, an infinite number of combinations exist and could manifest anytime a given set of vulnerabilities and unanticipated outlier events present themselves concurrently.
[0] The (Mis)Behavior Of Markets 2004
http://www.goodreads.com/book/show/665134.The_Mis_Behavior_o...
[1] 2006 reminder
http://www.ft.com/cms/s/2/5372968a-ba82-11da-980d-0000779e23...
[2]The Quants
Great Monday Morning QB analysis on what went wrong & why it will be repeated(hint:greed +ability +hubris)
http://www.goodreads.com/book/show/7495395-the-quants
edit: fixed links
Studies seem to indicate that studying and teaching ethics does not make you more inclined towards ethical behavior.[1] On top of that, from personal experience in business ethics courses, "ethical" has an extremely suspicious equivalence with "the thing which limits the company's exposure to liability."
[1] http://www.faculty.ucr.edu/~eschwitz/SchwitzPapers/BehEth-14...
As opposed to what? Subjective decisions made by humans based on biases and personal preferences with no semblance of reason or fairness?
> for example, who lives in an area targeted by crime fighting algorithms that add more police to his neighborhood because of higher violent crime rates will necessarily be more likely to be targeted for any petty violation, which adds to a digital profile that could subsequently limit his credit, his job prospects, and so on.
So having more cops in a high-crime neighborhood is somehow a bad thing, because the presence of more police actually confirms the fact that there is more of a need of police in this neighborhood?
> this technology was actually siloing people into online gated communities where they no longer had to even acknowledge the existence of the poor,
So targeted advertising to the needs of people is bad because... I can't even make sense of this one.
If we were to enforce all laws we would need massively more police in every neighborhood. When police are sent to flood a poor neighborhood they end up ticketing poor residents for a bunch of "violations" that they would also find if they went to the rich neighborhoods. But if they ever actually did issue all those ticky-tack tickets in the rich neighborhoods residents would be up in arms contacting their city council members and it would end quickly. In the poor neighborhoods nobody ever hears about it and the residents that can least afford it take the financial hits.
I'm not sure what data you have to support this. But check out violent crime rates. Larceny-thefts. Burglary/Property theft. Motor Vehicle Theft. Aggrevated assult. Robbery (mugging, stickups. etc).
These very popular crimes aren't so popular in nicer neighborhoods. Rich neighborhoods have their own problems, but let's not appeal to the spirit of lawlessness (cops are bad, rules are bad, unfairly arrested, etc.)
It makes sense for police to be where serious crime occurs.
Did you understand what you quoted? The author is saying that citations for petty offences (which practically everyone commits with some non-zero frequency, intentionally or not) have a higher chance of affecting future prospects of people in communities where there are more police. the Author isn't saying police are bad in general, rather that the negative effects of police presence are more significant in communities with higher police presence.
> So targeted advertising to the needs of people is bad because... I can't even make sense of this one.
The message isn't that hard to get, come on. The author is saying (assuming you think like them) that ignorance of suffering is a bad thing. You can say something like "why should I care about people I don't know", and I would understand and even agree to a large part with that statement, but this isn't James Joyce here, you can make sense of it.
Yes, but are there any positive effects of police presence? Does this outweight the negative effects?
> The author is saying (assuming you think like them) that ignorance of suffering is a bad thing
So seeing poorly targeted ads tells you that people are suffering? It is the role of corporations to inform you that there are people who have different needs than you?
Certainly I would agree that police have positive effects as well. Your second question really is the crux of the issue, and I can't answer that.
> So seeing poorly targeted ads tells you that people are suffering? It is the role of corporations to inform you that there are people who have different needs than you?
Suffering was perhaps poor word choice. Obviously it's not the responsibility of some faceless corporation that I be informed about what's going on in the world, but it would be nice if they didn't cut down incidental information the masses might receive so it fits in their little corner of the world.
This isn't Josef K in Kafka's Der Process. Most people don't have any trouble avoiding arrestable behavior.
Lawlessness is not a virtue.
>ignorance of suffering Ads for bad for-pay universitites are good to remind us of... suffering? And sorry, but even "awareness of suffering" isn't virtuous either.
It's what you DO about what you know.
I think we've pretty well established the past couple of years that there is no lower bound on what's considered "arrestable behavior" when it comes to minorities in the US.
Really? I wouldn't throw in the towel.
If there's corruption, that's still a crime. Why not figure out how to highlight it and stop it?
When you find bugs in your code do you throw up your hands and proclaim that it could be anywhere so it's too hard to find and fix?
Do you not fix the easy-to-see bugs because that's somehow unfair? Do you not fix the harder-to-find bugs becuase your biased?
You make the software operate according to your vision and purpose. You and go out of your way to ensure order and correctness.
No need to shrug off harder-to-solve social issues as impossible.
That sounds nice, but in practice, white collar crime, often committed by white people, has much lower incarceration rates than simple drug crimes. Even in Seattle, where its now legal to posses pot but not legal to smoke it publicly (yet this is routinely ignored), we found that black people were being incarcerated at a much higher rate than white people for public smoking of pot. An investigation found that there was one freaking cop who went around arresting tons of black people for this 'crime'.
This is the world we live in.
Edited to fix a typo
If you are having problems with drugs contact me and I'll give you details about a clinic in Seattle.
My rule of thumb for reasoning about things other than code: your code is like your code, and nothing else is quite like your code.
Ergo, don't expect that you can fix everything that's wrong with the Western civilisation just 'cause you can debug something you hacked together in a week or so.
For example, "breathing while black."
An interesting similar practical application was differential sentencing for drug dealing outside versus drug dealing inside. White and black Americans use illegal drugs at essentially the same rate, with white Americans pulling ahead in several categories historically. Suburban Americans more often deal drugs inside houses, while urban Americans are more likely to deal drugs outside than suburban Americans. Due in part to redlining practices, black people are more likely to live in urban areas and white people in suburban areas. So just send a bunch of cops to urban neighborhoods and you'll catch plenty of drug dealers and give them higher sentences, while letting all the white drug dealers go and giving them lower sentences if their behavior is egregious enough that you've got to arrest them.
This example doesn't even get into differential treatment for white-collar crime (scamming your investors) vs stealing meat from a grocery store. All those articles we saw on HN about scamming your employees and investors? MotionLoft, 1for.one? Those folks are "just doing business" and most commenters seemed to implicitly condone the scamming behavior -- sure it's unfortunate that it happens but you know business -- while no one here would condone stealing meat from a grocery store, and the meat-lifter would be more likely to get caught since there are cops at the doors of those grocery stores. Do we need cops in the startup scene if scams are happening, though?
The bad thing is not targeting "needs," as above, it's targeting perceived needs or exploiting differences in a way that contributes to injustice. I can see how you might not care about targeted advertising that allows you to ignore the lives of the poor -- that's fairly normal and in the US we really rely on it for political stability -- but ads that only show high-cost high-interest low-placement educational options for poor people push out ads from your local community college, targeted ads contribute to the political polarization we're seeing, targeted ads show higher-paying jobs to men than women regardless of qualifications. There is information asymmetry in a lot of these situations and in several ways: if the poor person trying to look up education doesn't already know about the low-cost legit CC and the information is buried after a page of for-profit ads, they may not understand all their options, and just as bad, we who earn a bit more may not know that's happening, because when we look up education to see if our teen could take multivariable calc nearby, we will get the CC right at the top of the page and the for-profit schools, not targeting us, won't show up! And then we can grumble with a clean conscience that poor people are so dumb they can't even look at the first search result...
Your answer for addressing indoor drug dealing is to not arrest outdoor drug dealing when you see it?
Allow stealing meat because somebody else got away with scamming?
We disallow theft because theft-free is a better way to live. We have decided that scamming your investors is also punishable as fraud.
Don't throw up your hands as if these are impossible issues. Live life and make your decisions to build your community into the places where you want to live.
That's the essence of the Security vs Liberty argument. Yes, on one hand it is a good thing to have more security in an area which needs it; but will you accept say, being filmed 24/7 for more security, with the tradeoff that this same data can be analysed by third parties for any reason whatsoever (ala Minority Report's shopping mall scene).
It is not an easy problem and we must decide for ourselves which tradeoffs we will accept, because it is seldom free technology with no strings attached.
However, for something like a credit approval algorithm, there is a valid issue of centralisation - if there are problems with the algorithm, a large number of people will be affected negatively by the same algorithm. While humans are biased and imperfect, you can find some who are less biased, or call their manager and see if they can justify themselves, etc. I can imagine situations where it's hard to escape or find recourse when you're faced with an algorithm that's put you in an unfavourable situation.
I'm hoping that, given enough data to train on and enough validation to make sure they are working optimally, this will not really be a big issue, but I can understand if some people are concerned about that aspect.
He'll also be less likely to be the victim of violent crime in an area with a proven record for violent crime. For the price of avoiding 'petty' crimes, it seems like a good trade-off.
What usually happens is that the police issue fines so frequently that they become considered an occupying force by the populace. In turn, the populace withholds information from the police, and both violent crimes and petty violations rise.
Police statistics make quantitative conclusions difficult, since they are so often fudged for political purposes. They rise when the police need more funding, and fall when they need political capital. Murder rates are generally considered the only reasonably accurate figures.
On the ground, however, the reality is very clear.
For the petty fines and eroded trust, this is from an advisory letter sent from the Justice Deparment to the judiciary of all fifty states:
“Individuals may confront escalating debt; face repeated, unnecessary incarceration for nonpayment despite posing no danger to the community; lose their jobs; and become trapped in cycles of poverty that can be nearly impossible to escape,” Gupta and Foster wrote. “Furthermore, in addition to being unlawful, to the extent that these practices are geared not toward addressing public safety, but rather toward raising revenue, they can cast doubt on the impartiality of the tribunal and erode trust between local governments and their constituents.”
http://mobile.nytimes.com/2012/06/29/nyregion/new-york-polic...
This seems to be what he is referring to.
(Also, the next bit of the quote, "Yet neighborhoods more likely to commit white collar crime aren’t targeted in this way", suggests a very loose grasp of how policing works. No, they're not, because patrolling the street in front of an office building isn't the least bit effective against white collar crime.)
Do you have any sources for this assertion? Have there been any studies?
Because nobody will make a study to verify that the sky is blue if there isn't at a minimum some dissenting opinion.
How much white-collar crime occurs on the street? Even violent crime off the street won't be helped much by street patrols.
I agree with you on principal but IMO society on Mars will be having this debate (I hope I'm wrong).
Police officers could be employed to periodically monitor a sample of the AI outputs and rate them. Humans can still be in the loop, just not at the lowest level of the loop. I'd have much more confidence in an AI that passes the tests and is independently monitored by humans, than in individuals taking the same decisions. An AI would be much more balanced, informed by more examples and tested to check its biases. Humans can always invoke discretion to justify their choices.
[1]: https://jeremykun.com/2015/07/13/what-does-it-mean-for-an-al...
Officers making themselves visible on the street and checking to make sure people are safe and finding out from them where dangers might come from, is different than directly treating everyone on the street like they're potential criminals.
They were found: http://www.military.com/daily-news/2015/04/03/army-seeks-to-...
When it comes to software, one policy to make things more transparent is open-sourcing. The question is opening what and to what degree. I feel strongly that organizations today could and should be much more open and are instead absurdly opaque.
It's our responsibility as a society to talk about transparency as key for trust and equality.
Consider when there are few companies within an industry, like rare metal mining. If they chose to use the same, open-source pricing model, and any changes to that pricing model would be reflected by all companies, the optimal pricing model would be to price at the monopoly level -- to raise the price to where it would be if there were no competition and a single firm in the marketplace.
In the case of lending, however, different firms have different levels of risk acceptance. Many firms (big banks) avoid high-risk borrowers. A few firms (pay-day lenders) are willing to lend to those borrowers, but only at a high interest rate. Somewhere in the middle is a collection of scrupulous lending opportunities.
Any regulation that would restrict the acceptable interest rates to a narrow band would reduce not just the (unscrupulous) loan sharks but also the banks lending to high-risk customers.
This would necessarily result in loss of borrowing ability to those labeled "high-risk".
Is this a desirable outcome? Maybe, maybe not. But it should certainly be discussed.
Regarding the borrower problem. This is a common fear of revealing one weaknesses by openness. I believe society will invent new alternatives to support the new needs. Again, it seems openness is positive overall, despite new challenges it presents.
"""
O’Neil also proposes updating existing civic rights oriented legislation to make it clear that it encompasses computerized algorithms. One thing that her book has already made quite clear – far from being coolly scientific, Big Data comes with all the biases of its creators. It’s time to stop pretending that people wielding the most numbers necessarily have the right answers.
Tongue and cheek response, but "wielding numbers"? Politicians don't misuse statistics to get what they want, people misuse statistics to get what we want. We shouldn't ban statistics and we shouldn't ignore the fact that both "sides" are equally susceptible.
Uhm, racial segregation is still a thing. Even if the signal of race is gone the geography of where a person lives could be signal enough to discriminate.
That sounds like a great idea, but it isn't as simple as it appears. Most systems won't have "race" directly encoded as a feature, but that is insufficient.
See slide 22 from [1], further in depth discussion discussion from [2] or just this feature (which many systems could automatically discover):
Feature6578 = Loc=EastOakland && Income<10k
I don't know what the solution to this is, but it's a pretty hard problem, and even the best intention are insufficient.Delip proposes in [1] to introduce a "fairness constraint" where the probability of a favourable outcome in the majority class is the same as in a protected class. This sounds sensible, but make optimising systems much harder.
[1] http://deliprao.com/archives/129
[2] https://www.chrisstucchio.com/blog/2016/alien_intelligences_...
The flip-side is that under a human-biased system, the first white person might get the mortgage based on being white, and the second black person might be denied. That would be reversed under an algorithmic system only considering legitimate factors, and surely that's progress?
The only think I think worth pointing out is that in the mortgage example the current location of a person isn't as major an influence as income. The fairness test here seems somewhat reasonable: a person on $10K/y who lives in Oakland should probably have the same outcome as someone on $10K/y from some very white area - all other things being equal.
The flip-side is that under a human-biased system, the first white person might get the mortgage based on being white, and the second black person might be denied. That would be reversed under an algorithmic system only considering legitimate factors, and surely that's progress?
Yes, I think so? Or at least it can be audited (in theory at least - although see sentencing guideline systems for the problems with this).
Race is often very predictive. Blacks don't perform worse simply because they have income < $10k. Blacks perform worse on many measures even if you hold all that stuff equal. Here are a few of the many studies reproducing this:
http://ftp.iza.org/dp8733.pdf http://www.mindingthecampus.org/2010/09/the_underperformance... https://randomcriticalanalysis.wordpress.com/2015/05/16/on-c... https://randomcriticalanalysis.wordpress.com/2015/11/22/on-t...
If you were right, then the machine learning algorithm would quickly determine that race doesn't matter. Redundant encoding, direct encoding, etc would be irrelevant - the algorithm would ignore it as noise.
The problem is that the algorithm is supposed to uncover hidden patterns that predict loan default. But then there is one specific hidden pattern that is socially and legally taboo to detect, yet also highly predictive.
(Incidentally, if you can solve this problem in lending, it's easily a unicorn startup.)
Many a time, algorithms infer race from secondary or tertiary variables. This problem is compounded by two more factors - (1) Unable to explain WHY an algorithm has reached a decision aka interpretability and (2) Focus of algorithm building to maximize accuracy regardless of other valid costs. There is some pretty interesting work going on this area.
The banks were holding higher tranches of the securitised mortgage pools. While those pools had traditional levels of defaults, the tranches had their expected value. As the defaults accelerated beyond historic levels, they plummeted in value. In both cases the same model said the instrument had/didn't have value. As these instruments are fair value accounted the fact that they were still bringing in cash didn't help any.
However, as the mortgage market started to go south, the default models underestimated the level of defaults, particularly in the equity release market. This underestimation was endemic in the industry at the time, from the banks to the regulators and ratings agencies.
There is, of course, a very strong argument to say that the ignorance was willful, everyone was on quite a nicer earner and reacted late to the evidence as it started appearing.
So there wasn't really a paradox. The banks were holding tranches while the expected defaults were within historic levels. Once it became clear that that wasn't a good predictor anymore, they tried to limit the effect but it turned out that so many participants were in the same boat that there weren't enough buyers.
I have to say I'm not entirely sure what this has to do with ad analytics but I'm not writing a book on the subject so maybe that's not too surprising.
I was in a situation like this in my own work and I regret not standing up for the right thing now...
Not to Godwin the thread ... but big data has been used for immoral purposes since well, big data [1].
The even larger danger, is that people who don't know better, will simply "train a model" and having no overt ill-will, will still create a biased algorithm.
[0] http://www.encyclopedia.chicagohistory.org/pages/1050.html
I discuss this in detail here, with numerical examples: https://www.chrisstucchio.com/blog/2016/alien_intelligences_...
The fact is that machine learning takes biased inputs and produces unbiased outputs, and this is a pretty normal occurrence. If I tell you that my algorithm to target advertisements treated `displayedInterest x isMobile` with 15% more weight than `displayedInterest x isDesktop` because it corrected for bias induced by latency on mobile connections, you'd think nothing of it.
If you train the algorithm to make the same decisions humans did (which is reasonably common when the humans are considered to be effective experts), it will make the same decisions humans did.
If you're comparing outcomes then you'll do better, but even then bias can affect outcomes - if members of group A was sold higher-interest mortgages than equivalent members of group B, then your dataset will show that group B has higher default rates and your algorithm will learn this, even though the difference was only there due to human bias.
I algorithmically trade the stock market. I don't train my model by asking whether it trades the way I would. The whole point is for it to do a better job than me! I train the model to maximize my returns [1] in backtests and simulations - if it trades differently than I do, so much the better!
[1] More precisely, volatility adjusted returns, and I also penalize model complexity to avoid overfitting.
In the idea case you would, sure. In practice mortgages take 25 years to deliver outcome data and the point isn't necessarily to get better outcomes, the mortgage provider would be content to get exactly the same outcomes (or even slightly worse outcomes) if it let them replace a large number of human employees with an automated system.
In most fields of human endeavour you can't backtest, because you don't have access to the outcomes of the decisions you didn't make. And the data often aren't as clear-cut or objective as market prices. The stock market is great but it's atypical of modern "big data" in many ways.
Mortgages are very specifically an area where you do. First of all, there is historical data.
Second of all, you can backtest well before 25 years. A couple of weeks ago I wrote a blog post explaining specifically how to make measurements in the presence of delayed reactions - I'm discussing a situation involving sensor networks, literally the same mathematics would work for mortgage default or refinance: https://www.chrisstucchio.com/blog/2016/delayed_reactions.ht...
Third, mortgage lenders can often backtest alternate decisions because there are pretty straightforward relationships between decisions. Some of them are even mechanistic, e.g. refinance_risk(interest_rate) and default_prob(interest_rate) are monotonically increasing.
You seem to think we are living in the exact specific dark age necessary to make your morality play poignant and relevant. That's about as silly as Star Trek landing on all sorts of alien planets, each one designed to highlight one specific social issue from USA 1966.
I'm willing to believe that theoretical solutions exist. I know from direct personal experience in the big data/lending industry that they are not always applied. If you are claiming that real-world lenders never train their models on human decisions then you are simply wrong.
Nice zinger - I hope you weren't saving it for too long. But I'm not going to get into witticisms or personal arguments.
I.e., if I know f(0) = 10, f(1) = 9, f(2) = 8, but I don't have data on f(3), it's worth running a bit of an experiment to see if f(3) is actually 7.
(I happen to know from experience I can't talk about that this analysis is regularly done, albeit keeping the Lucas Critique in mind.)
I don't claim that real world use of ML is perfect. I claim that "bias" (in the sense of making wrong decisions due to race) is a statistical problem and the solution is simply better algorithms rather than Cathy O'Neil's statistical nihilism.
For example, in the earliest days of the creation of the US national mortgage market, conforming to FHA standards meant following more or less explicit racial guidelines (not to mention car-promoting platting and urban planning guidelines). That amounted to a political decision to reward people who invested (and/or resided) in white suburbia with greater liquidity and hence price appreciation.
Focusing on the marginal profitability of individual loans ignore the bigger issue, which is that choosing the rules often encodes a much larger set of political preferences or goals. Then, the critique of using opaque algos is, for me, less that they hide the use of some verboten input variable like race or sex, but that by focusing on the algos we ignore the bigger question.
(it's a subclass of what I call the scalar fallacy, where any real world outcome being reduced to a scalar always collapses a multidimensional space into a single dimension, and that projection contains a huge set of silent preferences)
The whole benefit of using algorithms is that you make your goals explicit and non-opaque; you just look at the objective function. Opacity of the models isn't really a big deal - machine learning is fundamentally the mathematics of studying opaque models. It's the objective function that should be transparent.
If the objective function says "minimize defaults", that's your explicit goal. If the objective function says "maximize # of blacks getting loans", that's what you'll do. But it will do one or the other, unless you explicitly set your objective function to a x "minimize defaults" + b x "maximize blacks".
The tricky bit is that by doing this, we are explicitly and verifiably lowering the bar. Just witness how much flak people get for discussing this: https://techcrunch.com/2015/11/03/twitter-engineering-manage...
But unfortunately, an optimization system won't give you what you want unless you explicitly put what you want in the objective function, something our modern activists (probably including Cathy O'Neil) refuse to do.
That is not necessarily true, because it assumes that the biased inputs have no influence on the outputs.
Suppose the training data is a bank that has historically given higher-interest loans to minorities. Those loans will default more often, and so the algorithm could "correctly" conclude that minorities default more often.
Cause and effect are very hard for an algorithm to distinguish when it's the clustered input variables that are linked together for some opaque reason. The fact is, both statements in this example would be true: "higher interest is correlated with more defaults" and "minority borrowers are correlated with defaults." The algorithm doesn't know one from the other.
Sure, more training data could help the algorithm distinguish between the two, but there may not be enough training data that shows the inverse -- low-interest loans given to minorities -- if the historical data is biased enough.
But if the relationship is actually `defaultProb = a x interest_rate + b`, then the algorithm will reflect that. I discuss this case explicitly in the blog post I linked to in the section "What if black people don't perform as well?", and explicitly give an example of a linear model NOT doing what you claim it will do.
The fact is, both statements in this example would be true: "higher interest is correlated with more defaults" and "minority borrowers are correlated with defaults." The algorithm doesn't know one from the other.
No, it wouldn't because you'd also have cases of white people with high interest who defaulted. The model using race as a predictive variable would fail to fit those data points. I explicitly discuss this case in my blog post.
You are correct that in some cases you don't have enough training data. But in these cases, there is no particular reason for bias to have a particular sign. Again, as I discuss in the blog post I linked, why would Captain Kirk be biased against "black on left guy" vs "black on right guy"?
Insufficient training data causes variance, not bias, and could just as easily result in bias in favor of blacks. (In fact, I link to a number of examples where it does.)
Key point.
If one were to compare the two high-interest groups, a small sample for one group would be reflected in a large confidence interval around its mean, a greater overlap in the CIs of both group means, and less confidence that a real difference exists.
What's interesting is if at some point in future regulators extend "red line" to mean variable highly correlated to a protected class. Because if your using a variable and not checking how it correlates to a protected class and making financial based decisions on it. Couldn't you in affect transform the decision into a graph "women" on on side "men" on another(or what ever your favorite protected class is).
Maybe it's the cynic in me but someone who worked in a hedge fund, started yet another advertising tech company and now releases a book about the evils of "big data" abusing the poor... guess I'm biased.
I firmly believe in the power of data to help people make better decisions, but decision making is always a human process, even if much of the work is delegated to a machine.
Another example (from an old version of Picasa -- I'm not trying to pick on Google but I have firsthand experience of this) is when a friend of mine had photos of two different Asian people in his collection, but Picasa always recognized person B as being person A, presumably because it couldn't distinguish features unique to the faces of Asian people.
[1] http://www.cnet.com/news/google-apologizes-for-algorithm-mis...
Also, it's a bit strange that you suggest ML algos are racist against Asians. You do know that Asians are wildly overrepresented among people building such algos, right?
Consider the possibility that some classification problems are actually more difficult than others. As one example, darker objects are harder to distinguish than bright ones and there are very clear mathematical reasons for that. That's true even in fields where questions of "racism" don't make sense, e.g. in MRI (magnetic resonance imaging is non-optical, but contrast still matters): https://www.chrisstucchio.com/pubs/wf_segment.pdf
As a sibling comment stated, humans are deeply involved in all aspects of designing and interpreting an algorithm's results. To the extent that we are flawed, so will our algorithms be.
Another example I heard recently from a conference was that of analyzing blighted homes in areas. The underlying data was generated by human inspectors who turned out to have given passes to violations in more affluent areas and fines in less affluent areas. The fines compounded the problem of blight as the residents could not afford to pay them and ended up causing a feedback loop of blight increasing in neighborhoods identified as blighted.
I am frankly confused as to what is so controversial about the statement that algorithms created by humans are not free from human biases. I apologize if I have only added confusion rather than a productive analysis.
But that's not true. Inadequate training data will cause variance, not bias. This variance can have any direction.
In a linear prediction scenario (i.e. any go/no go decision process), which exactly what Cathy O'Neil is criticizing, the effect could just as easily be "more loans for black people" as "less loans for black people".
In a multidimensional predictor, it does directly decrease accuracy but that's all. "Black guy -> gorilla" is just as likely as "yummyfajitas -> tank".
Also, realistically, if we want to think about your black guy -> gorilla example, most likely if insufficient training data was the problem, then it was the number of gorillas that was too few. Google has vastly more black humans in their training set than gorillas - a quick google search suggests there are only 100,000 gorillas in the world, most of whom are probably unphotographed.
As a sibling comment stated, humans are deeply involved in all aspects of designing and interpreting an algorithm's results. To the extent that we are flawed, so will our algorithms be.
Sure, algorithms are flawed. But this does not imply that bias must be present proportionally to ours. Insofar as we do good statistics, human bias is mitigated and can even go away. Science works.
As a classical example, consider Morton's analysis of human skulls. Morton was a super racist guy trying to prove how white people are superior. But he did an unbiased analysis and accurately measured skull volume - filling a skull with buckshot is a pretty unbiased process.
http://journals.plos.org/plosbiology/article?id=10.1371/jour...
I am frankly confused as to what is so controversial about the statement that algorithms created by humans are not free from human biases.
What's controversial is that this is the idea of "original sin" but applied to science. And most of the proposed mechanisms are simply mathematically wrong. Also many of the examples cited by proponents of this original sin concept are also wrong.
Regarding big data and ML, I guess the question is then what kind of biases exist in data collection, something that seems dependent on the details and history each particular dataset.
http://www.theverge.com/2016/3/24/11297050/tay-microsoft-cha...
The only possible conclusion is that we should ban Time magazine. Or some such nonsense.