https://en.wikipedia.org/wiki/Money_laundering
So the idea is, you have some bias in some process (racial, religious, whatever). You set up some algorithm that relies on the bias to make predictions or classifications. Now you can say it's not you that's biased, it's just the algorithm. The algorithm is some process by which your bias is "made legitimate".
It's not that you set up the algorithm to rely on the bias - ML trains on data produced by a biased system and ends up building a model containing the same biases - you don't need to "set-up" anything - if you're fine with the existing biases you can "launder" them through ML.
https://www.technologyreview.com/2016/07/27/158634/how-vecto...
As another example, predictive policing tries to place police in places with higher crime rates. Those crime rates are determined by looking at past history of police reports and arrests. That past history has human bias already in it, with disproportionately higher arrest rates in places with racial minorities. The effect of the predictive policing is to justify overpolicing of minorities.
https://en.wikipedia.org/wiki/Predictive_policing#Criticisms
Ironically, this makes the original point nearly as well: we need to evaluate the hell out of machine learning systems to make sure that they’re doing what we think they are and that they’re not keying off something else instead, especially something biased. To date, the field has been...not great about this.
I'm being lazy, but the results for Gutenberg books you can check online at http://labs.statsbiblioteket.dk/dsc/
- man is to woman as doctor is to reprovingly (nurse is the first noun, on position 4) - woman is to man as doctor is to snodgrass (after a couple nonsense/rare words)
The most important thing that teaches us is that big corpora (bigger than PG) are essential for this method.
Man is to Woman as Doctor is to ___ gives 1) gynecologist 2) nurse 3) doctors 4) physician 5) pediatrician
Woman is to Man as Doctor is to ___ gives: 1) physician 2) doctors 3) surgeon 4) dentist 5) cardiologist
These are just generally near "Doctor" though: the ten nearest terms are physician, doctors, gynecologist, surgeon, dentist, pediatrician, pharmacist, neurologist, cardiologist, and nurse.
Some gender differences may persist (nurse is #2 for `woman`, but #68 for `man`, but it's also near `woman` generally and you could imagine it gets a bit of a boost from the verb ("to feed a baby") being attached exclusively to women too.
Anyway, my point is not that there's no bias (there certainly can be--seed GTP-3 with a prompt about Muslims) but that one should be wary of thinking they know what the model is doing.
Depending on the input “homemaker” may be a technically reasonable output.
Should we use such a model to suggest career paths to high school students? Or should we reevaluate the methodology?
https://www.slideshare.net/yuyomajadero/jobs-occupations-pro...
I'd like to add that there's also a danger in people trying too hard to avoid bias and losing important information. For example, a man who's a homemaker instead of a programmer is a less attractive partner for a woman. So such an occupation might harm both his quality of life and that of his partner. There is some useful information encoded in cultural bias. Even if that information turns out to be entirely socially constructed, it still has real harmful effects on real humans who go against it.
I know some people who, fairly strongly, believe that ML is the future of non-biased decision making, yes.
> a man who's a homemaker instead of a programmer is a less attractive partner for a woman
That's definitely going to be [CITATION NEEDED].
This is much to sweeping a generalization to have a place here. You might say a man who is a homemaker is less attractive than a programmer to you, but don't speak for everybody else.
Who are we to decide for the woman what she should feel about it? Is bias absolute or relative, objective or subjective? Is there a one true policy for dealing with bias or can there be one?
I've never done any, but my understanding is that machine learning is just correlation. It's good at figuring out "what", but not "why". Consider training an algorithm to recognize horses by feeding it millions of pictures of horses. Eventually, the algorithm "learns" what a horse is, but it's definition of a horse is based on the inputs it was given by a human.
So now, consider the scenario where the millions of pictures of horses were all brown. If you give the algorithm a picture of a white horse, it'll tell you it's not a horse. If you give it a picture of a brown donkey, it might think it's a horse because it's learned to put too much emphasis on the color brown.
If that algorithm becomes relied on to define a horse, "the system" will insist there are no white horses even though you can walk outside and see them plain as day.
Now, apply the same kind of idea and feed an algorithm mugshots of all criminals. It's going to develop the same bias and tell you that a black person is more likely to be a criminal than a white person. There's no nuance. The inputs used to train the AI were tainted by decades of systematic discrimination, but the AI doesn't know that.
Of course you could try to take that input bias into account, but the whole sales pitch of machine learning is that you feed it tons of data and it gives you an objective result. As far as I know, no one is trying to quantify, and correct, the biases in the inputs.
The phrase "money laundering for bias" means the machine learning algorithms are used to re-enforce incorrect opinions and assumptions because it gives the excuse that an "objective" computer used cold hard data to draw the same conclusion.
Machine learning is one of the scariest parts of tech right now because it's the equivalent of an extremely stupid person that only understands correlation and not causality and the systems being built are going to be making a lot of decisions at scale.
I would, but the people in charge won't. Some of my family members MUST chat with a crappy bot before they can get support from their mobile phone provider and that's in Canada where we pay an astronomical amount of money for our phones/plans.
A penny saved is a penny earned, even if it costs someone else a dollar.
No. Nobody training an ML system to detect criminals would train it only with pictures of criminals. And if you somehow did, it wouldn't determine that black people are more likely to be criminals, but that humans are more likely to be criminals than say ducks or fire engines.
Yes, ML models can end up reflecting prejudices in their training data, but this description is incorrect, reductionist and unhelpful.
The first is statistical bias - feeding algorithms training data that is somehow unrepresentative of the "real world" (or more specifically, the actual class of data for the intended use case). As an example, applying facial recognition to Caucasian faces when the model was solely trained on Chinese faces, you're going to have a bad time. The problem was that your data was "biased", because you actually wanted a model that recognizes "human faces", but you trained on the biased subset of "Chinese faces".
The second is the ethical/political notion of "bias" against individuals; more concretely, the idea that a society is "just" when people are judged as individuals, and not prejudiced by their gender/skin colour/etc. In this respect, when we say "we should not be biased against men", we really mean that "an individual man should not be treated any differently from a woman, even though men are overwhelmingly perpetrators (and victims) homicide".
The complication is that reality is inherently imbalanced/biased. Society can be chopped up into a lot of sub-views that skew towards particular demographics. Some are relatively harmless. "OnlyFans content creators" aren't 50-50 men-women, and men aren't charging the same as women either. Some are not - "murderers" are mostly men, black men are overrepresented in the "criminal" group, and so on.
This raises some obvious questions:
1) Why is this the case? Is this the result of systemic discrimination? Historical oppression? Innate preference? Cultural pressure to conform?
2) If you can answer (1), how does that influence your view of what a "just" society is? For example, do you consider it to be "unjust" to be wary of men (and men only) to protect yourself from random physical violence when you're out and about?
3) Does everyone share your view on what a "just" society is?
4) How do these answers dictate what you should "do" about it? As a voter? As an ML practitioner? As a CEO?
I'm not going to delve further into these questions, because they cause a lot of contention and deserve more time/consideration than I can justify right now in a HN post.
The main reason I decided to comment is that I've seen too much debate that tries to steamroll people into accepting conclusions without considering or answering these questions. Even worse, some people are actively trying to silence others who simply want to discuss these questions, rather than swallowing their conclusions uncritically.
(This is not levelled at you, by the way, your comment just presented an opportunity to lay out my thoughts.)
I don't think anyone can meaningfully discuss the (ethical) concept of "bias" without first laying out a very comprehensive perspective of "society" that touches on all of these points.
There was a joint paper from Google, Facebook (and possibly others) about 2 years ago that I thought handled this exceptionally well. The authors addressed many of these questions honestly and objectively, and most importantly, acknowledged the potential for disagreement.
Bias are subjective assumption you have of the world and people.
So this quote is saying that machine learning hides the source of bias.
Imagine a dataset about healthcare outcomes. You want to know whether airlifting a patient is good or bad, so you train a machine learning model to predict mortality given airlifting a patient versus not airlifting them. Turns out a lot more of the airlifted patients died, and the model picks up on that, deciding that airlifting is dangerous.
Obviously, we’re airlifting the patients who are in the most dire circumstances, and that’s why they die more, but maybe the model doesn’t have the context/circumstance variables to see that, or maybe it’s just regularized and thinks those context variables are noise. So the bias of “airlift = bad” gets stuck in the model, but then people defend it as mathematically precise so it can’t be biased like a person can. They say that the model just reflects reality, pretending the model or data is wrong is blindly rejecting reality for the sake of political correctness.
It’s worse with really human stuff like recidivism because the context variables that might be useful (like the patient’s dire circumstances) tend to be very human and complex and they’re unlikely to be captured in a simple form. Even if they were, interpreting them might require human level AI. By analogy, it is frequently impossible (or just too hard to be worth it) for current technology’s ML to tell from the data we feed it that the patients being airlifted were most near death to begin with.
So you end up with an algorithm biased against airlifting patients getting defended as mathematically bulletproof.
"Money laundering for bias" in this context can be translated as "Finding a plausible excuse/cover for a specific bias (your computer/model says the same!)".
You employ a machine learning algorithm, feed it some data that shows black people in poverty have high risk of default on loans and avoid giving it data that shows otherwise. Train your neural network hard.
Now when people ask about why you won’t don’t give out many loans to blacks, throw your hands up and say “I would but our advanced machine learning algorithms say these people are high risk, we’re not racist we’re just following the results.”