But the problem is larger than this single model, because this issue (or similar ones) are pervasive in the fields in which AI are being employed. If a neural net is helping a court hand down sentences, it is going to be trained on historical sentencing data, and will in turn reflect the biases present in that data. If you are still only seeing the one tree, you say "well we must correct for the historical bias," and absolve yourself of thinking of the larger problem. That forest problem is that we will always be feeding these algorithms biased inputs, unless we do the work to understand social biases and attempt to rectify them.
- court risk assessments
- loan risk assessment
- job application pre-screening
- law enforcement face recognition
- medical scan interpretation
...
So, in essence: justice, banking, jobs, policing and health care. Anything else? Seems like social bias can only affect a small proportion of the ML application domains.
This case is so pervasive that it gets taught in basically every university that teaches ML - and yet those students go on to repeat the exact problem they were educated against.
I guess you didn't watch Asian developments in face recognition closely in the last few years.
https://thegradient.pub/content/images/2020/06/image.png
PULSE is published as an art project, not suitable for face recognition or upscaling. It can only generate 'imaginary faces'.
They didn't even train the GAN they were using, it was borrowed from another paper. They probably used StyleGAN because it was a nice high quality generator and they invented a novel way to use GANs so they needed a toy model to showcase their algo.
Assume you want to train an AI to recognize shops or buildings for example for a Google car.
Well if you do it in the US with skyscraper, in Europe, in Africa, Middle-East or Asia, you will get completely different results and biases.
Also, I don't see how anyone has the resources to compensate such a social bias, unless they plan to do a shooting trip in hard to access locations and try to justify such as "Seriously officer, the reason I take all those photos is to make sure my machine learning model is unbiaised so I don't classify shops in your country as shacks."
And if we forget human activity, even looking at nature, the fauna and flora are different between areas their color, how sparse they are, etc.
Last example, applying sparsity to human activity, what is actually a town in some country might be classified as a settlement for example due to bias.
PS "Do the work" is a creepy phrase that is popping up everywhere in SocJus. I recommend describing the "work" that needs "doing" instead of just saying "the work" need be "done".
It has real world consequences that can drastically damage communities.
The current nightmare that is cyber security is caused by developers who do not understand that with great power comes great responsibility.
The culture of software “engineering” and development is fundamentally broken and based entirely on “its not my problem if someone else gets hurt, I just build things.”
There’s literally not a single other industry where this level of greed and willful neglect is acceptable.
Of course, this would be reflected in the AI community as well.
Blaming greed of engineers for security problems? Seriously?! That is like railing against EMTs as being responsible for the Coronavirus because they wanted to get rich without working hard. It conflates so many different areas and roles that it isn't even coherent logic and is nonsensical.
If you don’t understand that engineering is the fundamental bedrock of IT and the lack of security application during the engineering SDLC is a consistent failure then I don’t know what to tell you except maybe to gain more experience in software engineering and read more about data breaches.
Cyber security: https://en.m.wikipedia.org/wiki/Computer_security
But for arguments sake let’s just stick to network, application, mobile and IOT security.
- Why is MFA not on by default?
- How many devs have prod credentials published to Github.com right now?
- How many unsecured IOT devices are there?
- Why is email security such a dumpster fire?
- How many companies have an SSDLC?
- How many companies require separation of duties and approvals before a dev publishes another AWS bucket or some other unprotected data store to the web?
Go ahead and blame business but they don’t know jack about software engineering. We determine which corners to cut as opposed to stating to business that security is just a part of doing engineering. We’re the one’s who decide and thus it’s our responsibility despite denials of people like yourself.
Although while everyone intuitively understands the load bearing capacity of a footbridge, not many understand the capabilities of ML models. So perhaps better advertising the capabilities of AI would help inform decisions.
On the other hand, nothing an engineer or scientist can do will stop the Chinese government from using their technology to predict and suppress dissidents and minorities.
OT, but shall we stop throwing casual references to China as the example of everything bad that can happen?
On the other hand, China is in a conflict with the Uighur population of Xinjiang. As far as I understand it, there are elements both of cultural clash (China doesn't like religion in general, and the Uighurs are muslim) and the Uighurs' reaction to an influx of ethnically chinese population in the region. Anyway, Uighurs engaged in terrorism: this BBC article seems pretty balanced and lists a number of terror attacks by Uighurs as well as the repressive actions by China: https://www.bbc.com/news/world-asia-china-26414014
In this context, I think that China might be using AI not to "round up" Uighurs, but as an intelligence measure to prevent more terror attacks. Similar to how the US intelligence is (I have no doubts about it) profiling muslims and middle eastern immigrants- not because it has anything against those groups per se, but because it has reasons to believe terrorists might hide in their ranks.
I agree, but then it's telling how in an abstract argument about the potential misuse of technology seems natural to throw in an offhand accusation against a specific country. I am pretty worried by how quickly China has become the new boogeyman- everyone thinks it's perfectly reasonable to display anger towards a huge country that only a few years ago was seen, despite its obvious issues and shortcomings, as successful and dynamic.
And I remember how it started: when a mainstream, trusted news outlet reported about the existence of "spy chips" in hardware sold by Chinese companies to the US "according to extensive interviews with government and corporate sources". Which later turned out to be fake news. It gives pause for thought.
https://www.bloomberg.com/news/features/2018-10-04/the-big-h...
That is exactly what I claimed they were potentially using it for.
I acknowledge that I am very much biased against China as I live in a country within it's sphere of influence and have friends in HK.
But in any discussion of using AI for unethical purposes, China is the ur-example, as Nazi Germany is to facism: an authoritarian government with a history of tech surveillance, censorship and media control, and minority oppression, and the tech to back it up.
If they don't like that, maybe they should stop using technology to oppress their own citizens.
Abstracting away the ethical issues of your (elite) employment is a personal choice that is anything but objectively neutral.
Does this apply to doctors taking care a murderer, rapist, [insert felony here]?
Should this apply to the doctor of Yann LeCun?
Should these "personal choices" be valid, allowing the doctors to judge YOU, negating a cure?
You can extend 'doctor' to other professions too, like journalists.
PS: all those professions require to be objectively neutral.
No, this is only technically correct but actually wrong. In the example, they did in fact fail to build the AI, as is their job. "Recognize white faces" is a lame research goal, "recognize human faces" is the real thing. So if somebody builds an AI system that fails on out-of-sample data, then says that had they tried to do it properly they would have succeeded, that's a pretty lame excuse for a poor AI system. They didn't even do their job in the narrow sense that you're using, you don't even need to consider the "social biases" or whatever, it's just a system that didn't work. In fact, many years into this research program, "focussing on his task and correcting it" (working on all human faces) is still not done given the current performance of these systems, but they are quite sure it can be done if they tried.
As someone put on Twitter : we should be rethinking the meta learning algorithm of the ML academic sphere, and a leader like Yann is the kind of person who should be spearheading that.
For someone like Yann specifically? Publicly state that the ML scene is optimizing in a myopic way, and invest in doing so less myopically. For example, I think the translation space has a clear goal and the right goal and is making strides in improving language models in many useful ways.
Ultimately, if there aren't ethical and useful ways to apply facial recognition, leaders should be steering people away from those research topics.
The researcher's job is to make progress towards the research goal. Progress that doesn't solve the problem is still progress. Nobody's saying that the PULSE authors have solved their research goal of generating white faces. It's why papers have a "discussion and future work" section, which touched on the issue in the first version of their paper. It's why in their revised paper, they added a full discussion of bias and the issues in the current model. They did their job, but didn't solve the whole problem, which is too high a bar for any researcher. Science is incremental.
That's how I interpreted LeCun's distinction between the researcher's job and the product builder's job.
1) The model uses biased data when it could have used unbiased data.
2) The model uses biased data when no unbiased data exists and is incapable of correcting for these discrepancies.
The first case is most clearly and engineering/implementation issue, the second is obviously not. Biased data is a known failure case of ML, its the responsibility of researchers to design for it.
I think he has some responsibility to at least acknowledge the complexity of this issue in such cases. Not speaking to the public is also a option if he doesn't like the leader role, so him deleting twitter is a completely ok thing in my eyes.
"'Once the rockets are up, who cares where they come down? That's not my department' says Wernher von Braun"
The means calculated using much less diverse series of historical cohorts are still being fed back in to present IQ calculations, which thus retain a persistent corresponding "echo" of the biases in favor of those early test results.
The echo is gradually fading, but some say a clean reboot of the baselines is the right way to resolve historical sample biases which still skew IQ testing away from an accurate modeling of the diversity in present test populations.
But I just wanted to say that in the instance of a neural net that grants bail or not, there's an issue beyond either biased data or the neural reproducing previously biased opinions.
The modern notion of fairness implies that the individual be judged based on their personal merits rather than things roughly correlated with their surface characteristics. Being black is correlated with being poor and being poor is correlated with being a criminal and various other bad behaviors. But that doesn't mean it's fair to punish a given individual, who is only responsible for themselves, for such surface characteristic.
Which is to say that if an AI crunches the numbers in a objective fashion with the aim to make decisions based on various correlations, that can fundamentally problematic regardless of the bias of the original data or people.
For a scientist doing a ML system to reconstruct pixelated faces, trained with white faces, why is he/she responsible for "insert larger problem" outside of her/his field?
Do he/she has to also care about the brutal Tantalum Wars because there are some in the electronics they use?
As far as I can see, ML is pretty new, and there is a lot of room for improvement. And I think people need to stop thinking the entire world is racist, or doesn't care or don't want to improve things. Changes take time, and won't happen today, or tomorrow or the next decade.
There is a constant, furious need to point fingers, followed by fear of telling or not telling just the right words on the right order to the right audience. This is not how we are going to solve anything.
I don't think the researcher was being defensive just because it was about racism.
Many see these systems as a way to implement stop-and-frisk (quite racist outcomes) and worse across the country using AI/ML as cover. The higher error rate amongst darker skin people gives LE a new excuse to harass innocent historically disenfranchised people.
I expect this tech will be widely used long before the accuracy problem that affects ~60% of the global population is fixed.
By the way, ML is not nearly as new as you seem to think, and given the amount of resources poured into it by FAANG-type companies recently, even five years is a lot of time for ML nowadays.
I'm sorry but you can't know this. Maybe it was the real objective, or a first step for something bigger, or a drunk Saturday night project.
In any case is research, and is very valid. Unless you are the chief at their lab, I think you don't get to tell what to do or not to do and/or the scope of their job. Anything else is a conjecture.
> By the way, ML is not nearly as new as you seem to think
Well, the results are there, together with the polemic it generates TODAY about white and black faces. Tell me if ML and its adoption, generally accepted or not, is mature yet.
Because we're all responsible for how the tools we built are used and what they enable, and how they affect society at large. That's what ethics is about, something which seems to be absent in the education of the modern citizen and in particular engineer or scientist, who is supposed to come to work, program things and not think too much about the impact their products have on the world.
One, if datasets are biased --- if you build your system to only work on white males --- then it may have suboptimal results for other groups. This is a common problem: you use your company's faces, or college students enrolling in data gathering exercises, etc, who are not representative of the population at large. We can fix this by being careful about dataset bias.
But the second issue hits right at the heart of a major societal problem/debate. When we use AI to make decisions about people, will the system become racist--- even with representative datasets? If you train something to predict, say, odds to default on a loan-- will it figure out things that correlate to race and be making mostly racial decisions? Different races do have different default rates, but we've decided as a society that it is unfair to use race to determine an individual's probability of default. But if we choose things that are correlates of both default rates and race, when is that fair and when is it just veiled racism (redlining)? What things are measuring a causal relationship and what things are just racism in disguise?
This second problem is much worse with ML, because we have the ability to accept a whole bunch more things into our models and explaining the rationale of why decision are made is much harder.
And of course, the first problem-- both bias in datasets during use and biases during research and development -- makes the second problem worse. It can't even necessarily be addressed by broadening the dataset and retraining: if, in this case, you do your research and training with just white faces, and report positive results, it may not generalize to work as well for everyone with a broader dataset.
E.g. for a non-ML example -- redlining is theoretically "blind" to race, but makes extremely racist decisions.
ML will figure out both but can't explain the hidden variables: it'll figure out behaviors that directly indicate higher risk, and behaviors that indicate you're a minority that tends to be a higher risk.
This is a hard problem.
Many people would say that identifying people who spend a lot of money at bars or casinos as a credit risk isn't racist, even if it happens to pick up more minorities. The mechanism for the credit risk and behavior seems tightly correlated.
Many people would also say that identifying people who spend some money at clothing retailers that market to minorities as a credit risk is racist. Here, the relationship seems like a hidden way to spot minorities, who happen to be more of a credit risk.
When banks drew red lines around all the minority neighborhoods and didn't lend to people there, because they couldn't consider race anymore, most thought this behavior unacceptable, even though it did genuinely reduce credit risk for banks. ML can't explain its rationale, and very often does the exact same thing inside an opaque box.
My tendency is to think that since it is optimized to maximize profit, the more features about a person we get in the data, the less "racist" the system becomes under this definition. Yes it's more possible to "hide" racism in a complex method using tons of features, but increasingly less likely it would happen. If we can use income and debt and whatever other things to make either a good classifier that ignores race, or an inferior one that sneakily identifies race and bases the decision on that, mathematical optimization of the model should result in the former. ML can be trusted more than humans in this situation, not less.
Of course there's still the issue of the data being biased, which is where it all started.
I'm guessing it's illegal to discriminate based on race, gender or ability no matter how much it extra it costs a business, whether its deliberate or not.
This makes "buys fashion for minorities" quite an accurate predictor of high credit risk, but also a profoundly racist one.
Strawman. The question is whether predictions are improved by using race (or direct proxies for race), not whether it'd be wise to use proxies for race alone.
Loan default rates are higher for disadvantaged minorities, even after controlling for many, many other variables (income, neighborhood, level of education, etc). Therefore, using race (or inferring race) improves prediction quality, but is ethically dubious.
Can we skip this kind of snippy arguing please? Anyway I take it you agree the statement is true but don't want to.
So now you're saying those great minority loan customers are impossible to identify? I think you just need to figure out what information is still missing. What's the effect size at this point anyway?
> Can we skip this kind of snippy arguing please?
If you don't start off by willfully mischaracterizing your opponent's argument in order to be able to more easily refute it, I think you'll find that they accuse you of this less.
I think you're just willfully missing the point.
Ideally, you'd just lend money to the people who will pay you back. Unfortunately, we can't predict this perfectly. Adding race, or proxies for race, to the things you consider improves your predictions somewhat.
And what have I claimed to "easily refute" exactly? More like I ran with the definition, and considered how to address the problem as stated. I said more features were needed, I didn't say enough features were currently used. You keep pointing to a dichotomy between perfect and flawed, while I was talking about relative improvements. There isn't even necessarily a disagreement there.
It is mainly deep learning for which this is difficult. Which I'm sure is the method in the twitter discussion, but that is a different kind of problem where interpretability doesn't have a clear use.
They may be less inscrutable than deep learning, but can easily still draw a box around most of the black people in a clever, non-obvious way.
https://consumerist.com/2008/12/22/amex-lowers-your-credit-l...
Capital One has admitted doing the same, as well as a number of other lenders.
There's been analyses done since, that show that when controlled for all other variables in credit reports, credit issuers issue less credit to blacks.. e.g. Cohen-Cole, 2011, Credit Card Redlining, Review of Economics and Statistics
You asked for a source.
I provided a source of the same. Yes, things like purchase histories are used as proxies for race and used to deny credit to mostly people of color.
Then you go here:
> Credit Card applications do not ask for race.
which was never asserted.
> I suspect it’s more likely based on zip code or similar.
Yes, zip code can be another proxy for race.
As for the latter, no, it’s not a proxy for race. It might be highly correlated with race, but that is very different. I don’t know why everyone does these gymnastics to find a racist angle to everything, but it’s nonsense and it’s very harmful to spread that narrative.
Nobody is searching for black zipcodes and using that as an input to deny credit. The goal is to increase revenue by giving as much credit as possible with the least risk possible, so it wouldn’t even make sense.
Scoring based on where you shop is still behavioral analysis. I think it’s a bad policy, but it isn’t targeting people of color. Your zip code affects insurance rates too, and that’s based on claims. It doesn’t cost more to insure your car in south side chicago than beverly hills because there are black people there, it costs more because there are more claims. The same effect is seen in “white” zip codes that are in areas with severe winters.
Black people default on loans more, adjusted for income, education status, etc. But we've decided it's socially un-okay to ask people if they're black and adjust how much we're willing to loan based on the answer.
But instead, we measure all kinds of ancillary things, many of which are highly correlated with being black and not obviously correlated to ability to repay, and dump them into a model, where we come up with weights that make basically the same decisions as if we'd asked if they were black. Is this is fundamentally better somehow?
Mind you, I don't have a wonderful answer as to what to do: it's important to accurately price credit risk. But it's also important for society to treat minorities equitably, especially when inequitable treatment might reinforce the very problems/reasons why some minorities are worse credit risks.
It might be better to explicitly have the "is African American?" question in the model, because then regulators, policymakers, the public could perhaps eventually know the exact contribution of this factor... while approaching it obliquely makes the effect far less clear.
For example, from the outset, would you object to an AI that made decisions on how harshly to sentence someone based on age & number of prior crimes?
What about hiring or admitting people to college based on the results of an IQ test?
Yeah, those end up being badly biased, this first has been studied and the second is the reason for dropping SAT/ACT scores--every mental ability test correlates with IQ.
The more interesting thing, IMHO, is that the opposite is also helpful. For example, if you help all poor people equally, you help even the playing field by disproportionately helping out all disadvantaged groups.
Education is different. IQ shouldn’t be relevant but competency and aptitude should and in addition we should recognize some people are better off going to vocational school where they may do better economically during their lifetimes.
Yes. This is an overly simplistic system for which it would be hard to justify the cost of developing.
>What about hiring or admitting people to college based on the results of an IQ test?
Yes. For the same reason.
For the second, any mental ability test correlates with IQ, so you end up with a different set of difficulties. There have been attempts to, e.g. correct for cultural bias in the tests, but these actually made the problems worse.
I'm not presently aware of anything that makes that situation better.
If you're looking at things that correlate, you've got a lot of choices. Saying that your data set is representative, your algorithms are correct, doesn't mean your choices of what correlations to use are objectively or morally correct.
Suppose that you have a system that accurately tells you men have higher car insurance claims on average than women.
The same system might also be able to tell you people with low credit scores have higher claims on average than people with high scores.
Is it right or ok to use the first correlation and ignore the second? Or vice versa? Or maybe both of them are unacceptable? There's nothing objectively inherent in the correlations that tells you it's ok to charge men who are good credit risks the same as men who are bad credit risks. Or optimal economically! Those are independent and difficult questions.
You're worse off lending in traditionally black neighborhoods; the question is, is it ethical to make a map of all of these traditionally black neighborhoods and refuse to lend there? A model that takes this into account would be biased in the same way the actual world is, but many would consider it deeply racist and unfair.
ML has problems with all of these at times. The particular case here with the Obama picture is a particularly vivid illustration of one that can be used to bring attention to the overall problems.
I feel like you've missed at least half of what I was trying to communicate, because you're still presenting arbitrary discrimination as reality-based.
Discrimination as you describe is not only unfair, but it costs the lender money, because they are not identifying the true risks of individuals within the group. Because it's suboptimal, it can't be called an unique, objective reflection of the world.
All correlations are inherently imperfect and therefore unfair; if they weren't you'd have an definite causal relationship.
When you use one correlation, or set of correlations, maybe you have completely missed something that would be much better, not just from a moral perspective, but a business perspective.
If you discriminate based on gender, and haven't considered age (hypothetically it being legal), maybe the latter would be a much better proxy for the real causal factors. Your gender model can be correct in itself, and yet from a completely amoral perspective be terrible and uncompetitive. The model is not "the way the world is" just because it's technically correct in isolation.
No model is perfect. Odds are you can better predict risk by (directly or indirectly) attaching bonuses or penalties to given races, all else being equal.
e.g. Say, irrespective of all other measured quantities, people of race A are more likely to default because of other systemic racism against them. As a lender, it'd be completely rational to consider this and make "better" choices. And if you're not allowed to measure race, it'd be completely rational to find other variables that don't have an obvious causal relationship to credit risk, but predict race and thus have some information about credit risk. This is the exact kind of thing ML does.
Then, in turn, this becomes self-reinforcing. Because other institutions discriminate against race A, the risks going forward of dealing with race A increase...
But some are definitely better than others.
It sounds like you're looking at things purely from the point of view of getting the correct average for a group. But whether or not you get the correct average doesn't tell you if you're using a good enough or the best available model completely apart from fairness or justice.
If you do some type of testing and you know 5% of the tests should come back positive, is there a difference between reporting 5% at random and actually doing the tests? Of course!
No, I'm looking at things from the point of view of making a model that fits well to the original dataset and then is verified in the actual accuracy it makes over time.
If adding race-- or inferring race-- makes the model substantially better in predicting outcomes, is it right to do so? Credit default risk is correlated to race, even controlling for other variables. Hence, using race would help you make more accurate predictions.
This entire discussion is about ML models, which can often be fairly described as very fancy curve fitting.
> And when inserting a fitted curve into a feedback system you're very likely to just perpetuate the problem you're looking to eliminate.
This is not a problem for the individual credit issuer-- they're not looking to eliminate the problem. They've avoided some credit risk by taking race into account. They've improved their expected value, even if society is stuck with the cost of the problem getting worse.
And I repeat curve fitting isn't modeling. Because in that case it's not a model it's a prescribed outcome.
That is, what's your loss function and how does it prevent race from being considered in the credit decision?
- a first network take the input data and return a representation A (like an embedding vector): let's called it the "censor network"
- a second network take this embedding A as input and is trained to predict the class that should be censored (for example the gender of a person) : the "discriminator network"
- a third network take the same embedding A as input and is trained to predict the real task of interest (for example the probability of credit default) : the "predictor network"
The idea is that, by training the censor to make the discriminator fail (predict the wrong class) while making the predictor work, it will force the censor to learn a transformation of the input data that keeps the task related information in the embedding A, but removes the information correlated to the "censored class" (and that could be used to discriminate).
Here's a reference about this kind of methods, but it's still an active domain of study in ML and there are many papers that followed this one: https://arxiv.org/pdf/1801.07593.pdf
If neural nets provide a way maintain plausible deniability while breaking the rules, it's going to get used.
Imagine you have two groups of people that have different means and distributions of “input” (could be intelligence, wealth, interests, ...), and this input is somehow correlated with some output (insurance rate, criminality, university acceptance rate). Both input and output are one-dimensional.
Then you only have two choices. (1) You “unify” the distribution and use it as joint input, resulting in one group (the one with the higher mean) getting “more” - many people think this is unfair. (2) Or you can subtract the means first and so make the distributions as equal as possible, which can result in a person from the “more” group having worse outcome than an equivalent person from the “less” group (if the means are close together and the distributions are wide, there will be many such people) - many people would consider this unfair.
There are plenty of examples, current and historical, of society picking one or the other, for different groups and topics: wealth/success/income (old vs young, Jews&Asians vs other Americans, men vs women, whites vs blacks), university acceptance (Jews used to be discriminated against (look up “Jewish problems”), now Asians are, men vs women [1], “positive” discrimination of blacks) - just a few. In many examples historically, we consider (1) fundamentally better than (2) - e.g. discrimination against Jews at universities - but it seems it takes the society a long time to reach this conclusion in each instance.
[1] https://en.wikipedia.org/wiki/Simpson%27s_paradox#UC_Berkele...
Approach #1 is to just go to maximize expected value, using all variables, including race.
Approach #2 is to use race to adjust the distributions as you suggest. Call this "affirmative action."
The inbetween approach #1.5, that we usually follow, is to make a value judgment about whether any individual item is OK to use.
Income? Of course it's reasonable to use income to determine whether to lend-- even if it's correlated with race. Race? No, that seems wrong. Location? Sure--- wait, you've drawn a box around all the black neighborhoods, nope! Buys certain products? Sure-- wait, you don't mean "buys products that blacks stereotypically like," do you?
#1.5 is already iffy, but becomes completely untenable once model complexity gets high and we use approaches that do not do well at explaining their decision: it basically becomes #1.
A major problem with #1 is it becomes self-perpetuating and reinforcing: if everyone is going to be biased against you, you're going to do worse, and hence the models are validated/more models are created with these assumptions. Everyone using #1 may improve their own expected value at the expense of a worse outcome for society as a whole.
Do you have any proof of that? The world is improving on all levels, even the income gap between blacks and whites has been narrowing [1].
We're not talking about 100% segregated populations like during slavery or patriarchy, the means between the groups are very close and there's a huge overlap between the distributions. It's obvious that the vast majority of outcome isn't determined by race (otherwise you couldn't have white homeless addicts and black presidents) or any such "group" characteristic but instead by "individual" characteristics (skill, intelligence, drive, effort, hereditary wealth, chance, ...), so the evidence for this "self-perpetuating" unfairness is very weak.
[1] https://www.pewsocialtrends.org/2018/07/12/income-inequality... but you'll have to manually calculate the percentages
I'm completely overwhelmed by your evidence of median black income moving from 59% of whites in 1970 to 65% in 2016, and appearing to not change substantively in the last 15 years of the dataset. If the overall trend continues, we'll reach parity in another 269 years.
And it hardly has been happening in a world where considering race and proxies of race is considered "OK".
This is an absolutely comical piece of evidence you advance to try and justify your view.
When more variables come up and get more complicated it seems like it would be impossible to say "You're based on race" versus "You're based on a multitude of factors, several of which correlate to race, and consequently you give more or better loans to race X over Y."
You find criteria that are not inherent to race.
Reluctantly using your twisted example, imagine the left handed are more likely to default on loans because the checks they send are illegible.
Then you could use the acceptance rate of written checks as your criteria instead of handedness.
Yeah this is tough, answering the question of the odds a person may default on a loan will probably reflect historical bias. In fact I'd argue that it should show this if the model is accurate.
The moral issue is what do you do with that information? Do you use that knowledge to improve the outcomes of the disenfranchised or do you reinforce the bias based on the information you have?
Both might be right, for their respective time scales, which is why it's so hard for either one to back down.
Many many ways to talk past one another. We need a small book on techniques to short circuit these situations and salvage conversations.
This is pretty much the issue at its core.
The real question is, if Yann is so smart, why is he arguing with people on Twitter?
No decision is completely safe and no system gets everything right in the first pass.
All we can ask for is a serious effort to understand risks sufficiently well to reasonably believe that the system is safe.
Certainly treating PULSE as a simple engineering problem and fix-as-you-go isn't so bad, is it?
This has already led to numerous troubles in the applications for loans, job evaluation, criminal matters and so on.
Which I'm arguing about elsewhere... but demanding that university research about upscaling a face to be perfect to publish is a little much.
> . But success isn't a success when the dataset isn't representative and datasets used by academics as a rule are not representative
Yup, so, academics should never research this stuff or publish? That certainly is an ... interesting ... recipe for progress. And research progress is one of those things that could help us ultimately address some of these issues, because as the original argument makes clear: improving the quality of datasets is not enough.
LeCun - ML is biased when datasets are biased. But unlike deploying to real world problems, I don't see any ethical obligation to use "unbiased" datasets for pure research or tinkering with models.
Gebru - This is wrong, this is hurtful to marginalized people and you need to listen to them. Watch my tutorial for an explanation.
Headlines from tutorial (that Gebru didn't even link herself): The CV community is largely homogenous and has very few black people. Here's a bunch of startups that purport to use CV to predict IQ, hiring, etc. Marginalized people don't work on these platforms and there's no legal vetting for fairness before these platforms are deployed. Facial analysis has the highest rate of inaccuracy (gender classification) on fair-skinned men (?) and dark-skinned women. Datasets are usually white/male. Most object detection models are biased towards Western concepts (e.g. marriage). Crash test dummies are representative of males, so women and children are overrepresented in car crash injuries. Nearest neighbor image search is a unfair because of automation bias and surveillance bias. China is using face detection for surveilling ethnic minorities. Amazon's face recognition sold to police had the same biases (greater difficulty distinguishing between black women).
Now, I largely agree with what Gebru said in the tutorial. So does LeCun, who explicitly agreed a number of times that biased datasets/models should never be used for deployed solutions.
But it's a huge leap in logic to then demand that every research dataset be "unbiased". It's like criticizing someone for using exclusively male Lego figures to storyboard a movie shoot, or if I attacked a Chinese researcher because they only used Chinese faces to train a generative model, and none the outputs looked anything like me.
That being said, I'm open to being convinced if she had made any effort to show/prove that "use of biased datasets in research" is correlated with "biased outcomes in real world production deployments". But she didn't, which is why her criticism of LeCun smacks of cheap point-scoring rather than genuine debate (a criticism I made of Twitter generally the last time this topic came up).
Do you believe that industry uses pre-made datasets that researchers promote in their work?
Would yes to the above two question be sufficient to show "use of biased datasets in research" is correlated with "biased outcomes in real world production deployments"?
I'm not saying she's wrong, I'm saying we don't know because she defaulted to the argument that "you should listen to minorities", not "here is the evidence".
What's more, every single example of injustice in her tutorial was an image recognition/classification problem - entirely different from the generative model that originally sparked the debate.
And the point being made isn't "biased input data isn't responsible for a biased model", it's "you need to look one step further and ask why is the input data biased and how that impacts the world."
Like I said, LeCun (and myself) largely agree with most of Gebru's points. But when LeCun went to great lengths to defend/explain his position, Gebru then literally responded with "I don't have time for this". Even before that, she didn't even bother to link the presentation she referred to (which again, didn't even directly address any of the points that Yann was making!).
It's this complete lack of good-faith engagement that prompted LeCun to quit Twitter, not the underlying discussion on ethics itself. LeCun clearly feels that Twitter is not the place for reasonable discussion, and after this episode, I'm inclined to agree.
> If you think there needs to be an argument made to justify that, that's fine, but I don't think it's valuable to assume that people come into a discussion without a basic understanding of the software engineering ecosystem.
I'm not saying that people don't use off-the-shelf models. I'm saying that I don't know if forcing research datasets to be "unbiased" will make any difference to real-world injustice. I don't even know if any of the examples of bias she gave in her tutorial (HireQ/Microsoft/etc) could be ascribed to the use of pretrained models. She could be right. I don't know. You probably don't either.
Going beyond the empirical question, she'd also need to explicitly argue why responsibility should lie at the feet of the researcher, not the engineer. Gebru did neither, which is why I say it's a huge leap of logic.
That, fundamentally, is LeCun's position. He completely agrees that warning labels should be put on these kind of models that say "Model has been trained on biased data and unsuitable for use in real-world applications where racial fairness is expected". In fact, this is exactly what the authors did.
> And the point being made isn't "biased input data isn't responsible for a biased model", it's "you need to look one step further and ask why is the input data biased and how that impacts the world."
And I'd argue you need to account for the context in which your model is deployed. If I'm using StyleGAN to synthesize facial textures for a video game, biased datasets and models are desirable, not something to be eliminated. I'll use the appropriately biased model depending on whether I want to generate white faces, Chinese faces, or black faces.
It's the use case that dictates the risk, hence why LeCun (and I) believe it's the engineer's responsibility, not the researcher.
In fact, they did so after Timnit brought up her objection, and did so by taking advantage of Timnit's research (the model card they added is a direct result of research Timnit was involved with: https://arxiv.org/abs/1810.03993).
The issue at hand was her lack of good-faith engagement on Twitter and the subsequent pile-on from the mob. LeCun is quitting Twitter, he's not quitting ethical debates.
Even now, it's not clear to me that this is the case. LeCun still hasn't actually acknowledged any of the broader ethical arguments Gebru made, either on twitter or on his followup posts on facebook.
In fact, he makes no references to her research anywhere (beyond the vaguest "I value the research you're doing" in his apology tweet). I found that rather suspicious, I still do.
Like, having read through the entire conversation, I have no confidence that Yann could explain any of Timnit's research if asked about it, even in broad strokes. That's really, really weird given everything that happened.
Well, to be fair, she didn't provide Yann with any ethical arguments in this instance.
But being equally fair, you're right, I can't speak for Yann. I can only speak for myself, and I personally agree with most of what I've read of Gebru's work (though not all).
But the issue at hand wasn't the research itself - it's the way the dialogue was conducted on Twitter, and Gebru accounted for herself very poorly.
Yes and no. A lot of what I'm saying is directly from the tutorial she repeatedly suggested he watch. Because I took the time to watch it, because that's the reasonable thing to do when someone suggests that you aren't fully informed on a subject and suggests a resource to improve your understanding.
Should cryptography researchers backdoor their own papers, because terrorists or pedophiles might use it?
Building ML systems that are difficult to misuse is underexplored, and Timnit is one of the relatively few researchers actively doing work in this area.
I'm intrigued by this. Any names (projects/people/protocols) come to mind?
I'd call them both examples of applied cryptography research. I think these projects compare very, very closely to applied ML research:
They come out of industry research labs, are worked on by respected experts, usually involving some academics, ultimately you end up with an artifact beyond just a paper that is useful for something and improves upon the status quo.
I'm admittedly not a total expert, so I don't know how far down to the level of crypto "primitives" this kind of work goes, but I believe there is some effort to pick primitives that are difficult to "mess up" (think "bad primes") and I know tink actively prevents you from making bad choices in the cases where you are forced to make a choice.
Even more broadly, just consider any tptacek (who I should clarify is *not a researcher, lest he correct me) post on pgp/gpg email, or people like Matt Green (http://mattsmith.de/pdfs/DevelopersAreNotTheEnemy.pdf).
Edit: Some poking around also brought up this person: https://yaseminacar.de/, who has some interesting papers on similar subjects.
That doesn't mean what you think it means.
LeCun - "ML is biased when datasets are biased. It's not the responsibility of researchers to ensure that ML is used responsibly in all cases, but the responsibility of the engineers of a particular implementation who need to use the correct models for the task at hand."
> I'm open to being convinced if she had made any effort to show/prove that "use of biased datasets in research" is correlated with "biased outcomes in real world production deployments".
what does that mean? is there anything that would auto-magically eliminate bias if it were introduced into research?
Let me rephrase. Yann is basically saying "bias is the engineer's responsibility, not the researcher's". Gebru (presumably) disagrees.
Now I might agree with Gebru if:
(a) she can show empirically that "researchers releasing biased datasets/models" is correlated with "real-world deployment of said datasets/models that leads to injustice"; and (b) she can make a convincing argument why one person (a researcher) should be responsible for the actions of another (an engineer).
But she didn't address either these points on Twitter. She actually didn't bother to address anything on Twitter. Her whole argument was "You're wrong, I'm tired of explaining, you need to listen to minorities, I'm not going to engage".
That's not reasoned discussion or debate. It's posturing and point-scoring. The Twitter format only serves to encourage this type of interaction, so Yann basically gave up on the whole platform.
but because she did not explicitly state those on twitter, or because of the way she brought it up, we need to invalidate her whole argument?
i mean... how odd!
No-one said anything that could be remotely interpreted as "her whole argument is invalid".
I'm sure he'd be more than happy to discuss with Gebru where he agrees and where he differs on his Facebook page or at a conference panel. I think he explicitly said this.
He's just decided that Twitter is not the platform for that kind of reasoned debate. Gebru's attitude in this instance - providing nothing more than "I'm tired of this, you need to listen to marginalized communities" - was the straw that broke the camel's back.
Of course she's right about all the things that everyone agrees on. Everyone in the conversation is right about most points, if you break down their stance into a list of points.
It's not that the points of disagreement invalidate the correct points, it's that having a bunch of correct points doesn't really tell you much about the thesis.
> "...I don't see any ethical obligation to use "unbiased" datasets for pure research or tinkering with models..."
I don't think his comment was addressing the larger ethical discussion at all. I didn't interpret it as a discussion of ethical responsibilities, rather a strictly technical, matter-of-fact statement about the nature of ML training.
Please don't interpret my comment as an attack on yours, it was more pointing out I interpreted his statement differently.
Twitter user: “ML researchers need to be more careful selecting their data so that they don't encode biases like this.”
YLC: “Not so much ML researchers but ML engineers. The consequences of bias are considerably more dire in a deployed product than in an academic paper.”
Perhaps I’m wrong. That’s the whole problem with Twitter though - you can’t convey much nuance or sophistication in 140 characters.
I doubt that anybody was accusing the creators of that upsampling model to have intentionally tweaked the model to default to white people. In that sense, YLC is attacking a straw man.
Given the publicity such problems have gained in the community, one would expect publishers of any model to verify it doesn't fail with the most obvious examples. Not doing so is negligent at best.
If we lack datasets to train AI that doesn't spectacularly fail any test for racial biases, we lack datasets for anything that could be considered fit for use. At that point it doesn't matter if the model is flawed, or if it's "just" the available data. Such models shouldn't be published, and we should instead invest in either better data, or come up with better methods of training.
(And, as a minor point, his idea that Senegal is representative of "Africa" as a whole is also... let's say "unfortunate")
Senegal was just an example he gave in a tweet. No need to be so petty on every word.
Also: darker? Darker than what? Are you taking “white” as your baseline?
I mean, if you’re going to throw stones about how “you don’t understand”, you could try having a rational point.
> Not all black people are Africans or have African heritage. Far from it, in fact.
Exactly. That's the "rational" point: the mistake was training the data on anything but the target population. Nobody is contenting that it could have been trained on something else: it wasn't and that fact is the problem.
> Also: darker? Darker than what? Are you taking “white” as your baseline
Darker than the white people the algorithm turns most inputs into. Have you seen the results?
Here's what the authors say about their own work:
> PULSE makes imaginary faces of people who do not exist, which should not be confused for real people. It will not help identify or reconstruct the original image.
Furthermore, this author appears to describe themselves as a more of an artist and hobbyist. This isn't someone making some kind of statement about ML research. This is someone playing with computerized art, and the entire social media tech mob dogpiles his work over what exactly?
The negligence here is on the part of everyone getting their hackles raised over nothing.
If they are in here I hope they don't misconstrue my linking this image. I think they explained themselves very well and I feel bad that they were thrust into the middle o f this controversy.
My point is that while diversity might help, you need a lot more than that to address this problem.
To be clear no one (or at least no one of note) has their hackles raised over this specific dataset. PULSE is fine for what it is, and no one criticized PULSE for having these results.
It is however a great demonstration for laypeople about how ML models aren't magic and don't always do what you, as a human, would expect. This is true irrespective of the source of that unexpected behavior.
That said, I believe the disclaimer you mention was added only after the recent twitter discussion.
Philosophically I find ML’s tendency to reflect the biases we bring to it very revealing. In some ways it shows us what we’ve built, the underlying biases that we’d rather argue about and ignore. When an algorithm selects longer sentences for black men than white men, we rightly see that as racism. Some say, “use better data, we’ll then be objective!” But I wonder if a better initial reaction is, “wow, look how badly out system has failed that it would produce such a bad dataset.” Never mind that maybe it’s not actually possible to be objective and that’s the point. Math doesn’t lie, maybe when we make a racist model we’re failing to see the mirror it’s holding up for us.
That is a high bar to never discuss “unfortunate” ideas on social media
I cannot believe that so many people fall for this. Journalists, laypeople, even HN users. (OK not that surprised regarding journalists.)
The problem is not racial bias in AI policing. The problem is AI policing! Racial (and any other) bias is trivial to remove - just subtract the mean! But that doesn’t make predictive policing a good idea.
Imagine this:
> Hello, mister/lady, our 100% unbiased system has automatically determined that you are a potential future criminal. You are under arrest and sentenced to death.
(Edit: and same could be said for almost any situation where you have a bureaucrat making decisions about people’s lives, and you try to automate this decision with AI.)
Hello, mister/lady, our 100% unbiased system has automatically determined that you have commited XYZ crime, would seem more appropriate.
What it isn't an excuse for are the goddamned negligent idiots who tried to use it in law enforcement without through testing. It would be akin to a surgeon dipping every sterile sharp unstruments in yogurt cultures and foregoing antibiotics before use to see if good bacteria makes infections less likely and could prevent antibiotic resistant bacteria. Even if the theory is valid and the goal worthwhile the needless risk taking shows a callous disregard for human life especially when done in an utterly halfassed way like that.
the concern is that, if interpreted as a general solution to the problem of racial bias in ML, its incomplete.
after that it kind of devolves and everyone is talking past one another and flaming, its twitter after all
I mean, he was clearly referring to the specific model.
For the sake of ethical R&D, it's counter-productive to build a hierarchy of investment into the problem. Admittedly, the responsibility of end-results can differ, but the consensus that this ethical work is important should ideally be universal. That said, there should not be some ivory tower where you wash your hands of the ethical problems of your field.
Ok, so what is the solution then? Does she have a concrete set of steps or goals to address the problems she sees? Is there a list somewhere of things that would appease her, and in her mind make ML fair? Honestly interested.
As for Timnit and LeCun: a user above notes that the epistemological frameworks they're using to analyze this problem are not aligned. I found that comment pretty eye-opening, honestly.
Shouldn't "I care about ethics" just be assumed? How many people do you know who would say "I don't care about ethics"?
Computer science programs around the country promise their young, bright-eyed undergrads that they'll change the world. Very few of them teach them the ethics they'll need to do that in a thoughtful way.
The assumption that science and technology is inherently ethical has unfortunately led to dangerous ideas over the past 150 years. Some, like geographic determinism and eugenics, have directly led to the suffering of millions of people. I hope we tread carefully and take action when we see harmful models. [0]
[0] https://twitter.com/SpringerNature/status/127547736519656652...
But I didn't say that science and technology are inherently ethical. I said that most people care about ethics. They might have different priorities or philosophy than you, but almost nobody commenting on a social issue is doing so because they want the immoral thing to happen. Right? So asking everyone to say "of course I care, of course" before everything they say is laborious.
I think that the underlying assumption is that science and technology are inherently neutral (which isn't true) and that neutrality would be inherently ethical (which also isn't true).
Because it identifies which team you are on. People only recognize the existence of a nuanced point if it's made from someone on their own side.
This is purely a power game in which some individuals have managed to blackmail everyone else into recognising their role or being cancelled. Now and then someone prominent needs to be attacked to remind the others what they risk if they don't fall in step.
This is an extremely unpleasant position to take, if your point of view is empowered within the status quo. It is much extra work for no discernible benefit to the researcher.
If your point of view is subject to disproportionate suffering under the status quo, then reinforcing current practices by implicitly enshrining them in input datasets will make improving your situation even harder.
As an example, consider the case of public school funding. In a hypothetical system where school resources are provided proportionally based on student success, good schools will thrive and bad schools will get worse. If someone points out this isn't fixing the problem, you can reverse the proportions -- this will cause good schools to suffer while bad schools will get additional funding (disincentivizing student success). In cases like this, it's not enough to just have a purely abstract set of metrics on which to base resource allocation: it will always require actual investigation of why good schools produce good results and why students do poorly in specific schools.
This is sort of obvious, of course, but it isn't being translated into terms that some researchers can or will grasp. It's never enough to just tell someone to 'debias the dataset,' as determining that bias is a monumentally difficult challenge that people have failed to achieve for many generations. A key factor in fact is the propensity for this kind of research to get deployed, today, by people who are not experts in a given domain of investigation, with possibly disastrous results in policymaking. These tools are not abstractions that require a team of experts to translate from research paper to the real world; ML researchers put out results that you can shove into your nearest computer and run.
What Timnit and others are getting at is that it requires thoughtful and careful assessment to get real value out of this sort of research. Ideally, in Timnit's assessment, the researchers themselves would put effort into identifying possible calamities and put as much effort into mitigating them as they do into publicizing the work itself.
Yann LeCun and other researchers simply do not believe this is their responsibility; all they want to focus on is the mathematics themselves. I'm sympathetic to this position but I also very much do understand the opposition. One of my favorite movies from childhood, "Real Genius," deals with this sort of issue as the main plot line.
The problem here is that many stakeholders really really* want to reduce the required intervention to a mechanistic increase/decrease of some sort, never mind that different stakeholders want opposite interventions.
No matter how "intelligently" you go about it (whether the intelligence is natural or artificial), this desire to simply boil down policy making to reallocating resources or incentives/disincentives is fundamentally lazy. It reminds me of the mania for diversified conglomerates and corporate management of the firm-as-portfolio (ie. reducing management's function to determining financial allocation between corporate units and deemphasizing the need for operational knowledge) during the 1970s.
When it comes to problems, particularly in complex subjects which aren’t yet well understood, and where there aren’t yet an overabundance of high caliber researchers, this comes across as dismissive of the problem.
This can be even more concerning if we know there are investors lined up who will happily sell something to the world and who will intentionally hide or minimize known ethical concerns. And then play dumb and shocked later when the very same problems manifest.
I see this dismissal of justifiable and real concerns an awful lot in conversations of all kinds lately.
And from what I’ve seen, no one here is anti-ML, no one on either side here is a luddite. But like so many conversations online, we should quit talking past each other and likely need to quit trying to paint people as if their concerns don’t have very real ethical implications which will, if left unaddressed, manifest in all kinds of negative ways throughout society.
Again, a lack of a neat and tidy solution doesn’t mean the problem doesn’t exist.
Specific recommendations start at 17 min.
I don't think this criticism is fair. Presumably if someone with the title "researcher" has a hand in actually doing what LeCun consider to be the engineer's role, LeCun would say to treat them as an engineer for the purposes of his argument.
>For the sake of ethical R&D, it's counter-productive to build a hierarchy of investment into the problem. Admittedly, the responsibility for end-results can differ, but the consensus that this ethical work is important should ideally be universal. That said, there should not be some ivory tower where you wash your hands of the ethical problems of your field.
Suppose you're correct, should we define this threshold based on just LeCun's heuristic for defining it? Or would it likely be better to have a consensus that your role doesn't matter in acknowledging the important of fairness and ethics?
Would you prefer the world's most prolific researchers being mindful of these issues, even subconsciously? Or to care less, perhaps very little, because it can be deferred to engineers?
Because I would like for the field to unite against building harmful systems and to acknowledge the importance of this work throughout the academic hierarchy.
As for the ethical problems. Are you familiar with machine learning methods? It really is all about the data; that's not a dismissal, it's a fact about current technology. There are other kinds of A.I. which are not based entirely on raw data like this. Machine Learning has been simplistically described as "curve fitting", which I think isn't a bad description. So here you are arguing that the mathematician researching good ways to fit a smooth curve to a series of points in really high dimensions needs to somehow take into account what those points might represent in someone's use of the technology. It seems pretty unreasonable to me to require that they put ethical constraints on it.
I'm not advocating that purely theoretical work of every form should be focused on algorithmic bias. Certain realms of theory (adversarial and robust learning, deployable model theory, computer vision and NLP) lend themselves much more directly to societally-relevant bias than other realms (algorithm and complexity theory, hardware research, performance research, pure statistical learning theory). I work on faster graph neural networks, so I'm actually in latter group.
So? I just want everyone to be on the same page on the importance of work in algorithmic bias. I've seen people dismiss Gebru's work as "pure rhetoric" on Reddit - this is a cause for concern! Acknowledging the validity and importance of similar work is especially important for people who are leaders in the field and who have influence over priorities. Don't people on HN complain about FB's algorithms literally every day? Shall we forget the Myanmar incident?
Let me put it this way: in biology, Watson and Crick were researchers who participated in discovering the structure of DNA. For most of their careers, they were not practitioners. However, James Watson made a huge negative impact on the field by advocating against woman in science and advocating for eugenics (completely ignoring Rosalind Franklin). Setting an aggressive and toxic tone in genetics paid dividends during the Asilomar conference (1975), where reporters and scientists who were critical of big-name organizers got de-badged and escorted out of a conference on ethics.
Our leaders and researchers matter - let's not make the same mistakes. The message should be: "I might no longer tool in TensorFlow, but I care, and so should you." It was easy for Jeff Dean to do that, which I thought was awesome. No senseless purity testing of engineer versus scientist. I think small steps like the new NeurIPS broader impact statement are heading in the right direction.
Edit: I just realized this reads like I'm equating LeCun with Watson... that's not my intention. That would be incredibly insulting, my apologies. I just needed an example of leadership having ripple effects throughout a field. Mea culpa.
I am personally of the opinion that it is fundamentally impossible to advance technology in a one-sided way. Anything with the power to do good can do evil too. Power itself is the danger. There might be a logical proof of this somewhere. Step very far back, and try to describe what a technology like a ML algorithm provides to society: software that can perform tasks as well as a person? discriminate between similar things using noisy observations? extract information that is obscured? The technology which accomplishes this can always be used both ways.
BTW, Watson didn't ignore Franklin- her name is listed in the W&C Nature paper as providing data. What he said in his book was much worse than ignoring her.
which adds a ton of color, and also kind of supports the problem with overly enthusiastic science reporters publishing things irresponsibly early (covid reporting is a good example).
Even if sincere good faith and he has real expertise is assumed that approach kind of raises several "huckster alert" red flags.
Except for the part where it provides an actual workable solution to the problem at hand.
To me this is an argument between completely different mindsets, one that restricts itself to provable facts and one which restricts itself to political agendas. I don't see how the latter can also work in facts. Or belongs in a technical research discussion at all frankly. You want to make laws that force companies to produce identical/equivalent outcomes for every race somehow? Just go lobby for it. Perhaps it's a good idea. You aren't going to reprogram mathematicians to think in political terms instead of mathematical terms.
Sure, you can look at it that way if you like, that commonly results in hiring someone to be responsible for D/I without actually making any other changes.
A better response is more along the lines of "Not in MY Army" which makes it everyone's responsibility at every level.
Recently seems in support of a black scholar who cried racism because her non peer-reviewed work that was only posted on arxiv wasn't cited in a lecture on GANs ...
Voicing these opinions would probably label me as racist in their book ironically.
Don't worry, Twitter isn't real life. People have become experts at shouting down people who disagree with them on that platform, they aren't seeking proper arguments and make disingenuous attempts to present them as honest debates.
But fortunately what's popular on Twitter doesn't translate to the average population. Plenty of completely fringe ideas get 50-100k likes/retweets. It mostly just represents the voices of various super-niches living in bubbles.
Small but highly vocal groups can have a seemingly loud and powerful voice. Yet the results of polls and other public signals (even election outcomes) are frequent reminders that what is gospel on Twitter is often detached from 'real life'.
Using a president of a major country is a poor example in this context. But if anything Trump being one of the first major Twitter users strongly reinforces my point. Prior to election he tweeted plenty of things most mainstream US republicans wouldn't touch with a 10 foot poll. Let alone what an average American would say IRL (even right leaning ones).
Not to mention Twitter is a global platform so conversation around local politics can be heavily skewed by people not even in the country.
But otherwise I agree, it is infesting real life far more frequently these days. And it is worrying. Despite everything I said above, big corporations, the media, politicians, etc can't seem to make this distinction and take what is popular there as a direct reflection of the general public. And it creates a negative reinforcing spiral.
If you're lucky and the offense is mild you may repent, as it was suggested to Yann: https://twitter.com/le_roux_nicolas/status/12754857390238187...
I wonder how often apologies like this are genuine, versus simply bending the knee to the mob out of fear for one's livelihood.
It made me feel very uncomfortable, because I wanted to say something like "the issues you are pointing out are very important, but I don't think you are right in this instance" but I was afraid doing so might jeopardize my career at the company.
Engineering, as a discipline, has an unfortunate history of not wanting to engage with the social and political conditions under which its work occurs, but that becomes entirely untenable for how a lot of ML is put to use (if you’re trying to be honest anyway).
Are there any industries that actually do this well?
Why not just edit the end results to show what you really want?
And that’s the deeper problem here: “it’s just a biased dataset” is a misdiagnoses. It’s a whole system of biases that leads to people thinking they are training with balanced data when they manifestly are not.
You’re never really going to achieve this mythical “balanced training data” until you untangle all of the other implicit personal and organizational biases. There are a whole host of ethical discussions that need to happen to even begin to flesh out what “balanced” might even mean for, say, facial recognition software intended for use in law-enforcement, but the same biases that lead people to skip right past those discussions and begin training are often the very ones that result in the biased data to begin with.
It could also be balanced so that the evaluation metrics were similar for each subgroup (possibly ending up with a sample that's very different from the population). But what are the subgroups? For any commonly used definition of race, there is a lot of intra-group variety.
Maybe balancing the training data is enough. But figuring out what balance even means is a huge question.
I don't think there's a ready CS/stats solution to this problem, so it will require interdisciplinary engagement and listening to the people who have been on the wrong end of facial-recognition bias is likely a place to start.
BTW, what exactly is bias in this context?
My intuition is that no data set can take into account all aspects of humanity and therefore cannot be bias free.
However, I believe that exhaustive data sets have a great potential of being less biased that human beings.
Also, in the case of statistical models, the crafting of the trained features themselves.
Actually, this is also relevant for neural networks despite the fact that they learn their own features because some amount of "framing" of the raw data often takes place in order to focus the neural network on the portion of the input data the trainer sees as relevant. This removes noise, but also removes context.
You asked about the biases of the people building the model, which is what I answered.
You didn't ask about the biases that occur during the requirements specification stage, or the biases that occur during operational implementation of the trained model.
Those are just as important - and arguably even more important - than the choice of the training data and the technical implementation.
The responsibility for the ethics of using ML neither begins nor ends with the ML engineer who builds the machine, and there are serious questions arising from the application of ML in certain domains that cannot simply be addressed by "better training data".
And this, I think is the knife that separates the different schools of thought on the issue.
People who are judging whether an ML model is "good" or "bad" based on this criteria necessarily see the accusation of "bias" as a claim that their model is not successfully predicting things. They rightfully retort that they would do a better job with an unbiased dataset. To argue they are their models are always wrong on their terms is to argue that there is a Ken Thompson-like hack in their mathematics. [1]
On the other hand, people who judge ML models by criteria like how they might be used or interpreted by laypeople are fundamentally talking about something other than ML models-qua-mathematical models. To the modelers, you might as well be arguing that the theory of nuclear fission is biased against the Japanese. But you are not actually talking about the empirical quality of their model, and so on your own terms you are correct. The models can be used improperly, and researchers should be careful about how their findings are perceived.
I just don’t know how one would prove it, and as others have noted I don’t understand what the mitigating alternative in the short term should be other than just stopping the research.
ML tools are developed, and filter down into the general developer population where they are used without fully comprehending the biases they contain or can contain if used incorrectly. I'm pretty partial to the idea that ML is a systemic bias footgun. VERY hard to not use incorrectly, and with potentially HUGE social repercussions. E.g. the youtube funnel towards extremism that was well documented a few years ago - visitors start on some innocuous video, and get recommended more and more extreme things, funnelling traffic towards the alt-right in an unhealthy way.
For example - in the small town I'm from, there are practically no people of color. It's an extremely homogeneous town, in terms of ethnicity. Everyone's white - the only thing that seems to differentiate people, are their socioeconomic backgrounds. So in a way, ethnicity becomes a irrelevant feature, if we were to use it to say, profile and predict criminals.
And what's more - if we were to transfer that model, which is trained on data from my small homogeneous place, it would probably generalize very poorly in areas with more diversity.
On the flip side - if we were to make a dataset out of criminals in some large and diverse area, we'd need to get our sampling methodology right. Maybe it just so happens that the law enforcement tends to pour all their resources in policing poor areas, where some certain ethnicity is very over represented? As you can see, the further down we go, the more this problem moves from models -> datasets -> sampling -> policies, and so on.
And as for our imagined crime profiling model, you can see that the scientists are oftentimes forced to work with wildly different datasets, that may look very different from one another. You get a bunch of different models with great local optimization, but which fail to generalize on the population. And what's more, those locally optimized models may work perfectly well for their intended tasks, so there will be no complaints - until one of them are deployed to cover more general cases.
Yes - the ML engineers and scientists are responsible for building good generalized models.
Yes - the data engineers and scientists are responsible for building good datasets
But alas, there's only so much one can do about the above. In the end, the data actually stems from something real and tangible, and if there's a systemic bias in how that data is created, then that's going to be the greatest factor for everything which comes after.
It's a very deep and complex topic, which goes way beyond this. But sadly there are real-life consequences, when models are being deployed in the real world.
More likely, the reason that bias isn't routinely fixed is that it isn't easy, and these kinds of biases do make it into production systems. Which makes it a net positive in my view that the occasional shitstorm reminds society of this fact.
Can that be unpleasant for AI researchers? Sure... but if it bothers then, then perhaps they could focus their research on trying to fix the problem?
Physicists had and have unpleasant conversations about their moral responsibility for nuclear weapons. Other fields of research should take their moral responsibilities seriously as well (not just AI research, by the way).
Sometimes it is useful to just summarize a problem as clearly and simply as possible.
I also think there's value in simple true statements.
The issue is LeCun makes an argument about fundamental research, but is not exactly a fundamental researcher, and does not necessarily represent fundamental research.
As an analogy, if you are a researcher of general chemistry, persumably there is no issue. However, if your research is specifically about chemistry for improving bullets, and you produce working prototypes, then some might say you should be subject to some regulation as a part of the arms industry. I'm not saying this is the right thing, just that such a point could be made. LeCun is arguably much more the second kind of researcher than the first. The research he represents is "better ways to recognize faces", not "statistical properties of natural images".
To take another example, there is an enormous amount of regulation in, say, medical research. And in this case there are good reasons for that. Gebru could be possibly arguing for something similar in say face recognition.
To be clear, I do get arguments that ML might increase the efficiency of policing, and there is some inherent tension between the efficiency of the state and freedom, to the extent that imperfect enforcement is a tacit feature of our legal system. But, this is not an inherently racial point.
I really don't want my field to regress towards 19th-century pseudoscience when we can empower other scientific advancements that will yield immediate benefits.
This is a strange point. Are personality types subjective? Is IQ? Maybe they are, but we still sometimes find them useful to quantify. If it's ok to predict these things from your answers to questions...why isn't it ok to predict them from pictures of your face?
So yes, I think I can state pretty concretely that there's no value in trying to predict IQ from face shape.
The presence of error in a prediction makes that prediction valueless?
Can you describe a situation where the harm caused to someone by mis-classifying their IQ is outweighed by the increased efficiency from...being able to more quickly identify someone's IQ? Like, can you even describe a situation where this will be used ethically at all? What use is there for predicting someone's IQ with dubious accuracy, assuming you already have their permission?
Yes.
> Do you think their error rate is zero?
Well it depends. Do I think the error rate of IQ as a predictor of underlying intelligence is flawed? Yes.
Do I think that IQ as a measure of IQ is flawed? Still, often, yes.
> These things don't need to be perfect to be good.
To my point above, I generally am not a fan of IQ in general. Psychology researchers are consistently finding new and interesting ways that IQ tests are socially/culturally/environmentally influenced and that exams may not be fair.
If your claim is that you can somehow train an algorithm to look at someone's face and evaluate their underlying intelligence better than the tools we have today, honestly "delusional" is the first word that comes to mind. That's just not something that currently available tools can do. And thought experiments about a "perfect" classifier aren't interesting when we're discussing applied ethics.
And that comes with the huge, enormous, gigantic, asterisk that you can even measure "intelligence". Like, it comes with the caveat that you can even reasonably define intelligence. We generally measure intelligence as correlated with success at some metric: math problems, visual puzzles, chess, reading comprehension, whatever. If you're contending that a hypothetical model could predict underlying intelligence better than the measures we have today, how would we even know? When training a model you have an evaluation function.
If the evaluation function is flawed (which is essentially your contention, and which I absolutely agree with), the trained model will exhibit biases that reflect the flaws of the evaluation function. An ML model isn't suddenly going to solve the cultural issues we have with measuring IQ. It will encode the same biases that society does, because the model will try as hard as it can to do exactly what the human proctors would have done.
These are exactly the kinds of broader ethical questions that Timnit points out we need to worry about.
> If the evaluation function is flawed (which is essentially your contention, and which I absolutely agree with), the trained model will exhibit biases that reflect the flaws of the evaluation function. An ML model isn't suddenly going to solve the cultural issues we have with measuring IQ. It will encode the same biases that society does, because the model will try as hard as it can to do exactly what the human proctors would have done.
Of course. But so what? There are people using IQ now for things. An ML model isn't going to magically make those biases worse, either. What it's going to do is bring them to the surface, so that they are quantifiable and we can actually do something about them.
ML is the solution to the bias problem. Right now these evaluations are being made by other humans. Humans who we cannot statistically debias. Humans who's biases we can't even effectively interrogate. The reason people are making all these memes about biased AI is not that it's more biased than humans, it's that the bias is more measurable.
Well, actually, research shows that unless great care is taken it absolutely can. If you include for example race as a factor in a model, it can learn non-causal correlations between race and whatever the objective is. This can have a compounding effect in some cases. [0]
> What it's going to do is bring them to the surface, so that they are quantifiable and we can actually do something about them.
I don't follow this. If we have a biased objective function, the model won't surface any biases we weren't already cognizant of in the objective function. And they were already quantifiable: we had a function that we were using to evaluate the model. We could use that same function on whatever non-model evaluation we were doing.
> ML is the solution to the bias problem.
This is basically directly in contradiction to what leading experts on the subject say. ML cannot fix bias in human systems, unless we presuppose that those systems are biased, in which case we can often address the bias in the human systems directly without ML.
> Humans who we cannot statistically debias. Humans who's biases we can't even effectively interrogate.
You can still have decisions be made by objective expert systems without complex ML. If you want to learn someone's IQ, the best way is to debias the IQ test, not to try and infer it from their face bones.
> it's that the bias is more measurable
If we can measure the bias in the output of an ML model, we can equivalently measure the bias in the output of a human system. You're presupposing the existence of some unbiased objective function which we don't have, and that's at the core of the issue.
[0]: https://www.wired.com/story/ideas-joi-ito-insurance-algorith... has a few good examples here, like how naive bail and sentencing models encode racial bias that isn't present in humans. And to be clear the response here shouldn't be "well let's just build better models" but "why do we think a model will improve the situation here at all"? Removing agency from Judges has historically been bad for the average person convicted of a crime. This doesn't mean that individual judges can't make terrible rulings, but that the alternatives are usually worse on the whole.
We can actually follow the logic of the model. For instance, you can theoretically de-bias a dataset by building a racial classifier from it. What you need is an objective test for the presence of racial information, and that's easy to obtain: Build a classifier to explicitly predict race from your feature set. Train an adversarial model to reconstruct your dataset with maximum fidelity, subject to the constraint that race can no longer be predicted from it.
> This is basically directly in contradiction to what leading experts on the subject say. ML cannot fix bias in human systems, unless we presuppose that those systems are biased, in which case we can often address the bias in the human systems directly without ML.
These experts are just wrong, then. Naive ML won't fix bias in human systems, but that doesn't mean we can't use ML to fix it, if we do so thoughtfully.
> You can still have decisions be made by objective expert systems without complex ML. If you want to learn someone's IQ, the best way is to debias the IQ test, not to try and infer it from their face bones.
Sure, but there are a lot of things that we don't do in the best possible way because it's too expensive. There are lots of use cases for cheap, scalable, low precision models.
> If we can measure the bias in the output of an ML model, we can equivalently measure the bias in the output of a human system. You're presupposing the existence of some unbiased objective function which we don't have, and that's at the core of the issue.
Right, but we cannot fix the bias in a human. And humans are heterogenous and inconsistent. The same person may be more or less biased on different days. The ML model is consistent, and we can incrementally improve its bias in tangible and testable ways. The same is not true of humans.
And what does this get you? Let's look at a face recognition dataset. What happens when you debias it? Is it still useful? No. Because the faces no longer resemble real faces.
> These experts are just wrong, then
Perhaps, but you aren't making a strong case for that.
> There are lots of use cases for cheap, scalable, low precision models.
That involve facial recognition?
> Right, but we cannot fix the bias in a human
We don't need to. We just need to fix the bias in the system. And we absolutely can incrementally reduce bias in systems that involve humans.
Not to you. But you can remove the racial information without destroying all the information that a model can detect.
That's what the ethicists say: don't use facial recognition models. Don't work on them. Don't research them. They cannot be both unbiased and useful. And in general, there's few to no uses that are ethical, period.
Some of these correlate with race. But only part of the information correlates with race, not all of it. It is, in principle, possible to remove the information that identifies race without destroying the information that identifies the individual. It is true that part of an individual's essential characteristics are their racial characteristics, but it is not true that the only way to identify an individual is their racial characteristics. For instance, there is no way that i'm aware of to infer race from fingerprints, but you can absolutely identify a person by their fingerprints. So, the question is, can we extract a facial fingerprint that identifies a person, but not their race? I think the answer is almost certainly yes, and it is going to be up to a clever model design to do it. But essentially it would look like a GAN where the adversarial component is constantly trying to predict race, while the Generative component is trying to trick the race classifier without tricking the person-identifier.
Or perhaps they understand that this won't work in practice.
Personality, aptitude, criminality, intelligence, empathy, etcetera... that's almost surely pseudoscience.
Well, that's not a very full story. It was scientifically invalid because it modeled mental traits based on the shape of the skull, and that's the important part. If you trained a model to predict traits based on accurate data about skull shape and traits, I'm confident that the algorithm would happily do that. Phrenology would still be bad regardless of what technical methods are used to do phrenology.
> We have noticed a lot of concern that PULSE will be used to identify individuals whose faces have been blurred out. We want to emphasize that this is impossible - PULSE makes imaginary faces of people who do not exist, which should not be confused for real people. It will not help identify or reconstruct the original image.
This is just another way to generate realistic faces that match a pattern. The "bad use cases" I can imagine involve unmasking real people, which is not what the model does.
For example, there are already attempts to use expert systems in the courts - to "scientifically" determine flight risk (and thus bail), and even chance of re-offence (and thus punishments). Those are trained on datasets of convictions. But we know that both conviction rates, and arrest rates, are themselves indicative of bias - e.g. some people get arrested more often for a crime, for which other people might get a verbal warning from police.
So ML takes those biases, and entrenches them further. Worse yet, because people tend to believe that "computers cannot be biased", the fact that it's ML that's producing a specific result, is taken as prima facie evidence that said result is immune to criticism from that perspective. And with most advanced models, you can't really crack it open and make sense of what's inside - so you can't easily prove that it is biased.
I can think of at least one trick that might work. Build a classifier capable of identifying race from the data, and then reverse its gradient to remove its signals. This would deracialize a dataset, so that there is no detectable latent information in it.
However, whether or not my particular solution works is somewhat beside the point. It's an engineering challenge to figure out how to do this, but that isn't an argument that ML is fundamentally bad for these purposes.
> For example, there are already attempts to use expert systems in the courts - to "scientifically" determine flight risk (and thus bail), and even chance of re-offence (and thus punishments). Those are trained on datasets of convictions. But we know that both conviction rates, and arrest rates, are themselves indicative of bias - e.g. some people get arrested more often for a crime, for which other people might get a verbal warning from police.
I totally agree that this is a problem. But it seems to me like the solution is more effort on fixing the biases, not reverting to human judgment. I think it's always important to make sure we understand what the baseline is we're comparing to. Humans are racist. Humans created the biases in these datasets. These algorithmic tools just bring those biases to the surface and make them quantifiable. It seems to me that that represents an incredible opportunity to actually fix the biases in a measurable way.
For example imagine some ML algorithm that predicts the potential value of a customer (or their likelihood to steal from you) as they walk in the door of your business. How do you remove the racial bias from a system like that when racial biases exist in society that result in the average White person being in a higher socioeconomic class than the average Black person? The ML algorithm will tell you White customers are more valuable than Black customers. It is doing exactly what you told it to do, but the end result is that Black people get worse service in your business.
It's saying that the dataset isn't reflective of the world, and this is an issue because then you'll end up with voice tools[1], facial recognition software[2], and sentencing algorithms[3] end up performing disproportionately poorly on some minority groups. And this ends up reinforcing the poor positions many people are already in.
1) https://www.pnas.org/content/117/14/7684
2) http://proceedings.mlr.press/v81/buolamwini18a.html?mod=arti...
3) https://www.liebertpub.com/doi/full/10.1089/big.2016.0047
You can see in the responses however that many people were dissatisfied with his response.
The solution isn’t necessarily to fix the data or the model. It might be banning use of facial recognition by police, for example, which some places are doing.
ML isn't biased, its the people who keep loading them with biased data, and then running them against photos of people.
It isn't clear what a similar safety mechanism would look like for an ML algorithm. Modern ML algorithms are nothing more than generic pattern extractors. What would a safety mechanism look like?
Have government grants for ML research be contingent on training on diverse datasets. Fund research into more diverse benchmarks and push them as the goal to beat.
However, that framing of the discussion can easily be interpreted as absolving the ML community from the ethical discussions around the nature of the datasets.
Is it acceptable for the community to hand-wave these issues because they are "dataset problems"?
For what it's worth, this is a much larger debate that is happening in many fields. For example, a bunch of decisions around car safety were based on crash tests. Those crash tests are based on dummies that were designed around male body forms and thus don't really test how women would handle the crashes. To quote the author of Invisible Women:
> As a result, if they are involved in a car crash, women are more likely to be injured – 47 per cent more likely to be seriously injured and 17 per cent more likely to die.
This general questioning about the implications of bias in the datasets is happening in many fields
How was causation established?