LeCunn made repeated attempts to explain his position, only to be virtually shouted over about he wasn’t listening to people of color because he pointed out that a model trained on White people learned to produce images of White people, and that if it had been trained on Black people it would produce images of Black people. It was a research paper, not a production system, and Gebru and her acolytes were much more interested in scoring cheap points than in having a serious conversation about either ML fairness or the merits of this particular paper.
Seriously, the primary sources can be read by anyone.
Am I missing anything?
I think it’s clear Gebru was acting in bad faith and LeCunn was baffled and trying to deescalate. Gebru responds to each attempt to deescalate by a reply designed to further rile up Twitter, without actually talking to any of the points LeCunn makes. Virtually all of her replies are some variant of “you are wrong, but I won’t say why” and “listen to me because I am Black”.
The same is true of AI research Twitter.
> Virtually all of her replies are some variant of “you are wrong, but I won’t say why” and “listen to me because I am Black”.
I don't see that in the summary you posted.
I'd say she deserved to be fired for she is a racist who's damning to a normal society.
Whom did Gebru discriminate against, based on their race? No one.
That episode is unfortunate imho, but she only suspected racist rationales... Calling out racism does not make you a racist, just like calling something "fishy" does not make you grow fins and gills
Which she doesn't.
Most notably, people here have been raising concern about the discussion with Yann Lecun, but that was also just a civil, if tense, discussion.
You're better off not twisting the meaning of words to favor your argument, though. It'll just create more problems down the line
And her paper was rejected and her immediate reaction is to demand the company to reveal the identity of every reviewer? Yeah, right.
But this claim is certainly not necessarily correct, and he shouldn't have been so confident about it. Any part of a system can contribute to bias and that includes the model design, not just the data. If they want the system to work then they need to actually demonstrate it works, not just say it could with X change without testing that.
Though I remember people arguing it didn't work properly at the same time they were saying it was evil (because it contributed to surveillance). Which is odd because if you don't want it to exist, you shouldn't want it to work either.
This is not surprising. Black faces and White faces are not the same data manifold. This is like training a network to upsample oranges, then running it on an apple and being surprised when the result is an unusual color and texture for an apple.
I’m not sure what else there is possibly wrong here. Is the width of their convolutions racist? Their choice to work on super resolution? The fact that they released their work for reproducibility?
Since everyone else is just saying "you don't get it" without explaining what "it" is, I will provide some brief avenues of exploration. Because they are brief, they will be coarse and imperfect, and aiming to give you directional assistance on the topic.
So with that massive preamble because I'm not seeking to argue, here you go:
* Pre-trained models already encode much of this bias and if you don't use them you won't get very far very fast
* Large available datasets also reflect these biases
* AI model performance is generally measured against this dataset, further entrenching the bias. If you're better on Indian face generation it will give you barely any benefit on most scoring methods
* Saying "the outcome is only biased because the data is biased" misses the fact that the performance is on the biased data
There's a little more to it but I believe that's the meat of it. Anyway, not too keen on arguing this. Just sharing because it took me some work to figure out what they were talking about and I wish someone had explained it to me so doing so here.
For instance, if I were to champion metric A which purports to measure performance on human faces but it really only rewards performance on Senegalese then models that do better in general may not be recognized for being better.
In an isolated sense this is not a problem. However if the mainstream is that all metrics that are taken seriously are dataset-biased then we'll have an environment where the models will be trained on biased datasets in order to be successful.
For instance if everyone uses LFW to determine how good facial recognition is, then Senegalese fine-tuned facial recognition tech will not be recognized as being good at facial recognition.
So the argument is that dataset bias is built-in to our approach to the problem. I, personally, think that this isn't malice. We need benchmarks to judge approaches against each other. Benchmarks always have a first mover advantage and a massive path dependence issue. The first benchmarks aiming for things on humans do not reflect humanity accurately. These benchmarks became standard among the community. To be taken seriously you have to do well on benchmarks that are standard in the community. Dataset bias is then natural in newer approaches because the approaches are judged against how good they are against the inaccurate (if you will) benchmarks.
So no one need be racist or anything for the end result of the field to end up being discriminatory.
I don't think a successful approach is to call people racist over this. After all, it isn't malice that guides them. The discrimination comes from the sort of historical accident that has North facing up on a map. And no individual is really racist. It's sort of like the Bechdel test: no movie is crappy simply because it doesn't have two girls talking to each other, but if very few movies have two girls talking to each other about something other than boys, then it makes you think "hmmm, why's that the case".
That's, like, not even culture war, it's just basic correctness of the reference datasets. If a fruit classifier was missing oranges, we'd just fix it and move on.
Like, for instance, HN has people who will bring up privacy violations of big tech constantly. They see their role as making sure the conversation is happening. Not justifying this. Just aiming to understand it.
For my part, I prefer to take the approach you're talking about because I, too, think that the fastest path to this is getting the photos, labeling the photos, and then lobbying for inclusion. Ultimately, I think it's okay if things optimize fast for growth and then we fix up issues afterwards. So the people building the benchmark sets weren't able to get a set that's representative of humanity. Should they have waited till they could have done that? IMHO, no. Rapid release moves the state of the art forward and then we can put in all of these corrections as we move.
Then there's the question of whether all-humans dataset is a good thing or if instead a thing that is white-humans and another that is black-humans is better. Anyway, all said, my personal approach to this problem (if I cared about it a lot, which I don't) would be to say "Current benchmarks and training data available bias towards certain races. I'd like to build X/Y to solve that. Here's what I have so far" etc. etc. I think positive engagement like that yields better results because the vast majority of scientists actually aren't weird race supremacists and the vast majority of AI researchers will gobble up any more data you give them which is segmented and labeled differently, the greedy bastards :D
Some challenges that I can still see:
* Getting the data. Might not actually exist.
* Labeling the data. Probably needs some work.
* Getting it into the benchmarks. It'll invalidate old scores, so there's just a product adoption problem here. I don't know how the community handles newer releases.
As a last aside, I suspect that this conversation ended up the way it did because:
a. It's charged. It's race-based differing outcomes. That's a sensitive subject.
b. People feel unheard. This is natural. Like, this is not an 'interesting' problem. It's literally just a data error so the luminaries in the techniques part of the field aren't really that interested in it. And the techniques part is where the sexy is.
c. This sort of thing has a tendency to escalate. One side says "You're not listening to what I say" and the other side says "I'm not racist. I don't get why you're calling me that" and before you know it it becomes "You have to be racist to be ignoring me" and whatnot and de-escalation becomes impossible. Especially because everyone rewards the loudest on each side.
Honestly, I think it's quite interesting to observe and to understand as just a view into the human condition but we use AI models professionally in the GIS space and professionally we just don't go near this at all. No part of me finds it interesting to solve or to interact with the discussion in any way. I only sort of participated in this here because I think I managed some insight into what it is and I wanted to write that down because I wish someone else could have accelerated me into it.
Anyway, I think that's all the insight I have on the subject, so I'm going to just leave it there. Any more and I'll be ass-pulling.
Do something productive and uncontroversial, or something unproductive and very controversial?
Not that parameter specifically, but if you say the data is wrong then that means any other parameter might be wrong too. More data might mean the model is too small to fit it, or the hyperparameters might be wrong to train it, or you now have the wrong ratio of other phenotypes (let's say that instead of races…) in the training set and their results regress.
Also, if you're adding people with darker skin, that of course means the pixel values are lower. That matters for image processing code, things like SAD thresholds or noise reduction will work differently.
> Their choice to work on super resolution?
Superresolution is only useful as a toy and they should have mentioned that when they put up a live colab, yes. They added a disclaimer later on and it was good - there was a big issue where people were convinced this was somehow a surveillance technology because they watched too many TV shows, even though it literally can't work that way!
in particular because features contrast (dynamic range) is lower for darker faces.
>Is the width of their convolutions racist?
Not width. As a result of the above mentioned lower contrast, the racist here is the sensitivity of the resulting Gabor filters produced by the training in the first convolution layers and the Gauss filters in the next layer's. I suspect to deal with that problem one would have to normalize dynamic range of the faces, i.e. it would look something like kind of lightening of the dark faces and/or darkening of the white ones.
The most obvious and easiest way to clear this up would be for the original researches to train the model on a "Senegal" dataset as mentioned by LeCunn and see what the results are.
While it would be a symmetric situation a vacuum, we do not live in a vacuum. And acknowledging dataset bias by itself doesn't address the bias meaningfully. In practice, we often treat the bias as an exogenous factor when it is not, moving it outside the scope of our responsibility. But it is very much the product of our work, a reflection of our choices, values, and beliefs about what to prioritize. We can't abdicate our responsibility for it; the stakes are too high.
I am not sure why the wording around this has to be so abstract.
As you say, it's a reflection of our beliefs about what to prioritize. If you're interested in developing methods of image upscaling that generalize better and preserve facial properties that were underrepresented in the training data, that's an interesting area of research, and you're welcome to work on it, but you don't get to demand that people working on something else prioritize this instead, they get to choose their own priorities unless you're paying them for a particular direction of research (e.g. like Google should be able to direct Dr. Gebru while she was working for them).
There's a responsibility to correct for bias when implementing models in production (e.g. if you'd be actually deploying it in Senegal, then there would be a responsibility to use a Senegal-appropriate dataset), there's a responsibility to acknowledge the bias of a particular algorithm if one exists - for example if it highly relies on contrast values which would be different for different types of faces, then that's relevant, because it's a statement about the generalization of the algorithm; but there's no proactive duty to "address the bias meaningfully", that's nice thing to do, but it's like charity - a voluntary choice to factilitate a social goal, but not a requirement or responsibility to do that. There's a responsibility to correct harms you caused, there is no responsibility to correct harms caused by others, that's a good thing to do, but not a moral duty.
Someone who invests a lot into charity work or addressing bias is doing a good thing, but it doesn't mean that people who are doing other things are abdicating their responsibility - it's never was their responsibility in the first place to fix social issues in the wider society. You can't simply point at random people and declare that they're going to be responsible for something that they didn't personally cause and where they did nothing wrong, that accusatory behavior is unacceptable, so naturally there's a backlash to people who try to assert personal responsibility of others without a basis to do so.
This would be called “research”. That is what she was hired to do at Google. For everybody of every race ethnicity orientation able status and background that would be honored to give their best effort at the opportunity to do research at google, acceptance of resignation was the right thing.
My only point around her feud with LeCunn is that she threw a fit because people including LeCunn pointed out that it is not necessarily a race thing, and she wanted to give it a race spin without confirming that it was actually the case.
And I am not at all commenting on all this drama around her resignation/being fired without getting a look at the contents of the two emails: the one she sent to Brain Women and the one with her demands. But I know this: if I were to issue an ultimatum to my employer, I am treading dangerous waters and should be able to digest the outcome of that, including being fired.