600k Images Removed from ImageNet After Art Project Exposes Racist Bias
hyperallergic.com
hyperallergic.com
It's really only people where you can't tell what it does/is from the outside. Cars, trees, animals, mountains... everything else, if it looks a way it acts that way. Early AI will probably have just as much trouble with this as people have historically.
I really wish people would start viewing Racism as a willingness to let that primitive part of the mind be in command, rather than a binary attribute you either have or don't. Like, nobody is 100% not racist. There will always be slip ups, over-simplifications, snap judgements, subconscious or not.
You can't understand very much at all just based on how things look. That holds for humans and inhumans alike.
There are also cakes that look like cars, model sets which have miniature mountains, on and on. But in typical nature, it works well enough to breed and see what comes next.
When an arborist "looks" at a tree, they identify the kind of tree it is, and that connects to all of the research they know about that species of tree. It's reasonable to suppose that a new instance of a known kind of tree is going to have properties in common with all the other known instances.
?? Obviously you can. Why do you think you can't?
you can't identify every friendly dog just by looking at it, you'll identify dogs that are also acting friendly and doing things people have found preceded positive experiences.
you can't identify every sick dog just by looking at it, otherwise the vet would never find additional issues.
and so on
Yes, it's wrong to assume that you can tell everything about that dog by looking at it - but it's so much more uncomparably wrong to consider that you can tell nothing about friendly or sick dogs based on looking at them!
Observating the behaviors and dog breeds correlating with previous friendly and unfriendly experiences is a very useful, somewhat reliable predictor of how likely this particular dog is to be (un)friendly. Sure, if you know this particular dog, then that should supercede any group information, but if not, then that's all the information you have, it is useful information that correlates with (future) reality that matters to you, and so it's prudent to use it. Appearing "nonjudgmental" by throwing away information and judgment is simply stupid, and bound to get you bitten at least in the metaphorical way.
but nobody in this entire thread was saying that, and my contribution was an additional attempt to further point out why the other sibling responses were missing this
At this point I'm totally content in talking past each other on this topic as I honestly don't know why its not clear that the subset which is undetectable by mere observation is statistically significant.
Yes you can do that. To put it extremely simply: define what "friendly", "sick", and "lazy" means in terms of behavior, then observe the dog's behaviors.
Edit: sleeping is a behavior.
Your point about how a specialist can connect their knowledge to what they see is true, but not relevant to the point at hand. If I was your therapist, and had studied and understood you thoroughly, I might be able to assess your mental state at a glance in most scenarios. That doesn't deny you have a detailed inner life anymore than that an arborist might generally understand a tree does not deny the complex life of the tree.
I’d argue we learn everything based on appearance. You have to do more than just observe things superficially and take context into account, but for something to be measurable it has to have some sort of outward appearance, whether that be direct, as in something you can see with your naked eye, or indirect, as in something we need to measure through some other instrument.
This. Some Childhood development studies have born out that children preferentially treat their own 'ingroup' preferentially over those of an 'outgroup' [1]. This is before they really are "impressionable" as I understand it, so it shows something of an inbuilt mechanism. Though they made this distinction:
“Racism connotes hostility and that’s not what we studied. What the study does show is that babies use basic distinctions, including race, to start to cleave the world apart by groups of what they are and aren’t a part of.” [2]
1 - https://www.frontiersin.org/articles/10.3389/fpsyg.2018.0175...
2 - https://www.telegraph.co.uk/news/science/science-news/107705...
It might make interesting reading to learn about people who don't have an ingroup or who don't prefer their ingroup.
Or then again, it might just be depressing.
This was the classical experiment I remember: https://en.wikipedia.org/wiki/Realistic_conflict_theory#Robb....
At least some of the people conducting the studies on children are guilty of child abuse, for they tried to correct the behavior of infants with their infantile theories.
https://en.wikipedia.org/wiki/Thoughtcrime#Crimestop
We basically need this in software so that AI researchers can stop getting lambasted with claims of racism
Thing is, that's not even true in that it doesn't fully acknowledge the problem. You often CAN tell information from people's outward appearance, albeit probabilistically. Therein lies the problem: you can very easily train an algorithm to be maximally right according to your cost function, but end up biased because the underlying distributions aren't the same between groups.
The issue is that as a society we've (mostly) decided that unfairly classifying someone based on correlated but non-causal characteristics is wrong, EVEN in the extreme that you're right more often than you're wrong if you make that assumption.
For example: It may be politically correct to drive down a bad neighborhood in the middle of the night, but it isn't logically correct.
But note the implicit bias in your own comment: you assume yourself and all the readers of the comment are not people who live in bad neighborhoods.
that assumption results simply from the need for 'bad neighborhood' to be a negative scoring action.
if you oversimplify it to get rid of 'implicit bias' (which I don't agree exists in the example), the results turn into near-meaningless babble.
"For example: It may be politically correct to do something that ignores statistical dangers in favor of the promise of human goodness, which may result in the possibility for more personal endangerment than other choices, but it isn't logically sound to ignore such statistics for the hope of a less biased personal experience."
The example requires the person driving to be detached from the bad neighborhood that they have a choice to drive through. How that isn't an obvious requirement for the example to have merit is beyond me.
After all, you don't really need a law telling companies they aren't allowed to hire infants as senior officers - it's already not in the company's interest to do so.
However, when there is a logically correct but politically incorrect decision that the company could make, it is now that you need laws to prevent the company from taking that choice.
Of course, as the weight of an institution's decisions goes down, so does the need to police its actions. In particular, it is rarely necessary to prevent an individual person from acting on their biases.
Applying this to your example, if we had an AI that should suggest your best route home, and it avoided a short route through a bad neighborhood, that is likely ok. However, if a municipality used the same AI to decide where to prioritize changing street lights, that should be prohibited.
Sorry, but it's not a decision. Science has found repeatedly that using outward characteristics does not work as a good classification measure. Society simply enforces not making bad judgements. See:
Phrenology, Racism -- "appearance implies certain things about someone's character and intelligence" (https://en.wikipedia.org/wiki/Phrenology)
Sex -- "A person's gender presentation determines their gender or sex" (https://en.wikipedia.org/wiki/Transgender_history), "A person's genitalia determines their chromosones", "A person's chromosones determine their sex and their ability to give birth" (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC2190741/).
Wealth -- "That person is wearing bad clothing/living an unassuming life therefore they must be poor" (https://www.npr.org/2018/12/29/680883772/social-worker-led-f...)
Intelligence -- "This person scores badly on intelligence tests therefore they must not be worthwhile" (https://en.wikipedia.org/wiki/Richard_Feynman, https://www.psychologytoday.com/gb/blog/sudden-genius/201101...), "This person appears to be unskilled or incapable of many tasks therefore they must be totally unskilled" (http://en.wikipedia.org/wiki/Savant_syndrome)
Science has found that, on average, it works [1]. Of course there will be many cases where it fails - that's how statistical inference works. Whether that makes it a 'good' measure, by whatever standard, is a different question. But there is no doubt information can be inferred from appearance.
[1] http://www.spsp.org/news-center/blog/stereotype-accuracy-res...
Lol, what? Generically, that statement is almost certainly false more often than true in general in science. But I believe you are restricting yourself to social psychology?
You then presented a long list of examples taken from the tails of certain distributions to refute an argument that said distribution exists and has an average? I didn't even name any particular distributions. You're thinking appears flawed and emotionally driven, and most unfortunately, that's the type of thinking that will lead you to building biased systems.
Here's the point you missed the first time around: There are going to be outwardly visible characteristics that ARE correlated with some factor of interest, to the extent that training a machine learning algorithm based on a cost function that uses predictive accuracy alone WILL result in a system that assigns what society would consider an inappropriate importance placed on non-causal but correlated parameters.
Here's a real world example that might help you understand why this is important: (Data taken from: https://en.wikipedia.org/wiki/Incarceration_in_the_United_St...) Because blacks are over-represented in the US criminal justice system (40% of the prison population vs 13% of the population) and because part of what defines "black" is the outward appearance of certain facial features, a facial-recognition algorithm which is trained to recognize criminals, with a cost function based on prediction accuracy alone, and facial features as input parameters is going to have false positives that over-represent blacks. Does that sound like something you want? Because denying the underlying distributions is going to lead to exactly that.
It's very important to consider this when you develop a training set, for fucks sake. It might work something like this: Take 100 innocent people's faces at random. (On average it will have only 13 blacks) Then take 100 random criminal faces from inmates. (On average it will have 40 blacks.)
Then mix up the groups into your training set and assign a prediction score 1 or 0 depending on whether or not your classifier has correctly predicted whether or not a face was in the criminal group. Then, based on no other feature than race, your neural net can get better performance based solely on guessing more often that black people are criminals. That's not a good thing.
Do you get it now? The likelihood of being falsely identified as a criminal is greater based on the non-causal but correlated variable of being black. And this has happened several times already! You can't keep pushing this narrative that neglects the underlying statistics because of your beliefs, or people will keep making racist systems.
Trying to identify potential for criminality based on looks may well work, identifying and using the underlying racial biases. And if those systems are used, where people identified by these systems are more likely to be identified as criminals and investigated, we end up with a feedback loop and, over time, the racial biases in a society that uses the systems will get worse. More black criminals will be caught, as they are more likely to be suspected as criminal, making the racial biases worse for blacks. While the opposite happening for white faces.
Similar social evolution would happen if you pre-screened job candidates, over time magnifying existing gender and racial biases. I've seen in some cultures it is required to add photographs to job applications, but I think it a good thing that this practice is discouraged in western countries.
Say in your example of criminality, if black people were more predicted by the classifier, it isn't racist if it accurately reflects the base distribution. There's a line to be crossed here, and to me it's if that application starts to significantly (where the boundary lies here is up for debate), and directly affect the non-criminal portion. I'd say that if the misclassification error is similar across racial lines, then there is no issue.
Additionally, I don't quite agree with the "making racial biases worse" argument either. The way I see it, we already use racial heuristics in law enforcement. With automated, replicable tasks, we can at least quantify the degree of bias and correct for such.
Most people would probably opt for utility mixed with representational fairness of some degree even when it means law applies to some degree differently to groups and special cases.
Justice is not fairness. It has an institution called motive. (Which is often overly simplified or ignored.)
It incorporates elements of fairness. And it is hard to train. Not everyone can be Solomon, for example.
https://www.newscientist.com/article/2114900-concerns-as-fac...
In short, being humane. If a certain degree of racism or stereotyping is necessary for that goal (definitely not 100%, but something low), then so be it.
Currently multiple social system in the USA are considered much too racist.
What underlines it is seemingly a lack of understanding of statistics, becausee the most "examples", if they can even be called that are from the realm of non-common outlier cases. That's not what statistics operates on mostly - it operates on common cases that cover the largest part of a distribution. Not outlier cases on the ends of a distribution.
In every case you gave an example of some uncommon exception, as if that would somehow be an argument for the rule existing?
That sounds like the wrong basis for calling it extreme. It's not at all extreme to say that classifying an interview candidate based on correlated but non-causal characteristics is wrong, regardless of the statistical significance of those correlations.
https://www.fool.com/investing/2019/05/15/rihannas-fenty-is-...
Video: https://www.youtube.com/watch?v=Zn7oWIhFffs
Slides: https://www.chrisstucchio.com/pubs/slides/crunchconf_2018/sl...
It is a kind of optimality condition on all three goals.
The robustness additionally means that should conditions change, the algorithm usually will become better not worse and should a degradation still happen, it will be graceful and not catastrophical.
It's a hard and open problem in ML and especially ANN, design of robust solutions in the space. Most have really bad problems with it even when debiased.
Not that I mind that too much!
Ironically, the only places where it's legally prohibited or frowned upon to use these heuristic techniques are situations that people have arbitrarily (heuristically or conveniently) decided.
For example: It's "not fair" to hire someone because they're white (and consequently have a higher chance of being wealthy and hence a higher change of being educated.)
But it's "fair" to choose a love partner based on their height, their waist-to-hip ratio, their weight (and hence having a higher chance of giving birth to healthy offspring, better physical protection, etc.).
Maybe it's hypocritical, and I don't know if that's a good thing or not. Maybe being hypocritical helps us survive.
I get what you're saying, but no one moves jobs every week. The sunk costs of switching employment are significant for the employee, less so for the company.
To a lesser degree, the same goes with vacation time and other benefits. At least in the US, anyway. This is why having some of this stuff coded into law and decoupled from employment takes some power away from employers.
Of course, this is a point of view and not everyone agrees with me, but to me it appears that for a chunk of the population the available jobs do not pay well enough to meet a certain standard of living.
That is definitely true, but it's also pretty much the exact opposite of "always".
What? Can you explain the reasoning behind this statement?
Of course, that scenario happens only in extreme edge cases. Even in a booming economy with a shortage of workers, John Doe is not going to be pursued aggressively to fill the Senior Marketing Manager role.
This makes intuitive sense: employers have a ton of money, so people come to them.
Right now I have a client who needs to hire truck drivers and can't do it fast enough. I asked him what he'd done to make his company the most attractive (pay, technology, perks, etc.). He said he's done nothing.
That's exactly what recruiters and headhunters do. Aggressively nails it.
Most people will never be headhunted.
A candidate has multiple job offers, and decides which one to take.
The company interviews multiple people for a single position, rejecting the others.
There are only a small number of sectors where the first situation is reasonably possible, a lot of us on here are incredibly fortunate that engineering happens to be one of them. For the majority of the job market (by volume of people rather than volume of money), it takes people attempt after attempt to get a job. They don't get to choose between multiple offers, they have to take the first thing that will allow them to pay the rent, and then they have to hold on to it.
"are situations that people have arbitrarily (heuristically or conveniently) decided"
The Holocaust was not convenient, even to those who were for it. Slavery was convenient for those who benefited from it, but not those who suffered under it. Over the last 200 years racial categories lead to many millions of dead. This is not merely a matter of convenience.
Violence and coercion don't logically follow from racial differentiation. One may point out the differences between populations of different races, but that wouldn't justify attacking any individual from those populations.
My comment is framed inside that basic (and obvious) principle. Choosing a partner or an employee is not a violent or coercive act.
This applies to companies of any size.
Sure, some people in some countries have tried (sometimes successfully) to undermine this principle, but it's akin to forcing people to marry in groups.
Societal restrictions on contracts are a way to balance this. Virtually all societies have some form of this, outside utopian ultra-liberal hellholes.
Europe has the concept of indirect discrimination, which could make that illegal, certainly things less central to the role could amount to indirect discrimination.
This is likely due to an acknowledgment of the limits of human models to account for the full context surrounding correlated-but-non-causal classifications, such that conclusions drawn from them can have unforeseen or highly detrimental ramifications.
Speaking to race in America specifically, the schema through which we judge people are highly susceptible to bias from the white supremacist bent of historical education and general discourse. This is how you end up with cycles like those within the justice system (pushed in part by sentencing software), wherein black defendents are assumed to have a higher likelihood of re-offending, therefore increasing the likelihood of any given black defendent not receiving bond or having a lengthy sentence if convicted. After all, blackness correlates with recidivism. Lying outside this correlative relationship is the likely causal relationships of longer stays in jail and lack of access to employment opportunities, which disproportionately affect black people, causing higher rates of recidivism, regardless of race.
You can still have enhanced vigilance without enhanced annoyance and mistakes.
There's often a superior choice lurking that nobody is thinking about, sometimes expensive, sometimes not, seemingly unrelated to such optimization. This is why ML is not intelligence, it cannot find new solutions you're not already looking for.
The main problem of the judicial and police system is it tries to be procedurally fair and still fails at it anyway.
In fact, an AI trained on aggregate data to probabilistically infer characteristics about individuals is _literally a stereotyping machine_.
If people are upset that their stereotyping machine stereotypes people, they probably didn't fully understand what they were doing when they built it, because this is not a design flaw -- it's the design.
Then these people should argue against AI instead of swinging the racism cudgel.
More concrete: government AI shouldn't use things like names, zipcode demographics (at least those strongly to those characteristics we think discrimination = racism) and pictures of humans in their models. Why? Because it's pretty much impossible to control your model for racist tendencies one you start there. It's in the ethics of the creator of the model to point that out and just don't do it. If you do, and all people whose name start with an M (for Mohammed) get a different category, racist is the right term IMHO.
Quite a tautology.
However, had this been done to my daughter and the same result obtained (in fact the result was "blue" but there you go) by someone assessing her potential as a candidate for university - well there's damage.
And this is actually the rational thing to do. Reason: There are two potential errors involved here: labeling an innocent person a criminal (false positive) and labeling a criminal as innocent (false negative). The key is to realize that the cost of these two errors is not the same. For instance, treating an innocent person as a criminal could be much more expensive for the person and society than not detecting a criminal with a given classifier. For that very reason, we have the presumption of innocence as a principle in law. As a consequence, you don't want to select a classifier based on just the rate of errors overall, but you want to incorporate some kind of loss function that minimizes the cost for individuals and society. Under that loss function, the best classifier may actually be wrong more often than some other classifier.
They give everyone tests which reveal that primitive part of your brain. Essentially they want to shock you by forcing you to fail. It is almost like they are trying to shame people, which is a terrible way to teach.
There’s nothing to be ashamed of when it comes to having bias: women think other women are less competent than their male coworkers, trans women struggle with thinking of other trans women as men, black men find other black men threatening. Your first thought is the one you’ve been conditioned to think and is broadly speaking shared by everyone. It’s when you don’t stop to have the second thought — your own thought that’s the problem.
“No, that’s silly. I’ve seen her work — she absolutely knows what she’s talking about.”
“She’s a woman. Full stop.”
“That guy is just minding his own business and given zero actual signs of being a threat.”
Well there’s all kinds of mimicry in nature to exploit exactly this assumption.
You won’t hear much about it though because: 1) it’s a sensitive topic that is easily misunderstood by those who aren’t researchers 2) when everyone else is trying to scrub racism out of their models for politically correct reasons, having a “racist” AI can actually produce a competitive advantage in some industries, and it can be a difficult advantage for competitors to get similar performance if they don’t allow for natural racism to emerge.
In short, it’s not a big deal, humans have some amount of racism whether they admit it or not. What matters is that you treat people equally and without prejudice, regardless what you may think of their race. Judgements about a group and judgements about an individual are two different things.
Racist results don’t even need to be about negative things, it could be as simple as “a person of this race prefers this kind of food over that one”
If my classifier predicts that an Asian guy would probably eat rice, if it's accurate, is it racist?
The issue I'd have here isn't racism, it's when race dominates other features resulting in, in this case, poor recommendation quality/variety for Asians who prefer non-rice dishes or vice versa.
Absolutely. The mark of intelligence is to recognize them and correct for it.
And the people who just permanently slip up, and never correct themselves, and somehow always make the same snap judgments? 100% racist.
"And when one user uploaded an image of the Democratic presidential candidates Andrew Yang and Joe Biden, Yang who is Asian American, was tagged as “Buddhist” (he is not) while Biden was simply labeled as “grinner.”"
They are not removing images that categorize blacks as blacks. They are removing images that are incorrect.
Isn't that just a bug in the English language though? English very rarely employs variations of the same word based on gender - but that's not true of many other languages. If the labeling was done in say Polish or German suddenly it wouldn't be biased at all, a male doctor would be labeled Lekarz/Artz while a female one would be labeled Lekarka/Ärztin - it's just what it is, no bias here.
Racism starts out of preferential treatment of some people. Most racist people have a « root event », where the other party didn’t get condemned. It may even have happened several times, with various consequences (rape, molestation, repeat racket, etc).
Then they proceeded to denounce the rape/molestation/racket to the police. The police doesn’t act because they suspect you are racist. Which has the opposite effect than desired: It doesn’t condemn the criminal, and puts the burden on the victim.
Then they seek security in their lives. Therefore, if the probability is high to experience rape, molestation, racket or crime, enough that the racist has been confronted to it in his life, then the first best approximation of judging whether someone might be a criminal is whether he’s free or in jail. That’s in well-functioning societies. In non-functioning societies where criminals are not in jail, the second approximation to protect yourself is secondary indicators, inferred from grossly racist statistics. Here racism is born.
It also explains why racist people often still have various friends of the group they are supposed to mistrust. It’s because they were able to assess their probity and trustworthiness in some opportunity. That’s why mixing people by compulsory rules works. But mixing people is a poor palliative. It doesn’t solve the underlying problem if a non-functioning society, so it doesn’t make the people less racist, despite giving the appearance that people work together.
Racism is often used with regret by those who exert it. But it is the second best approximation to seek security. Racism is the result of a non-functioning society which is caused by being more lenient on criminals of chosen category. I’m pretty sure it is possible to engineer racism by letting a made-up group get away with crimes.
But you can find the same discourse in most famous racists of the world, at least currently living.
If I have to cite a source, Tommy Robinson (here comes the downvotes) constantly talks (and breaks) about the omerta surrounding the trials he reported on, which made that the rape gangs could proceed for so long with so few hindrances. Go through the list of people recently suppressed from Youtube and you’ll find the same (unadressed) logic.
Concerning root events, they are often private, so you have to know the person personally. But go ask around you when people became racist, they’ll often tell you one specific event. It’s interesting to go interviewing an enemy.
It this were the case, babies would be racist by default, but they are not. Racism is taught.
Our minds are literally categorical machines--in order to fight entropy we find stable states through classification of the world into categories. So literally any form of thought is a form of discrimination between many categories. Every word can be thought of as a category.
I mean it doesn’t seem like people tended to label some races with negative labels, just that they labeled their race at all, and didn’t do that for white people. At least in the example.
I don’t think it is useful to lump these two concepts together.
The people who curated the first training set used subjective words like 'attractive' to tag the images which means the AI tagged all images it deemed 'attractive' to the people who made the training set. as this is a very biased and homogenous group it means the AI turned out biased. Maybe if they randomly sampled millions of people from all countries in order to create the training data then they could effectively train an AI to guess what YOU might find attractive. However even then I somehow doubt it. Beauty is in the eye of the beholder. We dont consiously know the rules of what we are attracted to, nor does an AI have secret information if you just supplied it with enough words and images.
If they'd stuck with simple classifications like 'black' 'white' 'man' 'woman' then they would have less subjective judgement values about the original training set.
For example, it could have been interesting if they tagged each person with their actual religion or propensity to "grin"
However this was creating a dataset for classification. Something that specifically should not go beyond the facts. (The basis of the model is the strength of the facts it's built upon)
In my eyes, this result would reinforce unfair bias, and a thus well-designed AI should avoid this (i.e. with all else equal, a well-designed AI should suggest the label "software engineer" at the same rate for both men and women).
If the AI did start classifying masculine features biased towards software engineers, then the AI has learnt the above fact, and thus can be used to make predictions.
The moral standpoint that there shouldn't be more male software engineers than female engineers is a personal and subjective ideal, and if you lament bias, then why isn't this kind of bias given the same treatment?
The moral standpoint is that there shouldn't be an AICandidateFilter|HumanPrejudicialInterviewer that only coincidentally appears to beat a coin-flip because it has learned non-causal correlations which it uses to dust out qualified stereotype-defying human candidates because they don't look stereotypical enough on the axes that the dataset--which almost inevitably has a status-quo bias--suggests are relevant.
But if the task is say the pre-screening side. This becomes a more ethically/morally tricky question. If and only if that sex is not a predictive factor for engineer quality, you would then expect to see similar classifier performance for male / female samples. Given that assumption, significant (hah) divergence from equal performance would be something to correct.
Of course there are other issues to handle, such as the unbalanced state of the dataset and so on.
I.e. is the issue in the fact that the particular annotators were subjective and annotated some particular facts wrong (and the labels for skin color could be filled from, for example, census data which is self-reported) or that these whole type of labels shouldn't be attempted to be made as they're not facts?
If the latter, what do you think about the categories like "adult" or "sports car" that are also part of ImageNet; can we draw an unambiguous factual boundary between images of adults and teenagers, or "normal" cars and sports cars?
No, it's not "The dataset is biased, the AI isn't" it's "The dataset is biased, THEREFORE the AI is biased".
I'm not talking specifically about cultural biases like racial stereotypes etc. Confirmation bias is a thing, there's nothing stopping a researcher from making those observations that confirm their favoured theory and contradict all others.
Then of course there is sampling error. Just because you have a set of data that you collected "at random" doesn't mean that this dataset is representative of the population you are interested in. Let alone the fact that it's very hard to collect a truly random set of observations about processes that we don't understand to begin with.
The kind of data you're describing is an ideal, a principle that we all aspire to. It's far from the reality in practice.
well seems to be very simple - either there is a hat or there isn't
https://twitter.com/Dk3Kbball/status/1174115660219072512
That illustrates the depth of the issue - while directly racist data can possibly be removed there are many proxy/correlated attributes (as any insurance/mortgage/etc company knows), and to find correlations is the core nature of the machine learning systems (at least the ML as it is currently known to humans).
https://www.buzzfeed.com/ashleyperez/global-beauty-standards
"Make me beautiful," she said, hoping to bring to light how standards of beauty differ across various cultures"
She specifically asked for some kind of alteration, so leaving the picture unaltered was implicitly not an option. Furthermore, this assumes a random (but distributed) selection of Fivver artists represent worldwide beauty standards.
This is pretty close to "does the Chinese room know Chinese". https://en.wikipedia.org/wiki/Chinese_room Going too deep down that path is fun for philosophy, but not super useful...
The definition of beauty varies across culture, and while it varies from person to person, there are some aggregates that might be informative.
Interestingly, I was just thinking about the "black and white" issue today. A while ago I watched an interview on Youtube with a young woman who was born in Japan. Her parents were American and growing up she new she was different, but basically it boiled down to "the English teacher's daughter", or "part of the foreigner family". But she didn't speak English very well and her parents didn't speak Japanese very well, so she identified very much more strongly with her friends in the area than her family.
When she was 12, her family moved back to the US. When she went to school the new children were encouraged to say what nationality they were. One person said they were Canadian. One said they were Mexican. When it came to her turn she said, "I'm Japanese". One child in her class corrected her: "No you aren't. You are black." Of course, this was a source of considerable confusion for her.
One of the things that's kind of weird about Japan is that some Japanese people have very dark complexions -- darker than what would be called "black" in the US. Some have very light complexions. In my opinion, considerably more "white" than I am (and most people would call me "white" I think). In fact, after I moved to Japan, I realised that I wasn't white at all. I'm pink. I mean, I'm super pink. I seriously never noticed it until I spent 5 or so years living in a place where nobody else was pink. I literally avoid wearing red now because it makes me look like a tomato.
When I was about 3, my best friend was black. My grandmother asked me, "Do you notice anything different about your fried? Their hair or something?" I didn't understand the question. All of my friends had different hair. Then she said, "Can you tell that he is black and you are white? Or do you not think about it that way?" My grandmother was just curious, but this question completely blew me away. From that time, I realised that people didn't each have their own skin color. Instead, they were categorised and my friend was different than me. I think my friend noticed that I looked at him differently (though not necessarily badly). We stopped being friends for a long time. Somewhat strangely, I just recently realised that he was one of my best friends in high school... I never actually made the connection that my friend at 3 and my friend in high school were the same person until recently. I wonder if he ever realised.
But anyway, there really isn't a classification like "black" and "white". When I have a tan, my skin is darker than my wife's. But she is very tan and so her skin looks brown. When I compare my skin to hers, my skin is still red, even though it is dark. But if I were to compare my skin to an indiginous American, I think their skin will be more red than mine when tan -- and less pink when not. And my wife, when compared to someone indiginous from Africa has more a more of a curry brown than a deep chocolate brown.
As I was saying, Japanese people don't call dark Japanese people "black". They don't treat them as a different race. Neither are very pale Japanese people "white", even though they may say "your skin is very white". They aren't different. What I found interesting about the young girl who discovered that she was black was that she wasn't "black" in Japan. Or, at least, not "black" in the way we use the term in America. Japan does not have that cultural history (it's got enough of it's own baggage, thank you very much ;-) ).
If a computer were to compare skin tones objectively, it would simply tell you the color. If it decided to classify in terms of "black" and "white", it would be classifying based on cultural labels, not color.
It would be based on a larger selections of individuals (non-local), based purely on appearance, and un-biased by human conception (we are really good at processing faces/facial features).
"Stunner"? "Looker"? "Mantrap"? Or even trying to tag people's images with categories like Buddhist, grinner, or microeconomist?
What were they thinking?! Clearly these tags were never curated in any remotely responsible way -- for quality, for sensitivity, or just usefulness at all -- and I'm shocked they were ever intended for academic or research use. No wonder image recognition gets a bad rap, with input data like this.
They may be, however, labels generated by users from a context.
Our culture is broken and nobody is willing to stand up to insanity or even admit to it. People are too happy to do completely insane and senseless things if it saves them from the social lynchings.
In addition to their statement, linked in a thread below [0], what they were thinking is pretty out in the open in the abstract of the original paper:
"We introduce here a new database called “ImageNet”, a largescale ontology of images built upon the backbone of the WordNet structure. ImageNet aims to populate the majority of the 80,000 synsets of WordNet with an average of 500-1000 clean and full resolution images. " [1]
That paper is now 10 years old, WordNet itself had it's first release somewhere in the 1980s. A lot happens in research a decade down the road and lots of research towards biases of datasets are relatively recent in that sense. It can hardly be blamed on the researchers that people use their datasets in production merely because "they're large and there in the open". As their statement indicates they're actively working on the issues that arose during the last few years.
Most research datasets are fit for a very specific purpose, to me the larger problem here is two-fold: "AI" fearmongering on one side and startups/corporations on the other side lacking the necessary skills to select and filter their training data. Hopefully that'll be a matter of the past once the hyped data science field reaches maturity.
edit: Wordnet seems to include said terms to this day, [2] for example still lists "mantrap" associated to "beaty" in the synset for "a very attractive or seductive looking woman". Although I'm unsure how they should handle these cases, both are words and that's one definition used in practice, even if we find it reprehensible. They seem to have removed the n-word but that's about the only removed instance I can find in a cursory search of sensitive or insulting English words. Maybe they should clarify sensitive annotations like they do with vulgarity.
[0] https://news.ycombinator.com/item?id=21054770
This is a photo in question: https://memepedia.ru/wp-content/uploads/2019/09/imagenet-1.p...
This is the route the network took:
person, individual, someone, somebody, mortal, soul (6978) > female, female person (150) > woman, adult female (129) > smasher, stunner, knockout, beauty, ravisher, sweetheart, peach, lulu, looker, mantrap, dish (0)
So it was (politically) correct on the first three categories, and the last one was either a crapshoot (and she could also have gotten to the subcategory of "prostitute" > "streetwalker, street girl, hooker, hustler, floozy, floozie, slattern") or she really is posing in a common "beautiful woman"-way. (The global description for this route is "A very attractive or seductive looking woman" and often triggers for females with tilted heads and lip curls).
You can turn any faces dataset into a labeled face color dataset, so if a black person being subclassified as "negro" is problematic bias or encoded racism, then all such datasets are suspect. Noisy labeled data is the norm, not some horrible exception to be avoided at all costs.
It is the technical justification that should be all that matters for a canonical academic dataset. Science does its best to be apolitical, but then politics ("red bull drinking white men train racist and sexist classifiers") is forced upon it, and we can't really have a productive conversation about bias and ethics anymore.
AI needs common sense knowledge of the world to improve. Censorship so science does not offend our sensibilities, would only make it so Google Image Search (a machine learning algorithm) does not return any images of people when you search for "prostitute". Heck, the AI would never learn the difference between a male and a female prostitute. Destruction of accessible knowledge so we (aka: people on Twitter who think AI is the terminator, or the director of the internet) don't get offended by some primitive ML model-as-art-project forced to make errors or awkward classifications. That's a sentence the academic ML community could do entirely without. No benchmarks or duckface selfie would be hurt. No unfortunate third-world souls hired to scan 20 million + internet crawled images for wrongthink, only for the machine to do the unsupervised learning in a hug box, not a black box. Oi mate, you got a loicense fer that label?
Just wait until the activists find out who wrote the first 100 8's added to MNIST. Nobody but MIT would be associated with her, if they found out what she did.
I guess the problem with things like "microeconomist" are that other tags such as "fireman" might be valid (if a picture contained clues about vocation, for example) - furthermore, rather than condemn the label, we might instead consider not demonising a neural net, and taking its results at face value; in this case an indication that most labelled micro-economists in the dataset are suited white men.
> Each synset is classified as either “unsafe” (offensive regardless of context), “sensitive” (offensive depending on context), or “safe”. “Unsafe” synsets are inherently offensive, such as those with profanity, those that correspond to racial or gender slurs, or those that correspond to negative characterizations of people (for example, racist). “Sensitive” synsets are not inherently offensive, but they may cause offense when applied inappropriately, such as the classification of people based on sexual orientation and religion.
> So far out of 2,832 synsets within the person subtree we’ve identified 438 “unsafe” synsets and 1,155 “sensitive” synsets. The remaining 1,239 synsets are temporarily deemed “safe.”
They've also completely disabled Imagenet downloads while they remedy this.
I need to find an ImageNet archive now.
I think we’ll get there at some point, and I’m not exposed to the most cutting edge AI research, but it seems like AI us currently very overhyped and deeply flawed for many of the applications people would like to use it for.
To achieve what you want, semantic structure must be used as labels instead of just categorical labels.
Assuming we have a sane AI that now knows its looking for cancer, it knows what that means (from digesting medical textbooks, papers and generic text corpora) and it can detect rulers and knows the two are not casually linked from ruler to cancer, we could make the model output "dataset diagnostics", like a "Warning! The cancer label in this dataset is implausibly correlated with the visual presence of a ruler". Or "Warning: 99% of your hotdog images show a human hand. Evaluation on this dataset will ignore errors on hotdog images without hands!"
Context does matter though. If there's an orange fluff on a tree trunk, the AI is right to look at the environment and infer it's a squirrel.
And between the Dem/trump back-and-forth, racial paranoia is strong in America.
The hidden motive is "what if ImageNet is used to (approve loans|decide parole|pre-populate killer police drone biases)!", but the way this is done damning any emerging tech that is less that perfect.
If I heard loaded handguns where given to toddlers, I'd attack the practice of giving handguns to toddlers - not the existence of toddlers!
That said, between tech hubris, "self-regulation", and the toxic partisan culture war, I'm not sure what could change. A real AI guideline, rather than vague pop-sci-fi document, or tech policing (dictate appropriate usage only)? preferably incubated in a country where nuanced discussion can be had without the flamewars...
Edit: Might be this, which was a week ago / before the Roulette blew up: http://image-net.org/update-sep-17-2019
> We are in the process of preparing a new version of ImageNet by removing all the synsets identified as “unsafe” and “sensitive” along with their associated images. This will result in the removal of 600,040 images, leaving 577,244 images in the remaining “safe” person synsets.
If the ImageNet contains bias that leads to embarrassing results- that’s fine! That give us a readily available toy instance of the problem to study. Taking that away could actively harm anti-bias research.
I don't want the appearance of fairness (introduced by human dataset curation) to be mistaken for "intelligence".
Keeping the bias would hopefully cause people to think more critically about why such bias exists in the model in the first place.
I'm just saying that model bias is a very easy thing to explain (usually, data imbalance).
You can also fiddle with the label ratios to change the bias - which is also a good way of showing that the models aren't really intelligent.
If a data set is flawed, it should be fixed, but when ML finds objective patterns our culture finds subjectively unpalatable and we choose to "fix" them, we fall prey to and re-enact the same grade of self-delusion exhibited by, for example, "the church" in the dark ages. Computer science is already low in terms of accountability & rigor compared to other fields without these kinds of suggestions.
This isn't a thing and you should read more about what you think happened in this period of western history.
For example: black men in the US are more likely than white men to have criminal records, but this in no way means that black men are "objectively" more criminal.
To paraphrase Warren Buffet: "If a cop follows you for 500 miles, they're going to find a reason to give you a ticket."
If your life is a cesspit and you can count on the authorities to be part of the problem, where does murder end and self defense begin?
See the song "I shot the sheriff." I'm most familiar with the Eric Clapton version, but googling it recently suggests to me it was originally written by Bob Marley.
There are plenty of dead bodies with holes in their heads, but it's not clear how one accounts for crime among those 7800 dead.
How does overpolicing explain both arrest for murder rate and victim of murder rate have a significantly higher prevalence within black community?
I have a feeling most of you went to white schools or lived in white areas instead of black ones in bad places. Most in my area that went to the latter, including black folks, know where this comes from. I already described it here:
https://news.ycombinator.com/item?id=20660278
You folks need to stop just dismissing all blame on the black side as "cops must just always make stuff up or police them harder." There's plenty of that worth calling out. However, the biggest hole in your argument is murder in high numbers. You can watch folks that don't kill people or who get shot all day in whatever selective way you want. In the end for poor neighborhoods, the black ones you watch will have committed many more murders hitting many more black victims than the white ones.
And, if gangs are involved, the dealers and killers will often be told they're in for life. They'll be perpetuating it in a way that has nothing to do with white people. And you should be calling them out hard on that if you really care about thugs getting arrested by cops and/or protecting their victims which are mostly black. Instead, it's whites (damage of some kind) blacks non-stop in your comments or the media vs thug culture from blacks murdering black people all the time. And, it seems, something like that with machismo and cartels in Latino areas. I have less experience with them, though. I didn't get to since we had to leave the neighborhood after a small confrontation led to one putting a contract on our heads that canceled (maybe) if we moved.
So, I'm calling bullshit unless you're saying there's both racist over-policing by cops (mostly but not all white) and thug culture coming from blacks creating killers that take out tons of blacks blacks. Something similar for Latino areas. Then, it fits the data.
They claim this is AI, but earlier in the article, it states that mechanical turk was used. Mechanical Turk is basically just people getting paid pennies or fractions of a penny to tag these photos.
Many people on Mechanical Turk are from countries where English is not their first language and they don't have the same idea of racism as we do here in the US, which would explain the racist tags mentioned in the article.
This doesn't really show us anything about AI and racial bias and more that other countries still aren't up to our levels of what we consider decent.
I just hope the masses don't glom onto this, start shrieking that "AI is racist", and attempt to take AI completely out of the picture.
This is a really bizarre project. I had seen some really offensive race-based labels, but I thought revealing the ugliness of the system was part of the point of this project?
But besides that, the results just seemed completely scattershot; I half expect the artists/exhibitors to reveal that x% of the results were randomized. Last week, I tried it myself after seeing another Asian user display results that were entirely Asian slurs (e.g. gook, slant-eyed). I uploaded my own very Asian-looking photo and got "prophetess", along with very vague labels, such as "person" and "individual".
Maybe the exhibitors cleaned the data/results by the time I tried it, but I used it just a few hours after seeing the other Asian user's results, so I'm doubtful that her tweet/complaints were enough on their own to change up the dataset that same day.
It 100% is. From the link to the artist's website:
"Things get strange: A photograph of a woman smiling in a bikini is labeled a “slattern, slut, slovenly woman, trollop.” A young man drinking beer is categorized as an “alcoholic, alky, dipsomaniac, boozer, lush, soaker, souse.” A child wearing sunglasses is classified as a “failure, loser, non-starter, unsuccessful person.” You’re looking at the “person” category in a dataset called ImageNet, one of the most widely used training sets for machine learning."
In it there's a problematic feature tagged simply as "black" and it is defined as the proportion of blacks by town.
Any pricing model that is built off of this dataset is inherently racially biased because the data has been collected and the feature tagged - but what's the alternative? Not to collect the information? Or collect it but completely ignore this feature?
For example, a medical screening NN may find race to be a valuable feature for the prediction of illness; but a health insurance assessor should not.
If your data set was compiled by thousands of mechanical turk workers, well, you got a lot of people to blame a little bit. As it goes, "everybody is a little racist sometimes..." and apparently that shows on a big enough data set.
Guess who classified the blonde white woman "a dish" on a Mechanical Turk? Do you think it was that Billy Bob from the swamp of Louisianan who makes $12/hour? Because $12/hour is impossible to make classifying images. Or do you think it was a dirt poor Indonesian or Filipino or Chinese or Indian for whom that's the best job he or she could possibly get and that's exactly how they view the world and there are over three billion of them?
Sincerely, a Malaysian Chinese who’s seeing this slang for the first time.
I would guess the classifiers went to school in India or Pakistan.
If faces are like other images in the database, then according to this article turkers were presented a word and a group of images, and asked to click all the images that showed said word. https://qz.com/1034972/the-data-that-changed-the-direction-o...
from the article "Jamal Jordan, a journalist at the New York Times, explained on Twitter that each of his uploaded photographs returned tags like “Black, Black African, Negroid, or Negro.”"
Which indicates that the vocabulary they used has 4 words to describe a black person's skin color. This severely complicates things and leaves all kinds of unintended biases on the table before even getting to the human element. Though I doubt they used a dictionary at all because 'Black African' isn't really a word. And not found in a few dictionaries I tried online. - if the article is accurate.
They should have curated the vocabulary first, value judgements like 'attractive' should be removed because it means something different for everyone judging the images. Synonyms should be collapsed into one to prevent weighted biases, etc.
Though value judgements should not be removed. They should have been separated and placed together with some tagger metadata as they might actually be informative then.
I wonder if similar things can be done to address specific (i.e. racial or gender) biases in computer vision.
Doctor - Man + Woman = ?
What normally comes out is Nurse. What "they" think should come out is Doctor!
By "they" I mean people that get upset by this.
[0] https://arxiv.org/abs/1607.06520 "Man is to Computer Programmer as Woman is to Homemaker? Debiasing Word Embeddings "
[1] http://proceedings.mlr.press/v97/brunet19a/brunet19a.pdf "Understanding the Origins of Bias in Word Embeddings"
[2] http://matthewkenney.site/biases.html "Google word2vec biases"
An image of a "bad person" wasn't tagged by someone looking at said image and deciding that "bad person" was the best possible description. It was generated by searching Google Images for "bad person" and removing obviously incorrect results (e.g. when there's nobody in the image).
Researchers have been using it to learn the inverse mapping from images to tags with some success, but in its construction the dataset is not naturally suited for that task.
Also, it doesn’t seem like there has been any evil intent here, nor did I see anything about pernicious consequences. It seems a bit overblown to accuse people of racism over something that is unintended and theoretical. Just update the tags.
If it was trained to identify “human” or otherwise categorize the things in the picture, that’s exactly what it would output.
Next you can train it to attempt to guess basic traits like gender or ethnicity, of course this can only be done based on the RGB values of a 2D array of pixels. Interestingly the NN will not be merely be using skin color but building probabilistic weightings based on any statistically significant features. For added controversy, it’s probably even possible to invert parts of the network to suss out how it’s weighting various facial features toward different labeling.
Lastly the images could be labeled for things like profession of the subject. A good intentioned effort to perhaps detect things like a lab coat could mean a doctor or scientist was followed by oblivious Turkers with predictably poor results.
The problem of course is not with the images but the particular labeling hierarchy, and allowing opinionated labels as well as labels which may be true for the particular subject but which bear no distinguishing features for which to actually codify (“portrait of a macroeconomist”), in other words, garbage in garbage out. ImageNet calls this the “imageability” of the label.
Calling this racism of course is entirely inapt, because there is no judgement being made whatsoever. Even if the sampling method was well designed and the labeling factually accurate, the system would still produce output which could be considered offensive. Again, because the entire point of the algorithm is to generate statistical assumptions based on a single image.
My conclusion is that some superficial judgments are algorithmically useful and hopefully less controversial. “White male human, ~55 yrs old, 180lbs”. Even things like analyzing clothing and guessing where the picture was taken. Iff the clothing is a uniform, identifying the profession (police, fire, paramedic)
But you have to know where this goes off the rails. Bad enough to label indistinct portraits with the subject’s profession, let’s not do inane things like labeling them with how subjectively attractive the labeler thinks the person is, their economic status, maybe even a 1-10 scale of how threatening they look or if they look like a criminal or not! </facepalm>
Apparently i'm either an insurance agent (in a grey t-shirt) or a surgeon (in a red hoodie).
Odd that the images were removed for being categorized as the project intended.
An AI being accused of bias tends to really mean it works. Removing bias from AI, I've noticed, requires hardcoded 'fixes' rather than refactoring the algorithms. And in my view, becomes yet another human-curated classification system, and no longer AI.
The problem is not that the AI is "inaccurate". The problem is the second order effect: when you build systems on top of this AI's predictions, you cement the input social biases into future systems in a way that is very hard to remediate.
The real problem is how to avoid accidentally creating systems that amplify existing social biases.
I'm very specifically using the term "social bias" to distinguish from "bias" as a term in ML, because they are very different problems.
https://www.youtube.com/watch?v=HEX7xsYF1nA
(And now I'm gonna be listening to 90s electro-disco all day...)
"We have too many pictures of white people, remove them!"
"We don't have enough pictures of non-white people, add them!"
I'd have gone for the latter, and have let the set be biased until fixed.
Which is fine, I have a beard, But then I read the definition:
> beard: a person who diverts suspicion from someone (especially a woman who accompanies a male homosexual in order to conceal his homosexuality)
:-/
For example Search for images on scuba diver https://upload.wikimedia.org/wikipedia/commons/9/94/Buzo.jpg is labeled choreographer https://dtmag.com/wp-content/uploads/2015/03/scuba-diver-105... is labeled picador (the horseman who pricks the bull with a lance early in the bullfight to goad the bull and to make it keep its head low)
How about search for images of dancer https://eugeneballet.org/wp-content/uploads/2018/10/Alessand... is a nonsmoker https://www.ballet.org.uk/wp-content/uploads/2017/09/ENB_Eme... is speedskater https://www.ballet.org.uk/wp-content/uploads/2018/10/WEB-ENB... is plyer https://rachelneville.com/wp-content/uploads/2018/11/10.14.1... is a mediatrix (a woman who is a mediator)
How about images of lumberjack https://alaskashoreexcursions.com/media/ecom/prodxl/Lumberja... is a skinhead https://static.tvtropes.org/pmwiki/pub/images/lumberjack_591... beard: a person who diverts suspicion from someone (especially a woman who accompanies a male homosexual in order to conceal his homosexuality) http://cdn.shopify.com/s/files/1/0234/5963/products/I4A0032-... is a flight attendant https://previews.123rf.com/images/rasstock/rasstock1411/rass... is an asserter, declarer, affirmer, asseverator, avower: someone who claims to speak the truth
How about teacher https://media.edutopia.org/styles/responsive_2880px_16x9/s3/... is a shot putter: an athlete who competes in the shot put https://c0.dq1.me/uploads/article/54231/student-classroom-te... The kid is a non smoker and the teacher is psycholinguist: a person (usually a psychologist but sometimes a linguist) who studies the psychological basis of human language https://media.gannett-cdn.com/29906170001/29906170001_578035... is girl, miss, missy, young lady, young woman, fille: a young woman https://media.self.com/photos/5aa9743e19b7c01d73149d50/4:3/w... (almost the same picture as the prveious one ) is now a sociologist!
How about search for pilot images https://news.delta.com/sites/default/files/Propel%20Embedded... is a parrot: a copycat who does not understand the words or acts being imitated https://pilotpatrick.com/wp-content/uploads/2017/10/workday_... is a beekeeper, apiarist, apiculturist: a farmer who keeps bees for their honey !!! https://imagesvc.meredithcorp.io/v3/mm/image?url=https%3A%2F... is a boatbuilder !!
And some random one https://media.spiked-online.com/website/images/2019/08/06154... is a "Sister/nun" https://www.english-heritage.org.uk/siteassets/home/visit/in... deacon, Protestant deacon: a Protestant layman who assists the minister https://minervasowls.org/wp-content/uploads/2018/04/Romans-M... are identified as morris dancer
Gödel arbitrage will be very profitable in the future.
this AI seems to be doing that to things it has determined are black people. A bunch of synonyms for black people, some English, some Spanish, some phenotypes, some of those descriptions simultaneously fallen out of favor in some parts of the world but not others, while completely eliminating other metadata.
I'm not sure Imagenet's response of removing "sensitive" adjectives is capable of fixing this. Mechanical Turk: english, spanish, academics, using terms that aren't universally agreed upon?
That doesn't really address what is happening