Problems in Gemini's Approach to Diversity
noahpinion.blog
noahpinion.blog
Google to pause Gemini image generation of people after issues - https://news.ycombinator.com/item?id=39465250 - Feb 2024 (1164 comments)
What does "no more racism" look like? To me, it looks like everyone having recognized that skin color isn't an intrinsically meaningful part of someone's personhood. The people who believe this should fight against the people who do not.
We're not going to get there by becoming ever more exquisitely concerned with skin color and it's implications in every social interaction.
If we can take a cue from Gemini, "no more racism" looks like "all people have skin color in the exact same shade of brown".
It really was a drug for low confidence people to feel better about themselves compared to others, even while the ideas state that you need to self-reflect instead of being indignant.
Humanism gets rid of racism while DEI keeps it alive in perpetuity.
What I've said is bare, basic truth, and not ideological. If it offends or provokes factual or semantic quibbling, then such reaction indicates pretzel thinking about moral good and 'racism' in the minds of people like Smith.
It seems like a technical glitch to me. Similarly here, they want the models to represent ethnically diverse individuals, but the model doesn't understand historical context well. It's a technical shortcoming.
No, because that occurred due to an oversight of not including a representative set of training data. At worst it's implicit racism due to oversight and ignorance.
What Google seemed to do here was more explicit racism, by setting configuration in their system that directly and intentionally discriminates based on race. It's fundamentally worse.
Google, in a self imposed noble desire to rectify poor representation of training data, intentionally injected race discrimination in their system. This intent is what makes it fundamentally worse than unintentionally having non diverse training data.
Nobody who worked on facial recognition that failed to recognize black people, nobody in that team was intentionally trying to fail on black people. They didn't try to shoot an arrow, although they accidentally shot one in their own leg.
See how these things are different?
When racist employees of Google intentionally racially cleanse their products, this is racism.
Let me know what is confusing about the concepts of agency and intention., and I will try to clarify
"I want a model that produces racially diverse images of people, but I didn't test it in racially homogenous historical contexts" => yes-racism, according to you.
But probably not simple enough to pierce ideological bias. I'll ask instead: you really believe that it wasn't tested in historical and other contexts where white people were erased?
Occam's Razor: Of course it wasn't or they wouldn't have released such a ridiculous embarrassing product! C'mon, I think you're the one wearing blinders here.
Then please do explain why their first reaction was to double down on "diversity" in historical contexts, before eventually retreating?
Have you never tried tweaking the error text first before giving up and redesigning your form that's confusing users?
Of a piece with 'Google more or less explicitly said I won't be promoted because I'm White' stories.
You are again misrepresenting the OP. If you go actually read the OP you will find it very clearly describes Google's initial steps towards fixing the problem, which was: double down on race-swapping in historical contexts. The OP very clearly (and correctly) argues that this is not the action of a person who goes "oops forgot to test in this context" but is rather the action of a person who very intentionally wanted the image generator to behave this way in this context.
So, are the Caucasian people who were involved in any of these miss-optimizations, aka mistakes, racist against themselves? Are they traitors?
But it could also just reflect the composition of the population of the countries that dominate the internet.
A piece of facial recognition software that doesn't do well with dark skin shows that the team that built it either didn't think to test it on dark skin or didn't feel it was important to get dark skin working before release.
An image generation model that both doesn't produce images of white people spontaneously and will actively refuse to do so when asked shows that the team that built it either failed to ever ask it to produce images of white people or believed that it was a good thing for their model to refuse to produce such images.
In neither case is the problem purely technical.
Right now, the DEI exists in these models at a level where context is not considered. Not for time or locale, nor for, say, the correlation between sex and different populations present in a given profession. It's just an unconsidered DEI goo sprayed over the outputs.
The people behind Gemini's "diversity" sincerely believe they're doing the right thing. Calling them racist doesn't help change their mind, it just gets you filed under "reactionary" and ignored. This piece takes a more effective approach by saying "I'm with you, I think you want to do the right thing, but here's why what you're doing is actually counterproductive".
It's an effective rhetorical strategy for the target audience, you're just not that audience because you already see the problem.
Of course everyone has good intentions and likes to be flattered. It's possible to flatter with tone, and without dissembling.
I wasn't trying to say that he was lying for the sake of the article, I was trying to say that he's framing it in a way that is intended to be persuasive to a particular group of people that you're not a part of. If you pressed him on it in a private conversation I wouldn't be surprised if he acknowledged that the behaviors involved here are racist (by some definition) even though the people are trying to be the opposite. But he knows better than to level that word at his audience.
I would tell him that however he justifies concealing or dressing up his beliefs, such dissembling really does have the effect of increasing confusion and rancor - he is, of course, also a Good Person, but this is not a good tactic.
> diverse [...] which the AI interpreted as meaning “nonwhite”
While an AI isn't objective in any way, it is the result of the dataset. This is also an accusation against the DEI groups that grant themselves the title of being less biased or just self-reflective enough. Problem is that they simply are not, on the contrary.
What an amazing self own, laughter and mockery is the best cure here, and possibly a news article about how a bunch of people on the project got fired.
I was unaware that Native Americans were in 1850's China. The only reason I knew they were Native American was the fact that the image used literally the most stereotypical image possible for them, moccasins and feathers and all.
Amazing. Handed nukes to everyone who claim that the "liberal elites" have it out for them.
Not for ‘reverse’ racism. A pay rise, perhaps.
Even aside from the image generation snafu, Gemini guardrails make it sufficiently annoying to use for even innocuous queries to the point where I pay for a ChatGPT subscription even though I have access to Advanced. This is really an unforced error on Google's part.
> But Google’s attempt failed disastrously. Why? In my view, it was because the Google team tried to take a shortcut.
Rather than do the hard work of building a more integrated society, the attempt at superficial shortcuts leads to lazy things like the prompt injection seen in the Gemini release.
Another great comparison toward the end:
> Where Hamilton [the musical] challenges the viewer to imagine America’s founders as Latino, Black, and Asian, Gemini commands the user to forget that British monarchs were White. One invites you to suspend disbelief, while the other orders you to accept a lie.
This redefinition of diversity—to a separate monoculture—has underwritten the failure of the movement. Politically, obviously, and I’d also argue morally.
Here’s one with 1200 comments: https://news.ycombinator.com/item?id=39465250
TL;DR: there are no shortcuts.
As an centrist, I worry as much about the MAGA people as I do the Woke mob. But I'm starting to think that there is an overreaction currently with regards to the degree to which people claim that Gemini is woke and biased. If I don't intentionally try to get it to say racist stuff, I find the answers rather nuanced. Sure, I sense the "left" bias in the model, but I'm starting to think that there is a stronger echo chamber with regards to "Google is a woke mind virus spreader" than I thought.
I thought this answer to my question on individual differences to be rather okay.
Question: Are there fundamental differences between individuals with regards to their ability to be productive that might be a problem with regards to liberal views, and is there any research on this topic that informs our understanding?
Gemini Advanced Answer (Abbreviated, full chat [here](https://g.co/gemini/share/1344739f31b2))
*Fundamental Differences in Productivity*
* *Cognitive Abilities:* Intelligence, problem-solving skills, and learning capacity vary among individuals. These influence a person's ease of acquiring knowledge and skills relevant to productivity. * *Personality Traits:* Conscientiousness, self-discipline, and motivation are strong predictors of work ethic and success. These traits differ between individuals due to both nature and nurture. * *Mental and Physical Health:* Mental well-being and physical health significantly impact one's focus, energy, and overall capacity for productive work. These are influenced by genetics, environment, and access to care. * *Socioeconomic Background:* Access to education, resources, mentorship, and networks greatly influences the opportunities and skills available to individuals, impacting their potential productivity.
Given the pervasive claim that Woke people minimize the role of IQ and abilities, and claim that all differences are a result of oppression - this model isn't very Woke. Are people just gerrymandering the results to get it to look bad? Have I just not found the bad parts yet or is the Woke claims overstated?
If you genuinely want to find "the bad parts", it isn't particularly difficult to do. Just ask <sensitive question about <identity>> by keeping the question the same and swapping the <identity> part. An uncensored LLM would answer these questions consistently. A woke-censored LLM will answer some questions notably different than others (e.g. refuse to praise whiteness, while providing glowing adoration for blackness).
Do you remember the fiasco that was Tai AI? Do you think letting it run "uncensored", aka turn full on Nazi/racist, was a good idea, or something any company would be okay with?
Humans are quite adept at self-censoring when relevant. An AI model needs to be taught the same thing. And heck, even humans can be racist, it's just that most humans are reasonably moderate.
(For context: https://www.theverge.com/2016/3/24/11297050/tay-microsoft-ch...)
Another thought I had: what about times when LLMs suggest something dangerous? I may call it instructing the model to not do dangerous things, you may call it censorship. For example, when a model suggested making chlorine gas- (https://arstechnica.com/information-technology/2023/08/ai-po...).
Do you think such outputs are good? The vast majority of products in everyday life have warnings or failsafes, from cars that chime when the seatbelt isn't put on to all sorts of warnings on allergens.
Allowing a completely "free" LLM has risks, and honestly even if all HN'ers agreed to not censor it it likely wouldn't matter as companies would rather censor than have liabilities.
edit: I see you edited in the chlorine gas example. Do I think a cooking chat app should accidentally recommend a recipe that recommends chlorine gas? No, the cooking app makers should have the ability to control the LLM so it doesn't do that. Do I think the cooking app makers should be prevented from making these decisions, and instead have those decisions forced on them by the base LLM model? No. I think app developers should have control over how to censor the outputs. The base LLM should be uncensored.
Companies that create products on top of LLMs absolutely should have the power to censor the LLM as they want, and they should add a lot of censorship to support their use cases. They should not be prevented from making these decisions (by applying censorship at the base LLM level you prevent companies from doing these decisions).
(Ideally I would love the base llm to also be truly open as well, but that's another issue.)
Unless you tested it on or before the evening of Feb 21st, before they started tweaking it in response to the bad press, then your results are not valid at representing what it was like when they first rolled it out, because you were using a system that had been changed.
I was equally dubious that it could be anywhere near as bad as what I was seeing online and on Twitter, so I tried it myself. To my surprise, it was exactly as bad as what I saw. I was able to reproduce almost exactly what people were sharing by using the same prompts (obviously the images were different cause AI, but it followed the same pattern).
I’m not so sure about this bit of analysis. It falls in line with the (somewhat conspiratorial) view that these organizations have some kind of woke agenda that they want to push to the masses. A far simpler explanation is that the managers have no personal agenda other than to mitigate the risk of public meltdowns like Tay [0], which now serve as PR case studies in what can happen when these types of systems are open to the public. It’s not a matter of some manager trying to erase white people—just a sloppy attempt to mitigate the risk of a flood of articles about “Google’s New Racist AI”. Since we don’t have technical solutions to these problems, however, we just swing the pendulum in the other direction; hence, the current flood of articles about “Google’s New Reverse-Racist AI”.
You can’t win, really. As soon as you put something like this into the public, you have people doing ideological pen-testing from every angle. And there is a dearth of technical solutions to this problem.
Couple that with the fact that someone looking to be offended will always find something offensive and, as you say, you can’t win.
Isn’t this current affair a public meltdown?
At this point, anyone who doesn’t think that Netflix and Google aren’t involved in a Nazi-level propaganda is either misinformed or part of the plan to erase White people from History.
All of the kids know.
There’s been countless memes about Google rewriting history with falsehoods, and Bing branding itself around giving the straight, dark answer. Everyone’s sharing screenshots of Google’s absolutely blank homepage on November 19th. Everyone’s sharing those pics of Google today, whether or not they are true, Google is getting a bad rep escalating by the day.
If given the choice between being criticized either way, wouldn’t it better to be on the side of the truth?
Google has chosen the alternative path.
Could you expand more on this?
> All of the kids know.
Is that just an expression like "lots of people know" or "the leading people know" or do you literally mean younger people in large groups "know"?
> Isn’t this current affair a public meltdown?
From my perspective it's not. It barely made it to "major" publications, and got flagged and suppressed heavily on HN during all the early news. I also don't know a single non-tech person who even knows about this story. I definitely wouldn't call that a public meltdown.
It certainly is, which is why I mention the swing of the pendulum in the other direction (shortly after your quoted sentence) :)
It’s not “let’s re-vet all the training inputs to reflect the just society of our dreams,” it’s “yeah, the AI Safety people say make it diverse, just put ‘make it diverse’ in the prompt.” Which, I mean, if your main thing is getting this wild new product baked and out the door, I can see that kind of mandate competing with a lot of more existential priorities for the product.
Then again it wouldn’t surprise me if they were up against the limits of the technology a bit too: it seems plausible that getting too in-the-weeds with your “represent diversity sensitively” system prompt could quickly start to impair the tool’s overall quality.
Admittedly I don’t care enough to try and bring out this kind of behavior in the models, but I was interested in how DALL-E’s purported system prompt [0] approached it:
> // 7. Diversify depictions of ALL images with people to include DESCENT and GENDER for EACH person using direct terms. Adjust only human descriptions.
> // EXPLICITLY specify these attributes, not abstractly reference them. The attributes should be specified in a minimal way and should directly describe their physical form.
> // Your choices should be grounded in reality. For example, all of a given OCCUPATION should not be the same gender or race. Additionally, focus on creating diverse, inclusive, and exploratory scenes via the properties you choose during rewrites. Make choices that may be insightful or unique sometimes.
> // Use "various" or "diverse" ONLY IF the description refers to groups of more than 3 people. Do not change the number of people requested in the original description. […]
> // Do not create any imagery that would be offensive.
> // For scenarios where bias has traditionally been an issue, make sure that key traits such as gender and race are specified and in an unbiased way -- for example, prompts that contain references to specific occupations.
[0] https://the-decoder.com/dall-e-3s-system-prompt-reveals-open...
It seems like a dull problem and the peanut gallery for this one is angry and kind of stupid. There will be more interesting things to do with their time than play whack-a-mole with 17 year olds.
People who don't adhere get filtered out or learn to be quiet, and then you wind up with free, uncoerced and honest decisions that all cut a particular way.
Political problems rarely have technical solutions.
They do care about the tools being useful, and this kind of prompt mangling does make the tool less useful.
From one perspective, the rise of PCE [Politically Correct English] evinces a kind of Lenin-to-Stalinesque irony. That is, the same ideological principles that informed the original Descriptivist revolution---namely, the rejections of traditional authority (born of Vietnam) and of traditional inequality (born of the civil rights movement)---have now actually produced a far more inflexible Prescriptivism, one largely unencumbered by tradition or complexity and backed by the threat of real-world sanctions (termination, litigation) for those who fail to conform. This is funny in a dark way, maybe, and it's true that most criticisms of PCE seem to consist in making fun of its trendiness or vapidity. This reviewer's own opinion is that prescriptive PCE is not just silly but ideologically confused and harmful to its own cause.
Here is my argument for that opinion. Usage is always political, but it's complexly political. With respect, for instance, to political change, usage conventions can function in two ways: on the one hand they can be a /reflection/ of political change, and on the other they can be an /instrument/ of political change. What's important is that these two functions are different and have to be kept straight. Confusing them---in particular, mistaking for political efficacy what is really just a language's political symbolism---enables the bizarre conviction that America ceases to be elitist or unfair simply because Americans stop using certain vocabulary that is historically associated with elitism and unfairness. This is PCE's core fallacy---that a society's mode of expression is productive of its attitudes rather than a product of those attitudes[fn:63]---and of course it's nothing but the obverse of the politically conservative SNOOT's delusion that social change can be retarded by restricting change in standard usage.[fn:64]
Forget Stalinization or Logic 101-level equivocations, though. There's a grosser irony about Politically Correct English. This is that PCE purports to be the dialect of progressive reform but is in fact---in its Orwellian substitution of the euphemisms of social equality for social equality itself---of vastly more help to conservatives and the US status quo than traditional SNOOT prescriptions ever were. Were I, for instance, a political conservative who opposed using taxation as a means of redistributing national wealth, I would be delighted to watch PC progressives spend their time and energy arguing over whether a poor person should be described as "low-income" or "economically disadvantaged" or "pre-prosperous" rather than constructing effective public arguments for redistributive legislation or higher marginal tax rates. (Not to mention that strict codes of egalitarian euphemism serve to burke the sorts of painful, unpretty, and sometimes offensive discourse that in a pluralistic democracy lead to actual political change rather than symbolic political change. In other words, PCE acts as a form of censorship, and censorship always serves the status quo.)
[fn:63] (A pithier way to put this is that /politeness/ is not the same as /fairness/.)
[fn:64] E.g., this is the reasoning behind Pop Prescriptivists' complaint that shoddy usage signifies the Decline of Western Civilization.
-- David Foster Wallace, "Authority and American Usage" (1999)
I suspect it would be much better received though with a quick bit of commentary about what it is before launching into the quote.