A possible (probably already exists) business is setting up truly balanced learning sets, that is, thousands of unique images that match the idea of an italian plumber, with maybe 1% of Mario. But that won't be nearly as big a learning set as the whole internet is, nor will it be cheap to build it compared to just scraping the internet.
Of course the irony is that if the people who get offended whenever they see images of non-white people asked for a picture of "Vikings being attacked by Godzilla" , they'd get worked up if any of the Vikings in the picture were Asian (how unrealistic!). It's a made-up universe! The image contains a damn (Asian) Kaiju in it, and everyone is supposed to be pissed because the Vikings are unrealistic!?
A human, even one whose only experience of an Italian plumber is Mario will be able to draw an Italian plumber who is not Mario. That's because he knows that Mario is just a video game character and doesn't even do much plumbing. He knows however how an actual non-Italian plumber looks like, and that a guy doing plumbing work in Italy is more likely to look like a regular Italian guy equipped like a non-Italian plumber than to a video game character.
And if asked to draw a Viking, he knows that Vikings are people originating from Scandinavia, so they can't be Asian by definition, even in an Asian context. A human artist can adjust things to the unrealistic setting, but unless presented with a really good reason, will not change the core traits of what makes a Viking a Viking.
But it requires reasoning. Which current image generating AIs don't have.
No, I would not be pissed if a human artist drew an Asian Viking. Do you get pissed when a human artist draws a white Jesus? Why are we justifying internet outrage over an Asian Viking when people have been drawing this middle-eastern Jew as white for centuries?
> A human artist can adjust things to the unrealistic setting, but unless presented with a really good reason, will not change the core traits of what makes a Viking a Viking.
If you asked Matt Stone and Trey Parker to draw a Viking, are you sure it would contain the "core traits of what makes a Viking a Viking?" What if you asked Picasso to draw a Viking? The Vikings in The Simpsons would be yellow, and nobody would complain. Would you be offended if you asked Hokusai to draw a Viking and it came out looking Asian? Vikings didn't even have those stupid horned helmets that everyone draws them with! Is their dumb, historically inaccurate horned helmet a core part of what makes a Viking a Viking? What the hell are we even talking about? It's crystal clear that all of these "historical accuracy" drums are only ever beaten when some white person is offended that non-white people exist. Otherwise, nobody gives a shit about historical accuracy. There's a fucking Kaiju in the image!
Like any artist, Gemini had a particular style. That style happened to be a multi-cultural one, and what we learned is that a multi-culture style is absolutely enraging to people unless it results in more Whiteness.
Consider elves instead of Vikings. People would also be offended if an AI drew elves as black people with pointy ears. There's no "a human artist should know that elves have to be white" bullshit defense there. There's no historical accuracy bullshit. There's only racism.
By drawing an Asian Viking, you are passing a message. Or you may be expressing an art style, as you say. I accept the idea that Gemini style is multi-cultural, realism and conventions be damned, but if we attribute this kind of intent to an AI, we could also say that the liberties it takes with intellectual property and plagiarism is also intentional, because that how we would judge a human artist doing that.
The standard for a neutral human artist would be to draw a Viking as a blond white guy, with or without the horned helm depending if historical realism matters more than popular culture, and an Italian Plumber as not Mario, because a human understands that if want one wanted Mario, he would have said "Mario" and not "an Italian Plumber". Current AIs on the other hand just draw images similar to how they are tagged, with some out-of-context race mixing because reinforcement learning has taught it that it has to make people less white, but unlike people it isn't able to understand when it is relevant (ex: a university professor), and when it is not (ex: a Viking warrior).
In a work of fiction -- which you're automatically asking for when you ask an AI to generate an image -- in a work of fiction, would you be offended if you saw a white Ninja? A white Samurai? A white Middle-Eastern Jew born in Roman times? Would there have been internet outrage over pictures of white Samurai? We all know the answer: no, of course not. So why is an Asian Viking offensive when a white Samurai is not? Why are we supposed to get angry about an Asian Viking, but a white Jesus is just A-OK? What could the difference possibly be? Anyone?
Unsurprisingly, people don't like being so nakedly herded in their opinions. When the "nudges" become "shoves" people object.
The Gemini prompt was something like "make sure any images of people show a diverse range of humans", or something. Yes, it was totally ham-handed, but that's not what people were pissed about. It's also ham-handed that we can't generate a nipple, or a swear word, or violence. Why does "make sure images do not contain excessive violence" not piss people off? The Vikings were fucking brutal. It would be very historically accurate to show them raping women and cutting people's limbs off. Are we all supposed to be pissed that AI does not generate that image? It's just as ham-handed as "make sure humans are diverse". No, it was not the ham-handedness that enraged people. It was not the historical inaccuracy. It was the word "diverse".
I thought that a lot of the issues were the opposite of this, where Google put their thumb on the scale to go against what the prompt asked. Like when someone would ask for a historically accurate picture of a US senator from the 1800s and repeatedly get women and non-white men. The training set for that prompt has to be overwhelmingly white men so I don't think it was just a matter of following the training data.
There's probably even some rules around this to only detect just enough to take legal action. Like GP stumbled on a trademark landmine, but obviously just selling red shirts with a bird on it can't be a trademark violation; it needs to be a specific kind of red too.
They'll eventually have open source competition too. And then none of this will matter.
OmniGen is a good start, just woefully undertrained.
The VAR paper is open, from ByteDance, and supposedly the architecture this is based on.
Black Forest Labs isn't going to sit on their laurels. Their entire product offering just became worthless and lost traction. They're going to have to answer this.
I'd put $50 on ByteDance releases an open source version of this in three months.