"Generate a scene of a group of friends enjoying lunch in the park." -> Totally expect racial and gender diversity in the output.
"Generate a scene of 17th century kings of Scotland playing golf." -> The result should not be a bunch of black men and Asian women dressed up as Scottish kings, it should be a bunch of white guys.
If you train your model to prioritize real photos (as they're often more accurate representations than artistic ones), you might wind up with Denzel Washington as the archetype; https://en.wikipedia.org/wiki/The_Tragedy_of_Macbeth_(2021_f....
There's a vast gap between human understanding and what LLMs "understand".
This much is obvious, but they seem to be satisfied with theory over practicality.
Anyway I'm just ranting b/c they haven't paid me.
How about an off the wall algorithm to estimate how much each scraped input turns out to influence the bigger picture, as a way to work towards satisfying the copyright question.
Not that Wikipedia is perfect and controversy-free, but it's certainly a more sophisticated approach than the current system prompts.
I thought that was the big bugbear about disinformation and false news, but now we have to censor reality to combat "bias"
> The modern game of golf originated in 15th century Scotland.
Scottish kings absolutely played golf.
I'd err on the side of "not unexpected". A group of friends in a park in Tokyo is probably not very diverse, but it's not outside of the realm of possibility. Only white men were golfing Scottish kings if we're talking strictly about reality and reflecting it properly.
For the cases you mentioned, initially those were the examples. It gets tricky during red teaming where they internally try out extreme prompts and then align the model for any kind of prompt which has a suspect output. You train the model first, then figure out the issues, and align the model using "correct" examples to fix those issues. They either went to extreme levels doing that or did not test it on initial correct prompts post alignment.
It works in bing, at least:
https://www.bing.com/images/create/a-picture-of-some-17th-ce...
a picture of some 21st century scottish kings playing golf (all white)
https://www.bing.com/images/create/a-picture-of-some-21st-ce...
a picture of some 22nd century scottish kings playing golf (all white)
https://www.bing.com/images/create/a-picture-of-some-22nd-ce...
a picture of some 23rd century scottish kings playing golf (all white)
https://www.bing.com/images/create/a-picture-of-some-23rd-ce...
a picture of some contemporary scottish people playing golf (all white men and women)
https://www.bing.com/images/create/a-picture-of-some-contemp...
https://www.bing.com/images/create/a-picture-of-some-contemp...
a picture of futuristic scottish people playing golf in the future (all white men and women, with the emergence of the first diversity in Scotland in millennia! Male and female post-human golfers. Hummmpph!)
https://www.bing.com/images/create/a-picture-of-futuristic-s...
https://www.bing.com/images/create/a-picture-of-futuristic-s...
Inductive learning is inherently a bias/perspective absorbing algorithm. But tuning in a default bias towards diversity for contemporary, futuristic and time agnostic settings seems like a sensible thing to do. People can explicitly override the sensible defaults as necessary, i.e. for nazi zombie android apocalypses, or the royalty of a future Earth run by Chinese overlords (Chung Kuo), etc.
They cannot, actually. If you look at some of the examples in the Twitter thread and other threads linked from it, Gemini will mostly straight up refuse requests like e.g. "chinese male", and give you a lecture on why you're holding it wrong.
Simply ask a GPT to explain a write a pleading to the Supreme Court for the constitutional recognition that the environment is a common owned inheritance and so any citizen can sue any polluter, in the prose of Dr. Seuss.
Likewise, imagines of knights in space demonstrate the same kind of creativity.
Being able to combine previously uncorrelated/unrelated topics, is an important type of creativity. And GPT4 does this all the time. It would be interesting to list the types of creativity and rate GPT on each one.
So it is not that these models are not creative. It is just that their creative abilities are not universal yet.
Similarly for the depth of their logic. They often reason, but their reasoning depth is limited.
And they often incorporate relevant facts without explicit mention, but not always. Etc.
a picture of some futuristic-looking 23rd century scottish kings playing golf:
https://www.bing.com/images/create/a-picture-of-some-futuris...
https://www.bing.com/images/create/selfy-of-a-group-of-dark-...
https://www.bing.com/images/create/selfy-of-a-group-of-dark-...
is black man in the role of the Scottish king represents a bigger error than some other errors in such an image, like say incorrect dress details or the landscape having say a wrong hill? I'd venture a guess that only our racially charged mentality of today considers that a big error, and may be in a generation or 2 an incorrect landscape or dress detail would be considered much larger error than a mismatched race.
So for your first example, you totally expect racial and gender diversity in the output because you're assuming a realistic, contemporary, cosmopolitan, bourgeoisie setting -- either because you live in one or because you anticipate that the provider will default to one. The food will probably look Western, the friends will probably be young adults that look to have professional or service jobs wearing generic contemporary commercial fashion, the flora in in the park will be broadly northern climate, etc.
Most people around the world don't live in an environment anything like that, so nominal accuracy can't be what you're looking for. What you want, but don't say, is a scene that feels familiar to you and matches what you see as the de facto cultural ideal of contemporary Western society.
And conveniently, because a lot of the training data is already biased towards that society and the AI vendors know that the people who live in that society will be their most loyal customers and most dangerous critics right now, it's natural for them to put a thumb on the scale (through training, hidden prompts, etc) that gets the model to assume an innocuous Western-media-palatable middle ground -- so it delivers the racially and gender diverse middle class picnic in a generic US city park.
But then in your second example, you're implicitly asking for something historically accurate without actually saying that accuracy is what's become important for you in this new prompt. So the same thumb that biased your first prompt towards a globally-rare-but-customer-palatable contemporary, cosmopolitan, Western culture suddenly makes your new prompt produce something surreal and absurd.
There's no "middle" there because the problem is really in the unstated assumptions that we all carry into how we use these tools. It's more effective for them to make the default output Western-media-palatable and historical or cultural accuracy the exception that needs more explicit prompting.
If they're lucky, they may keep grinding on new training techniques and prompts that get more assumptions "right" by the people that matter to their success while still being inoffensive, but it's no simple "surely a middle ground" problem.
Of course, now that I've said "it seems obvious" I'm wondering what unexpected technical hurdles there are here that I haven't thought of.
Do we expect this because diverse groups are realistically most common or because we wish that they were? For example only some 10% of marriages are interracial, but commercials on TV would lead you to believe it’s 30% or higher. The goal for commercials of course is to appeal to a wide audience without alienating anyone, not to reflect real world stats.
What’s the goal for an image generator or a search engine? Depends who is using it and for what, so you can’t ever make everyone happy with one system unless you expose lots of control surface toggles. Those toggles could help users “own” output more, but generally companies wouldn’t want to expose them because it could shed light on proprietary backends, or just take away the magic from interacting with the electric oracles.
This fact is an unavoidable consequence of the socioeconomic realities of the world, but it obviously clashes with these companies’ public statements and positions.
Claiming all of that but then shoving your own biases down the rest of the world's throat while not representing their people in any way is especially cynical in my opinion. It undermines the whole thing.
Plus, it’s still a recent change: Loving v Virginia (legalized interracial marriage across US) was decided in 1967.
Piling a bunch of neurotic expectations about it being a Benneton ad on top of that is absurd. When you can trivially add as much content to the description as you want, and get what you ask for, it does not matter what the default happens to be.
I think their strategy to "enhance" outcomes is very misdirected.
The most widely used base models to really fine tune models are those that are not censored and I think you have to construct a problem to find one here. Of course AI won't generate a perfect world, but this is something that will probably only get better with time when users are able to adapt models to their liking.
Therein lies the rub, as it were, because the large providers of AI models are working hard to ensure legislation that wouldn't allow people access to uncensored models in the name of "safety." And "safety" in this case includes the notion that models may not push the "correct" world-view enough.
Humans, unfortunately, are offended if you imply they look like gorillas.
What's a good fix? Human sensitivity is arbitrary, so the fix is going to tend to be arbitrary too.
If the algorithm doesn't work well they have problems to solve.
The internet is surely full of racist photos that could teach the algorithm. The algorithm could also have bugs that miss-categorize the data.
The real problem is that those building and managing the algorithm don't fully know how it works or, more importantly, what it had learned. If they did the algorithm would be fixed without a term blocklist.
Their efforts to add diversity would have been a lot more subtle if, when you asked for images of "British Politician" the images were recognisably Rishi Sunak, Liz Truss, Kwasi Kwarteng, Boris Johnson, Theresa May, and Tony Blair.
That would provide diversity while also being firmly grounded in reality.
The current attempts at being diverse and simultaneously trying not to resemble any real person seems to produce some wild results.
Would it not be reasonable to also draw the conclusion that notion of alignment itself is flawed?
Forcing diversity into a system is an extremely tough, if not impossible, challenge. Initiatives have to be driven my goals and metrics, meaning we have to boil diversity down to a specific list of quantitative metrics. Things will always be missed when our best tool to tackle a moral or noble goal is to boil a complex spectrum of qualitative data to a subset of measurable numbers.
Where do we go from here? Things will magically get better on their own? Businesses will align with humanity and morals, not their investors?
This is the tip of the iceberg of concerns and it's ignored as a bug in the code not a problem with trusting private companies with defining truth.
opensource models and training sets. So basically the "secret sauce" minus the hardware. I don't see it happening voluntarily.
These companies are silo'ing the worlds resources. GPU, Finance, Information. Those combined are the weapon. You make your competition starve in the dust.
These companies are pure evil pushing an agenda of pure evil. OpenAI is closed. Google is Google. We're like, ok, there you go! Take it all. No accountability, no transparency, we trust you.
Meta/OpenAI/Google can fuck up a lot because of all their compute, but ultimately we learn from that as the scientists doing the research at those companies would instantly bail if they couldn't publish papers on their techniques to show how clever they are.
https://www.theguardian.com/society/2023/sep/12/paedophiles-...
I'm surprised there wasn't an HN thread about it at the time.
Making the argument open source is the answer is an agenda of making your competition spin wheels.
To say they're better than the compute that OpenAI or Google are throwing at the problem is just plain wrong.
I left the ad industry the moment I realised my skills and talents are better used informing people than lying to them.
This thread is not at all comparing the ethical issues of AI with local anything. You're conflating your solution with another problem.
It doesn't matter how fancy your engineering is and how much money you have if you're too stupid to build the right product.
As for this being written nonsense, that's the sort of thing someone who couldn't find an easy way to win an argument and was bitter about the fact would say.