It's deeply shameful that billions of dollars and the hard work of incredibly smart people is mangled for a 'feature' that most end users don't even want and can't turn off.
This is not a one off, it keeps happening with generative AI all the time. Silent prompt injections are visible for now with jailbreaks but who knows what level of stupidity goes on during training?
Look at this example from the Würstchen paper (which stable cascade is based on):
>This work uses the LAION 5-B dataset...
>As an additional precaution, we aggressively filter the dataset to 1.76% of its original size, to reduce the risk of harmful content being accidentally present (see Appendix G).
That’s the crux of what’s so off-putting about this whole thing. If Google or OpenAI told you your query was to be prepended with XYZ instructions, you could calibrate your expectations correctly. But they don’t want you to know they’re doing that.
Billions of dollars worth of data and manhours could only be justified for something that could turn a profit, and the obvious way an advertising company like Google could make money off a prompt handler like this would be "sponsored" prompts. (i.e. if I ask for images of Ben Franklin and Coke was bidding, then here's Ben Franklin drinking a refreshing diet coke)
AI as learning tool here feels misplaced to me.
The point is that those modifications should be reliable, so if you want a viking man/woman or an asian/african/greek viking then adding those modifiers should all just work.
Any other specific things we should not expect from AI or shouldn't ask AI to do?
This seems completely reasonable to me. I still don't trust computers.
However on the other hand that is a misuse of AI, since we already know that hallucinations exist, are common, and that AI output must be verified by a human.
So as a counterpoint, there are sound reasons for using AI to generate images based on history. The same reasons are why we use illustrations to demonstrate ideas where there is no photographic record.
A straightforward example is visualising the lifetime/lifestyle of long past historical figures.
I blame that decade of near zero interest rates. Companies could post record profits without working for them. I think in the coming years we will discover that that event functionally broke many companies.
But we are trying to create a tool where we can ask it questions and it gives us answers. It would be nice if it tried to make the answers accurate.
By lowering standards for black doctors do you think anyone in their right mind would pick black doctors? No I want the fat old jew. I know no one put him in the hospital to fill out a quota.
That’s what a movie going to be in the future. People are going to prompt characters that AI will animate.
Insane amounts of research go into creating historical movies, games etc that are serious about getting it right. But to try and please everyone, they take lots of liberties, because they're creating a product for the masses. For that very same reason, we get tons of historical depictions of New York and London, but none of the medium sized city where I live.
The effort/cost that goes into historical accuracy is not reasonable without catering to the mass market, so it seems like a conundrum only lots of free time for a lot of people or automation could possibly break.
Not holding my breath that it's ever going to be technically possible, but boy do I see the appeal!