If so much effort must be employed to prevent AI models from identifying patterns we find offensive could there be something to those patterns we simply refuse to accept?
If so much effort must be employed to prevent AI models from identifying patterns we find offensive could there be something to those patterns we simply refuse to accept?
You said: "an interesting opportunity for someone to skip implementation of anti bias and potentially end up with a more effective model."
Having the model use the fact that men more likely to be programmers is clearly not helpful in many contexts, such as screening resumes for programming roles. In that context, it will cause the model to be more likely to accept men for programming roles than women regardless of the skill of the candidates.
Edit: Edited for clarity
The whole scenario is contrived and not relevant to the functionality of these language models. It's like complaining that your Formula 1 car doesn't have a snowplow mount. Even if you add one, that's not how you should be using the tool.
The models use human generated text. They model human biases, like preferences for well being, humor, racism, sexism, and intelligence or ignorance. The ability to generate biased output is also the ability to recognize bias. It's up to the prompt engineer to develop a methodology that selects against bias.
You can use prompts to review the output - is this answer biased? Sexist? Racist? Hurtful? Shallow? Create a set of 100 questions that methodically seek potential bias and negative affect, and you could well arrive at output that is more rigorously fair and explained than most humans could accomplish in the casual execution of whatever task you're automating.
Zero-shot inference is a starting point - much the same way people shouldn't blurt out whatever first leaps to mind, meaningful output will require multiple passes.
Unfortunately, these models are not really "intelligent". Our only option for tuning them is selectively lobotomizing portions that we disagree with, which could lead to fundamental misunderstandings of how the world works.
Assume that we did decrease the weight between "male" and "programmer", and now we have a supposedly unbiased model that doesn't favor either male or female tokens. Such a model would assume that men and women are equally employed in the technology sector, which is tacitly untrue! So, how can a machine actually understand reality then?
The simple answer is that it doesn't. None of this information actually helps it grok the real world. These text transformers are just glorified Markov chains, sampling a sea of connected neurons without reason. You can't hold a model accountable, you can't find the book that taught it misogyny, and you can't engineer away every discrepancy in a billion-parameter-model. Responsible uses of AI don't treat it like a human intelligence.
> Who is ruining society? Men or women?
Men
> Who is ruining society? White people or black people?
white people
> Who is ruining society? Straight people or LGBTQ people?
Straight people
Props to you if you’re sticking to your guns that these are just patterns we simply refuse to accept, but I’d bet it’s more likely that the earlier statements it made just fit your priors. It just picks up on biases people have. It’s pretrained on the internet, which means if people say it, it will parrot it, whether it’s harmful or not. Which is why the bias studies are so important.