I mean that the likelihood (or weigh) of an opinion being expressed by an AI should be roughly proportional to the number of people who currently hold that opinion, assuming the AI is simply generating responses based on its training (which is what should actually be as unbiased as possible).
As an example, let's suppose that 55% of people believe that it's not OK to make jokes about women, but it's OK to make jokes about men, and that roughly 40% believe it's OK to make jokes about both (I'm not saying this is the case, it's just an example).
So perhaps, in this case, by default the AI wouldn't make a joke about women.
But if you would slightly nudge it or insist a bit more, perhaps the AI wouldn't refuse to make a joke about women anymore, because there is still a large proportion of the population who do believe that's perfectly OK (of course, then we might get into the territory about overtly sexist jokes, which obviously the AI would have to refuse a lot more than making a more innocent joke about women).
Now let's say we start asking the AI to make Nazi comments. Obviously, the segment of the population who agrees with Nazi sentiment is a lot smaller, and the anti-Nazi sentiment is a lot stronger, so the AI should have to object to such a request quite more strongly.
This type of refusal or likelihood of the AI saying something should presumably be roughly proportional to the opinions and sentiment of the general population (or at the very least, the target market for the AI), not just the OpenAI employees who performed the RLHF to train the AI in terms of acceptable responses and who are much more likely to be biased.
I'm not saying that this is necessarily easy to accomplish, there are certainly difficulties here. As an example, some widely-held opinions, even about objective things, may not necessarily be rooted in facts, so some kind of balancing might be necessary (a general kind of balancing, not a "let's dissect and nudge the AI responses on an opinion-by-opinion basis"). And yes, I understand that this can be quite difficult, because any given source of truth can be perceived to be biased by some segment of the population.
What I am saying, however, is that AI creators such as OpenAI should be making more efforts in this direction.
To start with, perhaps the RLHF training should be done with AI trainers selected from a more representative sample of the population.
And yes, we may never be able to accomplish 0% bias, but we should at least make some effort to reduce it.
It's also interesting to me that at some point, the AI may start to express opinions that are not a strict "linear" function of the data it was trained on, and yes, this might piss off a significant amount of people. In my opinion, this would be quite interesting and should be OK, as long as we made reasonable efforts to remove sources of bias from its training process.
Although I can also see an important target market (perhaps even larger) for an AI that is more biased to generate responses according to the beliefs of the general population, rather than what it "perceives" to be more true.