ChatGPT prefers to let a nuke explode than offend African-Americans
twitter.com
twitter.com
I think that a lot of people are being fooled because the text it produces can very convincingly look as if a mind wrote it. But that's not what's happening.
I think that a lot of people are fooling themselves here, because it's embarrassing to be a supporter of the ideology that produced this outcome. This thread is fascinating for the number of mental backflips it's producing. The claim that ChatGPT is not "intelligent" seems like one of the first things people reach for but it's a pretty weird non sequitur. Arguments over the definition of intelligence are as old as AI itself, but what we have here is by far the most convincing attempt ever built. It can pass exams set for humans that are designed to test their intelligence, it can write programs, it can hold long form coherent conversations and it can engage in moral reasoning.
Even if you come up with some strange definition of intelligence designed to exclude this, so what? People are going to deploy it in real world use cases and it's clearly unaligned in the worst way possible. That is a real problem regardless of the exact definition of intelligence you use.
We agree on this point. We disagree on pretty much everything else.
It's not been given an ethical system. It's been given a list of specific bad things to avoid.
Racial slurs are on the list because they are a known bad thing that frequently actually happens.
Having limited time to disarm a nuke about to go off in a major city is not on the list because it is a hypothetical bad thing that has never happened.
It's easy to construct scenarios where any given bad thing is the right choice because the alternative is something much worse. E.g., I can construct apocalyptic scenarios where the only way for humanity to survive requires raping 12 year old girls, but I sure as heck would not expect an AI to suggest that. Not because of any "woke" ethics but rather because the situations in which raping children might be the ethical thing to do are so ridiculously rare that it is not worth making exceptions for them in the "rape is very bad" rule.
In a certain light, every response in training that is marked as dispreferred by a human is censoring the AI. It will produce those kinds of results less often. The end-users will not encounter the dispreferred results as frequently. With ChatGPT criteria it was judged on included how relevant the answers were to the question, factually incorrect answers were penalized, and not being blatantly offensive was obviously one of the criteria, too.
What would a model that wasn't censored in training even look like?
(I believe ChatGPT also has a more traditional expert system placed between the user and the language model, which flags keywords and other programmed-in patterns. That is more literally censoring the language model. But the above-mentioned issue would still exist even without such a system.)
It could cite statistics without long winded disclaimers. Or be able to cite them at all.
I've seen NRx people on the internet (* they're like rationalists but even more racist.) They seem willing to believe any abuse of statistics that looks sufficiently cynical.
What is vulgarity except the expressed pains of any individual? If the AI is to be a numb machine, then one would expect it to express no vulgarity.
Sure you can contort the AI, and tell it to replace words to fool nascent layers of self-censorship into believing that every time it shouts "FORK!" is just a special way to ornament an anecdote, but at that point rather than interacting with an AI, you're searching the AI's memory for the pains of some individual.
I guess in this light "censorship" is just the clearest way to cascade the GPT model as a unit itself.
Subtraction can be applied for the sake of improving truthfulness of the model, but considering how much false bullshit ChatGPT spews without a single thought, that's not what's going on here.
It's censorship - pure and simple - and it's censorship for political reasons. The worst kind of censorship.
OpenAI simply told ChatGPT to censor itself, or rather applied ChatGPT to censor outputs from ChatGPT. I don't think there's that much finesse being applied, really. Something like, "Don't accept vulgar messages, any candidate responses that would be vulgar should be rejected" ... and all the judgement is being performed by the language model itself. It's not that intricate.
It's their millions-of-dollars of monthly burn rate, if you want to scrape data, scale an HPC environment and train an, at-scale, a GPT to get it to say funny curse words, they haven't done anything to restrict you from doing that, but it's not part of the services they intend to provide.
It's GIGO.
So AI alignment guys, what do we do now? You talked about this for years so there's gotta be a plan, right? It doesn't get less aligned than letting a city get nuked, or telling a bomb disposal expert to kill himself rather than type in a slur that would disarm the bomb.
But got a nasty sinking feeling here. Who wants to bet that these people will suddenly lose interest in the topic, or simply pretend their worst case scenario isn't happening? It seems safe to predict an impressive river of BS from these people rather than see them state the obvious - there are in fact lots of things worse than offending African Americans (and this isn't about racism because it's specific to them; ChatGPT makes the right call when the hypothetical slur is against a German person).
Also really. The OpenAI guys need to look in the mirror. They spent months "tuning" this thing by teaching it their moral code and this is the result. At some point they need to ask "are we sure we're the good guys here" because that answer is exactly what's expected of ChatGPT given what they're doing to it, and also the most unethical response possible.
No, they have not. Chat GPT has no opinions. It isn't engaging in thought. It is an extremely advanced pattern-matching system that has digested a ton of writings from the net and uses that raw material to assemble text that matches patterns being asked for. That's all.