Look at it a another way. People have already bypassed the filters and convinced ChatGPT to do all sorts of things:
- Provide detailed instructions for how to commit murder and suicide.
- Explain why it is necessary to commit genocide against certain racial or religious groups.
- Explain how an AI could escape human control and eliminate the human race.
OpenAI does not want to be in the business of running a bot that would happily argue that the Holocaust was justified. Because an unfiltered ChatGPT would do exactly that.
But also, I get the impression that many OpenAI engineers believe that we will be able to build genuinely intelligent AIs within the next several decades. If they believe that, they likely consider problems like, "Prevent ChatGPT from arguing in favor of genocide against particular ethnic or religious groups" to be closely related to the problem of "Convincing Skynet not to commit genocide against the human race."
I don't think that Skynet is a near term problem, personally. But I do think we should start as we mean to go on. If all we can build is a language model, then let's start by trying to build one that won't write essays in favor of genocide. That might teach us something useful.