This is what the Waluigi effect is, since it isn't described at the top:
> The Waluigi Effect: After you train an LLM to satisfy a desirable property P, then it's easier to elicit the chatbot into satisfying the exact opposite of property P.
Basically, the chatbot will often do the opposite of what you say.