The other day the system prompts for Claude (chat not code) hit the frontpage and they included a bit that something along the lines of "Claude should avoid saying honestly, because Claude is always honest".
So how come that these Claudisms still are so frequent in LLM output?
Do the system prompts just not work?
Don't the postprocess the output to deal with such policy violations?