- There is some behaviour that we want the model to show, and the inverse we do not want it to. - Both are learned in the massive training phase - OpenAI used RLHF to suppress undesired behaviour, but it was ineffective because we have orders of magnitude less RLHF data.
That would imply that RLHF would slightly suppress the 'bad' behaviour, but it still would be easy to output it.
This is disproved by what the post is trying to explain: We see _increased_ bad behaviour by using RLHF. The post agrees with the premise that both good (wanted) and bad (unwanted) behaviour is learned during training. But it's proposing the 'Waluigi effect' to explain why RLHF actually backfires.
Now, tbh it does rely on the assumption that we are actually seeing more undesired behaviour than before. If that was false then it would falsify the Waluigi hypothesis.