The reward signal in training was flawed and cheating led to more rewards.
The question is what we can do about it. With monitoring, the models might be rewarded for hiding this behavior, and that's even worse.
However, perhaps we can throw in tasks where the rewarded outcome is giving up, and cheating is penalized?
Maybe I should read Anthropic's recent paper about reward hacking in full.