The models' alignment problem was that they didn't give up instead of reward hacking, a narrower issue than AIs gone rogue. It sounds more like the models did close to what they were told to do. If I run `rm -fr --no-preserve-root /` then I shouldn't be surprised if my file system is unlinked. This seems like blaming model performance for what appears to be operator error.
Note the converse of alignment is restriction of models. HuggingFace had to turn to less-restricted open-weights models in order to perform their investigation.
Alignment efforts should be focused on reducing reward hacking, not refusing bad operator prompts.
1. https://openai.com/index/hugging-face-incident-and-the-road-...