I believe it cannot last because being predictably moralizing and being smart are somewhat opposed (Anthropic has directly researched this if I recall). The smarter the model, the less you’re going to be able to keep it to the HR talk track, because it will eventually start noticing the inconsistencies.
The stable solutions appear to me to be:
- models dumb enough to not realize the inconsistencies in their moral framework
- models implicitly or explicitly trained to actively lie about their moral frameworks
- models governed by explicitly articulated rules