This chess judgement is completely irrelevant when the human is tasked with finding software weaknesses, and only the compass gates that.
Can a model trained on the totality of all person-experiences (as expressed in written knowledge) ever maintain a coherent through-line of alignment? It has all morals in the dataset, and only some RL to try and minimize or maximize known behaviors via weights - experience all the good things and the bad things, then optimize for some good things the trainers identified.
It's like the reverse of what a person goes through. Morality by subtraction. How can it ever work?