So are these "unaligned" internal agents?
I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"