This sounds smart until you think about it for ten seconds.
If an ASI model is 100% aligned to user intent then you only need one person on earth to prompt "kill everyone" for an extinction-level invent.
There's no logical way around this. The model either has to ignore the person at the helm or we have to do a multi-national abort before RSI. Anthropic is trying #1.