- We are social animals, and take for granted that, all else being equal, it's better to be good to other creatures than bad to them, and to be truthful rather than lie, and such. However, if you select values uniformly at random from value space, "being nice" and "being truthful" are oddly specific. There's nothing universally special about deeply valuing human lives any more so than say deeply valuing regular heptagons. Our social instincts are very ingrained, though, making us systematically underestimate just how little a smart AI is likely to care whatsoever about our existence, except as a potential obstacle to its goals.
- Inner alignment failure is a thing, and AFAIK we don't really have any way to deal with that. For those that don't know the phrase, here it is explained via a meme: https://astralcodexten.substack.com/p/deceptively-aligned-me...
So here's hoping you're right about (a). The harder AGI is, the longer we have to figure out AI alignment by trial and error, before we get something that's truly dangerous or that learns deception.