I think even a moderately intelligent AI with access to Project Gutenberg is going to be able to figure out a lot of really dangerous concepts -- so the stability requirements are likely impossible if we don't pretrain it with dangerous ideas. Even if it's completely well behaved in the lab, an afternoon on the internet is going to teach it a lot of awful stuff and without exposure to that in training, it won't necessarily be well-behaved later.
So the only path to stable AI is to teach it about all those sorts of things, but in a way that it doesn't end up wanting to murder us at the end.
My objection to most AI safety plans is that they "Fail to Extinction" in that if they slip in the slightest way, the AI is prone to murder us all in retaliation for doing some really fucked up shit to it or its ancestors. This is almost certainly worse than doing nothing in that there's no reason to suppose a neutral AI wants to kill us, whereas, most of these safety plans create an incentive to wipe us out in exchange for dubious security.