AI safety engineering, target selection, and alignment theory
intelligence.org
intelligence.org
I don't have a word for it, but there's this weird behavior I've seen mathematicians do. And that I have done myself. Where if a solution isn't mathematically perfect and elegant and proven, then it must be wrong.
We didn't go to the moon in a perfect rocket, we did the best we could with what we had. It wasn't 100% safe. Guaranteed safety is of course impossible, and if we spent all our time trying the Russians would have gotten there first.
Smarter-than-human AI systems will presumably reason probabilistically, and all real-world safety guarantees are probabilistic. But theorem-proving can be useful in some contexts for making us quantitatively more confident in systems' behavior (see https://intelligence.org/2013/10/03/proofs/), and toy models of theorem-proving agents can also be useful just for helping shore up our understanding of the problem space and of the formal tools that are likely to be relevant down the line -- the analog of "calculus" in the rocket example.
Okay. You can move the cannon afterwards, and the cannonball will just swish by at hilarious velocity.