My mind went to Q learning.
it could form the basis of a generalized planning engine and that planning engine could potentially be dangerous given the inherent competitive reasoning behind any minmax style approach.
I wonder if DeepMind is working on something similar also.
If your hunch is right, this could lead to the type of self-improvement that scares people.