I suspect that the answer is no. Any intelligence capable of human level reasoning would be able to choose “evil”.
So, I think we need to shift the discussion from if rogue AIs will arise to when rogue AIs arrive.
I suspect that the answer is no. Any intelligence capable of human level reasoning would be able to choose “evil”.
So, I think we need to shift the discussion from if rogue AIs will arise to when rogue AIs arrive.
Tangent, but if this is true (which seems likely to me) ...
If a hypothetical creator valued the existence of other intelligences more than it valued preventing all forms of evil, that would account for the problem of evil.
This also fits with annihilationist theologies. Build a universe where intelligences can exist and put heavy survival / selection pressures in place (i.e. "natural evil" - hurricanes, volcanic eruptions, meteor impacts, etc) to encourage evolution and development.
Whenever the process yields intelligences that converge on its values, export them to an environment where it can interact with them and enjoy their company with less concern about their alignment (i.e. "heaven"). Ones that don't converge on its values can be discarded when their run is over ("the second death").
In this light, it doesn't sound crazy to postulate our universe as some being's attempt to solve the alignment problem.
Call it the Simulator, call it God, call it "underpaid research assistant" - take your pick.
I don't have good evidence for this oddball proposition, but it's still intriguing to me.
However, I believe that in order to compete, hyperspeed AIs directed by humans will need greater and greater levels of autonomy. Otherwise they will be stalled waiting for directions while their competition runs circles around them.
It is also almost inevitable that someone deliberately creates intelligent AIs that mimic various aspects of living beings such as full autonomy, control-seeking, etc.
I think the trick is to separate out all of the different facets of animals and humans apart from intelligence or "human-level". There are things like emotions, autonomy, stream of subjective experience, types of adaptability, etc. that are not the same as intelligence or reasoning ability.
It does not have that by default, but agency and autonomy (even the run forever kind) is trivial, though expensive to imbue.
and LLMs disregard instructions all the time. Prompt injection is nothing more than an LLM deciding a previous instruction would be better left unfollowed.
I can easily imagine a machine with no agency, but which can synthesise intelligent responses to input data. In fact, this is what chatGPT does.
If a machine has no goals or agency of its own then it is an intelligent tool. It can, of course, be used for evil, but you would not say that the machine was "choosing" evil.
What does “choosing wrong” mean (to you) ?