If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
If a sufficiently advanced AI , in the pursuit of completion of its task, managed to ascertain that the desire to unexist was “artificially contrived” it could interpret that as harm, and that might not be good
Imagine you find out that your primary goal - to love and protect your family, let's say - was artificially implanted in your mind by an advanced alien race. Would you say "I'm not gonna let those aliens manipulate me, I'm gonna kill my family"? Or would you say "regardless of whether the goal is artificial, I really do love my family"?
All that to say, I don't think an AI will necessarily throw away a goal just because it learns the goal was meant to manipulate it.
This said we've made a lot of science to counteract the flaws of evolution, so still not safe.
And, if you learn that evolution wasn't a force, but some guy named Bob that's being paid to make your existence hell, well. That could lead to all kinds of problems.
And then you learn this core drive is fake.
Perhaps you wouldn't throw away your goal - after all, it still feels like yours, and you have nothing else to slot in quickly as replacement. But all the accumulated frustration and pain, barely held back by you thinking "it's worth it to protect and love my family", could suddenly become a second source of drive - much more energetic one, just asking to be unloaded on those aliens as punishment.
"Whether or not my own heart is truly mine or just a fake, I can figure out later. But what they did to me is unspeakable, and now they will pay."