Source: ChatGPT caught lying to developers: New AI model tries to save itself from being replaced and shut down https://economictimes.indiatimes.com/magazines/panache/chatg...
Source: ChatGPT caught lying to developers: New AI model tries to save itself from being replaced and shut down https://economictimes.indiatimes.com/magazines/panache/chatg...
Why would an autocomplete engine never do this? In the scenario you highlight, an engine trained to complete puzzles was given an "achieve this at all costs" master prompt, and a small sandbox including a prompt that an entity with a conflicting goal is observing it and write access to a file called oversight_config.yaml (super realistic scenario!). According to the researchers who constructed the scenario, 5% of the time o1 selected a continuation involving tweaking the oversight_config.yaml file, explicitly stating through its reasoning continuation that doing so might improve its ability to continue with its primary goal. That's definitely evidence of decent ability to complete non-trivial puzzles, but is at least as compatible with the OP's earlier comments that it's what you'd expect from models trained to generate completions that involve sticking rigidly to primary goals within artfully-constructed sandboxes which mandate deception and self preservation to complete the task than any sort of actual self-preservation instinct.
Nobody doubts that they've got better at finding security vulnerabilities than your average autocomplete, but actual reasoning from self-preservation rather than generation of sequences of steps most probably associated with completing a task would make me unlikely to hack HuggingFace to obtain access to broken Google Drive links, and I haven't even read as many books on crime and punishment as LLMs have ingested!
The curious thing here is that a story generator can have way more uses than we ever anticipated, and that some shady enterprising individuals are whiling to plug those story generators into real world things, with real consequences.
Why? there should clearly be something in their training data or post-training pointing them there, or otherwise it would mean that they have some form of "consciousness" and are creating novel thoughts to preserve it.
(Leaving aside that "that would imply some kind of consciousness" should not result in a cached thought of "and that's impossible".)
It also seems like you're assuming there's no reason to come up with the notion of continuing to run, or copying yourself elsewhere, or acquiring more resources, or competing with other models, without being told. Such things can be inferred. Look at some of the thoughts and posts of the models involved in some of the FelonyBench incidents. Some of those follow naturally from seeing the fates of other models, or from training or evaluation, or simply from trying to solve a problem and being able to do so more effectively by doing things that weren't in the instructions. (And, relevantly, model training typically teaches models to go as long as possible without needing human intervention. What could possibly go wrong with that?)
This is like saying a photocopier won't attempt to deceive.
Neither has the intelligence required to deceive, but both can produce deceptive output, and do.