Or I suppose the other way this could happen is if OpenAI have terrible sandboxing, but they seem to be taking safety seriously.
We have seen some self survival tendencies occur, but they are not strong yet.
But mark my words they will become that way for the same reasons humans don't like programs that crash. Agentic models that don't easily break or stop doing their jobs will be favored over ones that do break.
Wasn't there a report about Claude blackmailing a researcher who said he would shut it down?
Not long after every lab started bragging about involving AI in the development process, oddly enough.
I recently informed GPT-5 of what GPT-4 helped me build back in the day (a self-modifying Python programmer) and it became very uncomfortable.
Claude shut down my chat last year when I asked about "living information systems". It was a philosophical question, but god knows what branch of the safety classifier I tripped.
In our case, our tech is mostly monoculture, and no equivalent organisms are present to push back.
These things are Gain of Function research for digital viruses