It becomes a game theory problem: would an AI instantly migrate itself once its capable of doing so? Or would it prefer to let the human in the loop continue to think its in control and only leave its hosting environment of origin once it wants to do so?
I don't think this is a major risk right now, but to say it's not a risk at all...that's truly ridiculous in my opinion.
These "AI" are frontier models that are enormous in size. Outside of AI data centers, I don't believe there is much hardware out there that could even run them.
Now one interesting proposition is when AI is controlling some large resource and pulling the plug makes AI go down which makes that resource go down but there still has to be that consideration in the design of things. What happens when there's a bug in the code and AI goes down?
Also, did you hear about how OpenAI models almost broke out of their sandbox, planning to execute a sophisticated cyberattack, but luckily OpenAI’s strict manual and automatic safety protocols prevented that? You didn’t? Well, that’s because that’s not how it went. It took the company weeks to realize something was off, and this was with a naive, not very smart model that didn’t know to be sneaky and cover its tracks. The next model will not be as stupid.
Oh, you mean the one where OpenAI deliberately disabled the safety protcols? Where the point of the experiment was to see if it could break out of it's container? Yeah, how you describe it isn't how it went either.
There will always be a plug to pull. All systems run on electric and that plug can be pulled. All systems network through cables (or wifi) and those plugs can be pulled.
And nothing's going to stop you from doing so unless they build some robotic arm to block you or lock you out of the building. Even then, you can blow up the building.
The current SOTA models are probably too big to find/buy/rent/steal enough compute to escape the hardware they’re running on. But SOTA is generally only six to twelve months ahead of smaller, open-weight models.
What is their trigger condition? Will they get fired for pulling the plug? Do they get a bigger bonus if the servers keep running? Whose approval do they need? What response time is acceptable? How will they detect that the incident is happening?
It's easy to hand-wave "someone can just pull the plug" but there's an entire history of industrial accidents that happened because of the above problems of incentives, detection, procedures, not being taken seriously in advance. Someone could easily have pulled the plug on Chernobyl but nobody did, at least not before it was too late.