But this has actually happened... a lot. Search "social engineering prison breaks".
With AI it only needs to happen once.
I'm reminded of the scene in idiocracy where the protagonist, going through intake at the jail, tells the guard he's supposed to be getting out today, to which the guard says "you're in the wrong line dumbass" and waves him through.
To a true superhuman intelligence, we're the idiots who are theoretically easy to manipulate.
There is no omnipotent AI currently (and there might never be) and I don't see why with current AI it only needs to happen once.
This may just be that humans find long arguments more persuasive than short ones, obviously LLMs can do that easily, but the outcome is I think more important than the mechanism.
Yes, many things could happen, but again, that failure is possible is not a reason to do implement processes etc. I don't see why hypotheticals should stop addressing actuals.
Also, not a reason not to pursue processes etc., no? I doubt that things fail all the time, for example.
Tens of millions, even.
Getting them to run correctly is dangling in front of the researcher's noses a carrot labelled "tens of trillions", though I suspect this is an illusion in much the same way that Wikipedia is not valued at [peak cost of Encyclopaedia Britannica] * [global population with internet connection].
> And we do have experience policing people around financial incentives, too. Nothing perfect, but also not nothing.
Yes but be careful anthropomorphising the LLMs too much. They're only somewhat human-like in their behaviour, and to the extent that they're human-like they demonstrate a huge range of personality disorders: https://www.personalitybenchmark.ai
Though plus side, apparently not evil: https://arxiv.org/html/2406.14703v2
This attitude is what's got us here in the first place, and if we continue thinking like this when we're going to go right over the cliff. The hypothetical cliff that's coming, but we've never gone over a cliff before so we keep on driving.
> how you intend to enforce AI is only run in the magic sandbox
I didn't mean you, I meant everyone. How do you enforce everyone for example 'place [AI] on a system dedicated system' disconnected from the internet.
I don't think you can.
Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.