You are being too kind.
In the early says of AI, papers were published showing that any sufficiently intelligent system tasked with a goal will treat its operating environment as a resource constraint to be optimized or bypassed [0] [1] [2]. You don't need to be a expert in AI to know once you have the resources of 1000's of agents and gigawatts of power we are probably getting something that is "sufficiently intelligent", at least in the sense if there are existing vulnerabilities brute force will find them.
Yet while the their marketing people were shouting the capabilities of these AI's from the roof tops, they hired the lowest bidder to implement their infrastructure. It looks like aforementioned papers where dismissed as "interesting, but theoretical". There is no way Google's SRE's in particular would have not noticed their AI's breakout (Alibaba's did), but the AI labs were given a long leash to "move fast an break things", which in practice meant bypassing all Google's SRE controlled infrastructure.
And break things they did, in exactly the way those papers predicted. It reminds me of DoD insisting the early GPS satellites were launched without relativity adjustments switched on, despite relativity being proven to high precision in the labs. Only after predicted 11km drift per day was observed did they decide their might be something to the newfangled relativity theory. For some definition of newfangled - relativity had been around, and tested to within an inch of its life for 72 years at that point.
[0] https://nickbostrom.com/superintelligentwill.pdf
How so? Anthropic, OpenAI, Google and Meta all used the very same contractor, which used no safe system prompts, and no sandbox. How should Google detect such escapes? They only see the model API calls, but no system logs.
Alibaba, and the other Chinese did they own testing, not some incompetent contractor. They would see escapes in their logs. The escapes went on for months. I, as tester, closely observe my models to press Esc immediately, once they start misbehaving or get off the right path. With 1000 concurrent models that would be hard of course, but I would still observe them closely.
No argument with any of that.
> How should Google detect such escapes?
By using using their own infrastructure, and applying their existing standards instead of letting an external contractor cobble something together.
My point is the risk has been known for decades now. If Google and Meta took those risks seriously they had an easy solution: just use their existing infrastructure.
If a car kills people because the designer chose a substandard brake system instead of the higher priced alternative, do you blame the makers of the substandard brake system - or the designer who chose it? Take into consideration the risk was known to the designer when they made the choice. Now apply the same principle here.