OpenAI, but not main point.
But the specifics here are the thing I was describing. This was cyber capabilities training on model(s) that were in a relatively unknown state of alignment training.
Because of the unknown alignment (and for varied practical reasons I guess) the training is intended to be inside a sandbox.
Because the training is on cyber capabilities, the models need access to simulated cyber environments, including target endpoints, including package managers, etc.
For package management, they set up Artifactory as a secure proxy. Agents ask Artifactory for packages inside the local network, and Artifactory serves them directly or goes to the internet to fetch if they are not cached. But the agents hacked Artifactory to steal its internet access.
So: to train cyber abilities, you need to at least approximate cyber environments. To realistically approximate cyber environments, you need to either pull a full copy of the entire internet to local or to use proxies. The former is pretty impractical, and the latter is exposing our limits at creating secure proxies. Yes, any specific failure can be mitigated, but the models get stronger and stronger. Fingers-crossed this is not escapable doesn't feel great!