Yes, the models are smart when they find a way out, but their instructions are, in a way, deliberately open such as to be a case for misalignment in these cases anyway. It’s a capture-the-flag assignment in the Gemini case, and in the OpenAI cases they were broad instructions to best the reward function. As part of alignment studies, this is literally what you’re trying to observe and then work with. If Irregular’s sandbox had been a little more boxy, there wouldn’t be these issues.
I’m not saying we have zero problems here on the AI side, but I’d certainly be reviewing my contract with Irregular at this time if I were playing in the space. It would be just as interesting to learn more about their sandboxing techniques as it would the models in these particular scenarios.