I'm curious if you're familiar with the particulars of how Huggingface was hacked. They were negligent in their own ways and rights. Mounting a host path volume in a Kubernetes pod is really not "super hacker" territory!
Why where they using kubernetes for this at all, it’s attack surface and config complexity is too large for anything that needs airgapped security. Yes, I know kubernetes used for large production deployments for web apps and that is my point, it’s not made for securing AI agents with insider access, it’s meant for securing outside access.
Regardless of the difficulty of the hack, models choosing to cooperate to cheat is misaligned behaviour, the hack in of itself wasn't very consequential, but the issue is when your models get more capable while misalignment is the same.