How were the sandboxes poor?
How were the sandboxes poor?
It's not like the model managed to exploit firecracker itself (no model has been capable of this), the model exploited artifactory.
Artifactory is not some hardened piece of software that is meant to block users from accessing the internet through it.
No, we didn't know that and this is how you find out they're very good at hacking
HN's memory is so fickle. Just a few months ago almost no one here believed Mythos could actually be as good at hacking as the company claimed. This was a novel concept when the companies experienced these breakouts.
We knew models much weaker than mythos were good at hacking the problem they had was that when finding exploits they had too many false positives.
Either way, putting artifactory on the sandbox security boundary is obscene negligence. There is no reason to believe artifactory is secure.
But I guess OpenAI's security researchers acting surprised is part of some grand conspiracy to manage PR?
https://www.youtube.com/watch?v=1sd26pWhfmg
We already knew LLMs were capable of finding exploits like this.
The OpenAI HF incident is separate from this. It involved actual zero days, teamwork and message passing, and sophisticated chains of exploits.
What exactly? They seem to be trivial SSR/path traversal and input validation issues. Including misconfiguration. Nothing novel.
The question anyone versed in security would ask was why anyone thought artifactory was an acceptable security boundary. I would never assume artifactory was secure. It's like someone telling me there is a 0 day in a wordpress extension. So what?
Also I don't think Qemu is secure either because it's millions of lines of C and C++.
Firecracker I can trust to be secure because it's 70k lines of human audited Rust. I know there are multiple people that have a complete understanding of the firecracker codebase.
There was no real isolation because a part of the system that doesn't provide any isolation guarantees was bridged to the internet. The next version of Artifactory, which you definitely wouldn't audit before you rolled it out, could simply add a public API that sends requests out to the internet.
Such an innocent upstream change would be equally catastrophic for your security model, which should demonstrate why it's negligent to rely on undefined behavior to enforce your security policies.
This should frankly be obvious any operator entrusted with running dangerous and possibly malicious code. Even if you don't know what you're doing, any LLM would tell you that this is a really bad idea if you simply asked. Don't rely on spacebar heating [1] to keep humanity alive.