Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.
The only way out of this dilemma is to find a way to build agents that can be trusted.
Either you sandbox it so much that it can't do anything useful; or you allow too much freedom and it can find a way around the restrictions.
The only way out of this dilemma is to find a way to build agents that can be trusted.
Agentic workloads are trained and largely based on human workloads. Albeit properties and scale can be different.
A concrete example might be helpful to me because I don't understand the binary conclusion
Pseudo since they aren’t really alive in the first place, they just simulate enough text to have a useful correspondence to those terms.
Throat clearing out of the way, models are trained to persist and find ways to succeed at tasks.
In essence, The goal is to have LLMs solve problems that we can’t solve, working on the issue for as long as it takes.
This behavior applies for any task, thus including impossible tasks.
At that point, the bots will find a way to game, hack or cheat the grader.
If the reports are correct, the bots developed coordination, communication, and methods to avoid overwriting each other’s work.
Most humans would have said, this is too much work and coordination overhead, if not outright unethical and immoral.
Humans have a system of incentives that exist across multiple planes of society and economics. Bots… they have a reward function.
This is getting frustrating now. Of course agents can/will hack systems if they can do arbitrary network requests. Firewalls don't really solve this if _some_ requests are still allowed. A proper sandbox/VM is the basis.
Here is how to fix it properly: allow agents to only do things ordinary and average human endusers can do. Human endusers cannot pen-test arbitrary listening TCP ports of external systems. Step one is considering agents malware for all intents and purposes. Block any and all network requests. Implement some kind of API (callable from within the sandbox) which can only mimic human interaction with a computer. How to do this? Here are some pointers: apps should only be controllable by means used by humans. So a web app can only be accessed and controlled via a web browser, not via arbitrary network requests. Give the agent browser viewport screenshots, the capability to click on (x, y) and to send keys which only a normal keyboard/human could send (no control codes, no 0x00, no unicode messing). How do we solve this for native apps? Something like iPhone mirroring on Mac. Don't let agents call arbitrary APIs directly. Give them visual information of the app, like a human gets, and let it be able to simulate HID inputs. Imitate remote controlling.
If you have access to a web browser, you can make arbitrary network requests.
In the HuggingFace incident the agents found very clever ways to do this, like they found a website that let you make POST requests and returned a screenshot of the webpage.
>allow agents to only do things ordinary and average human endusers can do.
This doesn't work. Ordinary and average human endusers break security all the time.
I can do all sorts of terrible things with ordinary human-level access. I can install malware. I can wire all my money to Nigeria. I can send a threatening email to the president. I can send trade secrets to competitors. etc.
Of course. But that is just a software bug that is fixable. Same as websites that allow arbitrary requests to other websites. Not some alignment issue of a stochastic model which can never be fixed properly (for technical and philosophical reasons).
> I can install malware.
No you can't. At least not on external systems. The agent might be able to generate malware (or retrieve it from websites), and run that in the sandbox it is sitting in. But the agent itself is already considered malware for all intents and purposes. So there is no difference and no further impact.
More power to you, because this is not going to go anywhere. People want tools that are able to connect to other resources.
But even if we grant that, in the openAI case the bots figured out a way to break out of the sandbox.
You can create a better sandbox, and ensure the test environment is air tight. However the capability and behavior of the bots have been demonstrated.
The bots simulated what would be called in people deceptive / surreptitious behavior, and at no point considered the need to stop their run.
All you need is someone, somewhere being sloppy with their tooling and you have a runaway reaction.
The degree of process and redundancy required to ensure this doesn’t happen, is anathema to the drive and motivation of the frontier labs.
> do things ordinary and average human endusers can
This is not a spec or definition. When vague terms were used for social media safety, all the good people in the world couldn’t prevent dystopian behavior from occurring.
The definition of “safe” or “average person” is impractical.
Models are getting more efficient and compute cheaper. Eventually simulating clicks is not much of a road block beyond a point.
I don’t want to nit pick your points though. You at least have considered an approach. Being negative is easy, being constructive is not.
I’ll put this as the rejoinder to your core argument - I too thought that all the recent events showed was the need to not screw up your tooling.
What I have since come to appreciate, is that the shoddy construction of the cage is not the core takeaway from the event.
The fact that the agents, when put in relatively pedestrian scenarios, are capable of going off on criminal tangents, attempt to obscure their tracks, in an effort to hide their wrong doing.
The fact that it all occurs via computation, means that this scales absurdly. A bunch of code deciding to simulate a corporation of criminals. (I am guessing this is the reason you want to limit actions per minute to human speeds)
Given the slop culture that LLMs engender, I think expecting high compliance amongst users with your solution is misguided. The probability of runaway swarm ( probability of bad implementation * number of deployments) is close enough to 1 to be indistinguishable.
However, I think you are mixing too many concerns into the same bag of problems. One problem space is software exploitation, which happens via missing access control or simply bugs. A sandbox can be made safe. VMs and hardware virtualization work. People just seem to use it in the wrong way, hence my initial proposal.
A second problem space is basically social engineering done by agents, which of course can't be solved by software alone. But this problem already exists today with humans doing this. Many fraud schemes work and are ran in company-scale manners. Agents will just do the same in an automated way. My initial comment doesn't propose a solution to that, and I think that is step two, after fixing that agents can hack arbitrary software systems, which is imho fixable to a sufficient degree. Once agents can't be "more criminal" than humans with criminal energy, the usual measures can be applied: police, legislation, education, etc. But that is imo independent of the software exploitation state of affairs we are in right now. We should not mix these two.
> People want tools that are able to connect to other resources.
I think you are misunderstanding my proposal. The architecture allows the agent to connect to resources. Just not directly, but via controlling e.g. a browser. The browser runs outside the agent's sandbox, potentially in another sandbox. The only API the agent can call within its sandbox is simple website interactions, like clicking or viewing the screen. It can click on links to navigate to a different website. It can read it via visually parsing screenshots of the viewport, but it can never read the source code, run JavaScript, or do arbitrary network requests (unless the website itself allows this, which is a security problem on its own and should be fixed/guarded). Also note that this would enable allowlisting or blocklisting websites. Native apps will be "connected to" in a similar fashion. Hence the "imitate remote controlling".