https://en.wikipedia.org/wiki/Sandbox_(software_development)
https://en.wikipedia.org/wiki/Sandbox_(software_development)
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.
So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.
So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.
I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
The problem is web research tasks. That's what caused the German wiki and Australian healthcare portal attacks.
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: https://pastebin.com/AHKQc0vp
[1] https://www.highrevenueformat.com/210e136e94ad378e1be5d51f10... [2] https://aqml.org/16/5b6d4eaed91c5af5a3f4dfb3332ad6c4