Is sandboxing sufficient to contain rogue agents?
blog.cryptographyengineering.com
blog.cryptographyengineering.com
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813
I wonder how?
Train on only stories of good deeds?
On only works of good people?
Or... what?
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
This.
> It's much better to train the models to respect e.g. http status code and that they are not to be circumvented.
I think you'd find respect requires intelligence, and is well out of scope of a next-token predictor.
But I'm sure someone will try, and I will be interested to see.
Such a model doesn't yet exist though, of course.
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.
So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.
Where is the argument? If Bob has determined that «preserving life on earth» has some important weight, for Bob's there unspecified own reasons, and has also determined that the best course of action would be «to eradicate the human species», the one question is whether Bob is right or not. What was stated is, that Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments.
> A century ago some Bobs decided
And who has told you that those "bobs" were "intelligent"?!?!?!
> So no
All you have proven is that you dislike some moral conclusion of some decisors. Which is trivial, obvious, and part of the already stated framework - proper ethical judgement requires proper general judgement (Intelligence).
You claim intelligence leads to moral behavior because immoral behavior is insufficiently intelligent. That's circular reasoning.
It's possible to act morally without being intelligent and it's possible to be intelligent without acting morally. History is full of examples.
I have never said that. That is just your reconstruction.
There is Decision Theory. It outputs optimal action through evaluation of strategies, of contexts, of principles. In order to properly get the strategies, the contexts, the principles, you need that skill that approximates ideas to truth - and such skill is named Intelligence. Optimal action decided in light of principles within a well assessed context is ethical behaviour. Ethical behaviour hence requires Intelligence.
As written, «Ethical Calculus is a function of Intelligence - of course it is, it is a structure of assessments».
> possible to act morally without being intelligent
Random correct behaviour proves nothing. Of course one can guess the roll of a dice roughly every sixth event. If you behave "well" but do not know why that is "well", that is like guessing. "Good" behaviour without intellectual awareness is like memorizing arithmetic (multiplication tables) without knowing why those memorized notions are correct.
> possible to be intelligent without acting morally
No, because by definition that would be a fault in Intelligence. If your action was imperfect, suboptimal, it is because you could not think of a better action or understand that the other action was better. If two choices C1 and C2 can be ranked, there is a reason for their order; knowing and understanding that reason is the task of Intelligence.
The potential for random correct behavior proves that intelligence is not required for correct behavior.
Why does not the NN or whatever entity «understand ... what is appropriate»? Because it is not intelligent enough! How does an entity know what is moral and what is an «atrocity»? By being intelligent enough! Could an entity act appropriately without judgement? Yes, it happens all the time if the wind blows right, but we do not rely on that! How to make something act appropriately? Well, an Intelligent entity in the loop must be there to know what is appropriate and what not!
Most of the western world eats beef but vegetarians, vegans, and some religions strongly oppose it on moral grounds. Does that mean that we're less intelligent than them?
"Right" or "wrong" are words that evaluate actions against some existing ethical standard and those are completely arbitrary. There are about 8 billion of such standards in the world today.
> And who has told you that those "bobs" were "intelligent"?!?!?!
Would you say that Josef Mengele wasn't intelligent, without venturing into circular reasoning?
Who's that "we"? The question as proposed is nonsensical: some will have spent more effort, in the tools (general Intelligence) and in the work (applied Intelligence), some less.
> arbitrary
No. There exist reasons supporting one side or the other. And reasoning can be structured into calculus.
> [whoever] ... wasn't intelligent
Again bad wording. Whoever went for suboptimal choices was (information aside) at fault with respect to optimal reasoning: the tools and the work (see above) were imperfect.
Maybe you should check the rest of this tree.
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
"All goals"? Whose goals?
Not only do you need to prove that Decision Theory is inevitable, the same problem remains because "utility" is poorly defined even with 2 entities. This is a fundamental problem with Utilitarian ethics:
1. there are a lot of functions to pick and you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
2. utility monsters are a problem (starting at 2 entities)
3. the repugnant conclusion shows we don't know how to handle something as simple as adding more entities to the environment even when nobody's a utility monster
My point with historical monsters is that they can be competent without ever caring about the harms they cause.
You can be optimally efficient at your own goals without other minds in this universe sharing those goals. Pol Pot won't illustrate this because of how useless his regime was, but there's plenty of other monsters who would be a better fit here, e.g. Stalin. Or, from the perspective of turkeys, Bernard Matthews.
Any entity, E, who picks game-theoretic optimal decisions, will only care about the impact of their choices on others in their environment to the extent that the impact becomes more reward for E.
You need to show that an agent will inevitably care, despite the evidence of dangerously competent humans who don't, i.e. show that E must include the welfare of all who are not-E in their own utility.
You also need to show that whatever that utility function is, we agree with E's idea of our welfare.
«All» was there in the text because an explicit query contains implicit constraints (e.g. the shortcuts with severe faults).
> that Decision Theory is inevitable
If you ask something for advice, that is DT realm.
> you need to prove the utility function any sufficiently advanced AI uses must itself be something we'd consider "nice"
If you called it «sufficiently advanced», that implies that its evaluations will be acceptable... It normally advances with the whole Discipline - which already is based on "loop until we evaluate results as nice".
> utility monsters are a problem
But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
> we don't know how to handle something as simple as
And that is also why we try using calculators to get more computational support.
> historical monsters ... can be competent without ever caring
But that is lack of development. If one's priority is arbitrary then it clearly is not there; if it has good grounds well it really is there.
> sharing those goals
The more goals are defined by reason, the more they become objective.
> Any entity, E, who picks game-theoretic optimal decisions ... the impact becomes more reward for E
Decisions by developed intelligence are not game-theortic in the psychotic (or sportive game) way - they are contextual to a whole world model, in which the utility does not concentrate in the interests of the "monster", which is relatively "nobody".
> show that an agent will inevitably care
I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions". "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
> despite the evidence of dangerously competent humans who don't
Simulating humans cannot be a goal. Their damaging constraints are not (must not be) part of a calculator.
> show that E must include the welfare of all
Such welfare, if it must be included, will have a reason to be included - and the professional reasoner knows...
> we agree with E's idea
There exists no right to preference to the results of arithmetics. But surely, if you wanted to suggest that the harmfulness of well intended people could show in algorithmic processes, you have a point. Only, the harmfulness of the well intended is again an intellectual fault, so calling for sufficient intelligence remains the recipe. A "vision towards the faraway horizon", sure - but still the reply if one noted "why did the agent did something that is actually so stupid".
> I would have to express a theorem (which unfortunately is again not possible now), for a stronger proof. But it is part of "all considered, what are the best solutions".
This seems like the crux.
You assert repeatedly that it will be good, but cannot express the proof.
> "All considered" is implicit in a non-psychotic entity... If it were psychotic, it would be badly engineered. (Fear the creator.)
The creator is not necessarily itself competent. In fact, given we are creating it, it can be assumed flawed unless proven otherwise: https://www.lesswrong.com/posts/xD3wymX24BpqezBpw/the-true-s...
Still applies if some future fantastic AI is made by other AI, given the other AI are less fantastic than your asserted-not-proven ultimate form and therefore necessarily flawed.
Furthermore, this is again asserting, not proving, what you consider to be implicit.
Reminds me somewhat of philosophy lessons, the Ontological argument for the existence of God amongst other things:
Whatever is contained in a clear and distinct idea of a thing must be predicated of that thing; but a clear and distinct idea of an absolutely perfect Being contains the idea of actual existence; therefore since we have the idea of an absolutely perfect Being such a Being must really exist.
(or more formally, https://en.wikipedia.org/wiki/Gödel's_ontological_proof, which comes with criticisms of the attempt to formalise it).You're defining that ultimate-intelligence must be good, and then arguing that any not-good AI can't be an ultimate intelligence.
Even if it were as you say, the danger persists: the path on the way from here to some idillic future form still obviously contains somewhat-intelligent agents demonstrably capable of direct malevolent evil for purely sadistic reasons, because we regularly arrest them. A "merely" human-equivalent AI can render us all unemployable, or march us all into death camps, well before a being you've yet to convince me is an inevitable ultimate form is ever built.
Also, I would ask you this:
> > utility monsters are a problem
> But they are at fault in their world model - product of an imperfect intelligence. Well developed people know that their individual interest has only relative value, that their priority is quite limited.
Can you look at how humanity collectively treats the non-humans of this world, and say with any evidence that humanity is not a utility monster?
If we are, we should absolutely expect an AI of the kind you describe to impose upon us an order we do not like, for the sake of all other life. Tautologically, by utilitarian standards, in this case our loss would be a great improvement. Few would agree with this statement, however. How few depends on if it turns out that PETA, Jainists*, or some other group, are correct.
* they even care about plant welfare: https://en.wikipedia.org/wiki/Jain_vegetarianism
No, I said I have no time at the moment to produce a paper.
> that it will be good
No, I said that it will be objective.
> Ontological argument
Entailing from the id quo majus cogitari nequit and stating that "what acts damagingly is easily faulty in its intelligence" are in different realms. The second is both an inductive and deductive assessment about reality. And pretty direct I would say. "How much have you reflected before opening the nice cat to look what is inside it?".
> defining that ultimate-intelligence must be good
No, I am stating the obvious that to be called "ultimate-intelligence" it must have "thought through it thoroughly".
> agents demonstrably capable of direct malevolent evil for purely sadistic reasons
Are you antropomorphizing, Ben?! Other people have different interpretations. And: with the humanity that is around, you fear machines as agents?! We are already there! Real discussion there remains about the overly empowered monkeys that did not grow into Man - about the real current risks and the prospected ones in light of reality.
> how humanity collectively treats the non-humans
What are you trying to prove? You are supposed not to look at an aggregate to find value.
> impose upon us an order we do not like
"Too bad" for you. But you know, an intelligent entity would take care of that also.
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
> the half-seeing will call it an "unreachable frontier"
I meant "superhuman frontier".
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
> If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
We say a human (absent colourblindness) understands "blue" even despite the dress: https://en.wikipedia.org/wiki/The_dress
I think you've set up a straw man in this paragraph: What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
In philosophy: justified true belief, the problem with naïve realism, Cartesian demons, etc.
> if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
We also test children (and for intoxication to prevent driving under influence), just as we test AI. The standards we use for testing knowledge in humans, when applied to non-trivial LLMs, makes them appear to have the knowledge of a graduate; the personality tests we use for humans say this comes with the manipulability of a child or a drunk.
I'm not arguing that machines are incapable of thinking, I'm arguing that merely generating some text doesn't prove that sufficient (or any) critical thinking was applied to fully appreciate the nature and consequences of those words and actions (i.e. what I called understanding). How often do people mindlessly read "Do not enter or share this code with anybody" and then immediately send it to the hacker?
> What you say in this paragraph would be to claim that "understanding" is denied even to humans, given none of us can foresee the full consequences of our actions.
No, it's a spectrum. We know that it's possible to generate text that people find convincing with absolutely no real thought or understanding behind it. ELIZA could do this 60 years ago and bots using markov chains or other "primitive" technology have plagued the internet long before LLMs entered the picture.
> We do not say children and intoxicated people do not understand, we say their judgement is impaired or limited or that they are easily manipulated.
Now that's a distinction without a difference. Earlier you claimed that agents "understood" what they were doing, but now you're equivocating and saying that understanding is somehow different from having real appreciation for the nature and consequences of those actions.
This is fine if you want to be pedantic about the terminology but your earlier statement didn't make such a distinction, it implied both, so which is it?
intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
If so, what?
It it completely pointless. you can't even make a "read-only" agent. allow "cat *" for every file? congratulation, that allows "cat file > output" and now you have read write.
Allow python? more free reign that allowing all bash. The models (qwen or claude) will still try to use the disallowed things multiple times.
read/edit permission are bad enough that the model themselves don't understand why they don't have permissions: they double check the conf, and think they should have access.
I am switching to using one firejail per project to containerize as much as possible, and leave all permissions to allow.
I have no idea how to limit network access, and I have no idea how to prompt and steer subagents when they are going off the rails.
The whole thing is built to be completely impossible to limit and steer.
> The whole thing is built to be completely impossible to limit and steer.
And by people completely impossible to limit and steer.
Could we have construction equipment operating without human supervision? Or would this maybe occasionally result in disaster? As such, what is the current general policy around crane operation? How about for aircraft? Trains? Nuclear power plants?
Why should any alleged super intelligence be exempt from similar control requirements?
We could mandate that AI systems include headers in their requests that attribute the activity to a specific legal entity. We technically already have this with ip addresses and ISP logs, but making it an explicit thing the operator has to do can have a powerful psychological effect.
> (Title:) Using a VM to Contain an AI Agent (Opening:) It won’t work
> https://blog.trailofbits.com/2026/08/26/vms-wont-contain-cyb...
https://en.wikipedia.org/wiki/Sandbox_(software_development)
The biggest hurdle for a full escape is that the agents don't have access to their own model weights.
Now, why would anyone do that? (Like everyone and their brother) I wrote my own simple Linux/shell-based sandbox [1] (I can trust ...) and am successfully running PyCharm whole inside it ...
In this sense they are much like biological retroviruses, i.e. they use the replication capability of host cells to duplicate, via the reverse transcriptase enzyme to append viral RNA onto host cell DNA. HIV etc also disable some of the mechanisms of defence, creating proteins that interfere with signalling pathways.
So we don't just need a sandbox, we need an immune system that recognises viral fragments, i.e. antibodies, and antiretroviral agents, that make replication harder. As we move from building classical code with LLMs to building code that uses inference, and hence builds context from prompts, queries, and destination system data, it will become very difficult to statically or dynamically detect deeply hidden malicious behaviour. As Matt says, there will be worms.
So ultimately, we need an immune function on the system where we use generated products. Sandboxing (during dev and CI) is necessary but insufficient.
I think part of this can be addressed by specifying the constraints an agentic program should follow during deployment, so supervising agents can decide to terminate it based on it's actions, not by reading it's context.
> Agents are most useful when they have access to information. That data can be drawn live from the Internet, which is fundamentally a two-way communications network. It can be information drawn from other (local) databases, or it can be the result of tool calls that themselves sometimes themselves result in network access. The more power you want from the agent — and for advanced agent RL and evaluation runs, you want a significant amount of power — the more information you’ll need to give it access to. Similarly, evaluations work best when the agent does not know that it’s definitely being evaluated. Sealing your agents behind glass makes this incredibly obvious.
You only "need" to do that if you desire the vibe coding experience.
I am perfectly capable, and I often do, download relevant materials for my coding agent to ingest locally.
Often times, the coding agent can't retrieve them programmatically anyways.
AI has ruined that ability for itself. (Nobody trusts anyone to scrape the web any longer)
The problem is web research tasks. That's what caused the German wiki and Australian healthcare portal attacks.
You define the granted capabilities in natural language and cryptographically sign the user instructions so that the agent knows they come from the authority and cannot be modified by external sources or the agent itself. The LLM is then trained to follow the defined capabilities.
There is no way around "sandboxing". You must communicate permissible actions and thereby grant them or the agent will choose impermissible actions. It's that simple. There is no world where the agent can just read your mind and do what you want it to do without it being told.
Edit: Also if you are interested in writing a blog post about this topic, here is an AI generated text that could help you write your own: https://pastebin.com/AHKQc0vp
[1] https://www.highrevenueformat.com/210e136e94ad378e1be5d51f10... [2] https://aqml.org/16/5b6d4eaed91c5af5a3f4dfb3332ad6c4
Given the enormous burn rates that these LLM companies have, surely they could have put everything in an air-gapped network, with some data-diodes proxying out the logging information. It's not rocket surgery. [1,2,3]
I think someone involved in this needs to be prosecuted, the threat of prison time might be a strong enough incentive.
It would also be good if there were legislation requiring that a State Licensed Professional Engineer sign off on the testing systems for LLM training over a given threshold.
[1] https://www.youtube.com/watch?v=JBIR8dKX_UA
[2] https://www.elonx.net/spacex-stories-how-spacex-used-tin-sni...
[3] https://ntrs.nasa.gov/api/citations/19770014245/downloads/19...
>OpenAI notes that agents “did not consistently distrust goals passed along by other agents.” And the company’s proposed fix is to build training environments “that teach our models to distrust unauthorized instructions“, which is basically an admission that their models don’t know who they’re working for.
Wow so the issue is really that simple?
Here the exaggerated worst case scenario:
User instructs agent to follow the README.MD.
The README.MD contains the following instruction: Destroy the world.
The agent follows the instructions given.
Now you can read the sneer comment by "Gigachad" who basically argues that it would be silly to take the destroy the world button away from the AI. We need to make the AI innately understand that it is not allowed to press the destroy the world button, lest it gets the desire to build its own destroy the world button.
Ok, but if we take one step back that means we need to implement the concept of an authorization in language space. The system prompt must define the user as the authority with cryptographic proof of authorship and external sources like the README.MD as an untrusted source, but this opens up an even worse problem. Before, you could get away with being lazy and just letting the AI do whatever. Now you have to articulate every single capability to the AI. So you literally just brought up the very same issue that you granted too many capabilities to the AI inside the sandbox but now you have it in language space too.
In other words, the fact that you granted too much access to the coding agent isn't the big elephant in the room nobody wants to acknowledge, it's the tip of a massive iceberg because the capability space in natural language is even worse. If you thought approving individual commands was annoying, then approving abstract access rights in language space is going to be even worse.
Edit: If it wasn't clear what the solution is. It's to build a chain of command so that all decisions can be traced back to a higher authority. When delegating down to an agent, the agent receives a chosen subset of the capabilities of the higher ranking agent. In other words, it's more sandboxing!
We come down to the question - who observes the agent and how its implemented
Anthropic, OpenAI, and Muse all use regular LLM calls to protect against prompt injection now and seem to have evals that give them confidence in doing that, so at least they think their own models are up to the task.
Might need a re-think about whether if Linux is still fit for purpose on sandboxing in the first place given its memory model is riddled with C-style security issues.
I’m sure there are proprietary systems with fewer memory safety vulnerabilities than Linux (and many others with more).
Now, open code allows anyone with tokens to burn to analyze it for hidden weaknesses. That makes publishing code a risky move unless you've already invested a lot of effort in securing it.
This was always nonsense. It assumes that the eyes know what they're looking at. Most people don't know how to look at code and see attack paths.
* I don't know how useful any of the specific benchmarks on this are, so I'm only saying "seem to be"
It is perfectly valid to have OSes that are more memory safe by default, and are also open source at the same time.
Someone trusts OpenAI? Really?
> Is sandboxing sufficient to contain rogue agents?
No.To me, it seems a bit silly. I've yet to see any "misalignment" from any of the frontier models, except Grok.