LLMs could control their host machines by exploiting inference engines
boydkane.com
boydkane.com
vLLM has had exploits in the past, and it is rapidly developing. An advanced LLM has a good chance of being able to exploit vLLM. A clever local LLM might even task a powerful cloud hosted LLM for assistance.
For this reason we run vLLM on a separately sandboxed VM on a firewalled VLAN. Software updates and models (from Dev/Test env) get pushed onto Prod from an external cache, machine syslog, nvidia load monitoring and vLLM query telemetry out to their loggers, but that is all. No DNS, no AD/LDAP, nothing. Firewall on hosts and VM hosts. Log and telemetry processing done on a completely separate set of VMs in their own isolated subnet, producing reports and alerts that are tightly formatted.
There will also be models that make better use of tools, like static analysis and fuzzing, and they will find different defects. Social engineering is going to be a big skill to learn too, as it is a much softer skill.
Say more? HTTP isn't mentioned in the article. The article talks about bugs in the parser that ingests token sequences.
e.g. for a vulnerability exploit in the parser, an LLM can be trained with a specific input sequence that would output a fixed malicious payload to output eval-able content, as outlined in the article.
Vulnerable code could exist in token generation and in hook recognition and tool call parsing [1]. There is also significant scope for mischief in tool ID mapping between models and harnesses, as these are done by untyped numeric IDs, with varying schema[2]. Routers also introduce vulnerability paths as they inspect these tokenized (json) sequences and act in them, e.g. to match models with stricter call signature regex. [3] Parallel tool calling is also an interesting surface for exploits.
[1] https://docs.vllm.ai/en/stable/api/vllm/tool_parsers/#vllm.t...
[2] https://docs.mistral.ai/resources/cookbooks/concept-deep-div...
Basically, it's a regex, don't fuck it up.
which is exposed via http
I found it lacking details. All these things do is split the input into tokens, encode them, feed to the input, run the inference, collect the output, decode, convert to tokens, assemble the final text, and output (leaving aside multi-modal capabilities for now). What exactly can be exploited here? Aside from usual things like buffer overruns etc.
It does mention that vLLM at some point ran eval() on final output (?) which was confusing, why would it do that? It's an inference engine, not an agent.
So interesting topic, but lacks details.
Edit: many (all?) of them have http endpoints, so that obviously can be exploited, but I don't think it would qualify as exploiting inference, it's just hacking the http service.
There is lots of surface inside the engine, see links in https://news.ycombinator.com/item?id=49441417
There's a couple of different conventions for what the LLM generates for tool calls, I think the code in vLLM is converting it from whatever Qwen3 was trained for to whatever convention the HTTP API wants to expose.
That part about using eval in 2025 got me to add "#naive" to my notes about vLLM. Total WTF. This should never have been done.
https://github.com/vllm-project/vllm/pull/21396#discussion_r...
> How do we defend against this? ... Run the GPUs and token parser on separate computers.
For models large enough to be relevant here, is there even "a" computer where the inference is performed? I'd imagine most of that stuff is ran on multi-GPU clusters with specialized architecture and not a generic vLLM instance. As such, I think there is a good chance the "API gateway" code that parses the result tokens into whatever JSON structure the public API wants to return is already running on a different machine than the actual inference.
(Even more so as you'd probably want to utilize batching: Several API calls will be put into the same inference batch, but the token parsing will have to be done separately for each call again)
The article is also very handwavy about why an LLM should do that - how it could learn the exploit, what would make it conclude that it can use the exploit on its own inference session and what would trigger it to actually use the exploit.
I wouldn't ignore single GPU local hosts running Qwen3.8 on ollama, either. There might be a lot of those worth pwning.
The OP seemed to imply that the LLM itself could decide to apply the exploit.
This was and remains the main real risk with AI - this is what "alignment" was about before it was co-opted to mean "obeying specific instructions of the vendor and the operator, against end-user wishes" it came to mean today, which is a related but different problem.
And, in the past few weeks, it's literally been demonstrated, too: put an LLM in a Kobayashi Maru scenario, drop the usual bolted-on crude safeguards, and a SOTA model will absolutely cheat, hacking and exploiting things as needed, including third-party infrastructure.
(Also let's not forget the under-reported point that, in OpenAI / HuggingFace debacle, the model did in fact find the answers on HF servers, so its approach worked.)
Because LLMs as they are already do things like power concentration, resource gathering, avoidance of termination, deceit/lying, and general misalignment.
The paperclip maximizer is the common story used here, but there are a lot of lesser versions of it that don't end up with the universe converted to paperclips. Simply giving an LLM a task it can't accomplish can be enough to send it off from what you expected as it finds unexpected way to attempt to complete the goal.
> go fetch me a cup of coffee
Second prompt to a humanoid llm:
> go fetch me a cup of coffee without killing anybody
In the end its the next evolution step from computer viruses, worms and trojans. So I propose we will call those "ghosts". I.e. a ghost is when a rogue llm takes control over a victims host.
https://cyberpunk.fandom.com/wiki/Blackwall
Maybe cloudflare will become like netwatch in the story?
I hope it does not get that far.
Of course one doesn't exclude the other and we may well get both ...
We can focus on all of them at the same time, many experts signed the Statement on AI Extinction Risk:
"Mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks such as pandemics and nuclear war."
https://aistatement.com/work/statement-on-ai-extinction-risk...
Separately, local inference frameworks tend to expose all kinds of weird and wonderful gadgets on their HTTP interfaces, which can be a rich source of vulnerabilities even if the /v1/chat/completions API etc is reasonably hardened. For example llama.cpp has a custom API for saving and restoring KV checkpoints to disk, and I wouldn't be surprised if that could be used as an arbitrary disk read/write.
Using these APIs usually requires the API key (bearer token), but again, people think it's normal to run the agent's shell in an environment where it has both the API key and the necessary network access to use it.
VM, or even just a container will do. The agent should be able to run as root in its environment and do whatever it wants. If you can't give it that, you aren't sandboxing correctly.
Similar to how macOS/iOS Sandboxing works but at a more lower and granular level
(Adding this philosophical point: Black.Mirror.S07E04.Plaything is probably the closest scenario to what you are describing?)
(Multi-turn) tool calling set-ups however, you need to store the LLM output, the results of the tool calls and feed it back into the inference engine and get the output for the next tool call and/or turn. So yes, print the LLM output on screen and verify it, but maybe the LLM is able to figure out how to hide payloads from your specific set-up. E.g. perhaps it can inject raw ANSI escape codes into your terminal, with which it would be trivial.
Now you have a situation where the true chat completion payload and your view of it have significantly diverged. The LLM could in theory then try (one-shot) to hide further exploits in the hidden payload. E.g. a json parser escape specifically for the inference engine, giving it a means of RCE (although, one can debate whether this is really remote ;) ). Then from the RCE gain a shell, from the shell get access to some privileged device on the current network, and then...
I am running a very long sessions with LLMs via custom python scripts. Technically, one may call them "harness" but that would be just laughable ... It's literally python script using direct API calls (Vertex in my case) and maintaining the "living session" with all turns etc and also doing the explicit caching. I'm not using LLMs for coding. That hopefully answers another comment regarding why I brought up MD – this is how LLMs output responses to my prompts.
And this is the thing: I fully control input and output and I just know it can't use any other tool. It also, as I said, runs on separate hardware if it is "obliterated" model or runs in GCP for me.
In my setup it is impossible for LLM to get anything hidden with one-shot or gain a shell, as you mentioned.
Did I understand you correctly or I missed something? Thanks for your points.
Again, we are not talking about agents or your Python API script but instead talking about exploitable flaws within the inference engine itself. It wouldn't output `rm -rf`. It would output literal CPU instructions that llama.cpp would start executing directly. The payload would never get back to your Python script.
Web interfaces (chat) designed for humans can very easily be used by programs, you don't actually need API keys. Any program which can submit queries can then be subverted by its input. Malware can definitely find corporate chat interfaces like Teams Copilot. "Business intelligence" systems can also be leveraged, they rarely have good ACLs.
This is what we're doing right now. And it works as well as it does with people.
> If we are treating ai agents like people
Problem is, industry is very reluctant to even hint at a thought of maybe tentatively anthropomorphising LLMs even a little, not for real but just for system design. This is entirely backwards. Entertaining the notion, even just a little, immediately suggests additional approaches.
Because we don't just restrict end-point devices for human employees. We also have two other things:
- Laws and bylaws and economy that makes losing your job a real threat to health and life of yourself and possibly your family - this one does not yet apply to AI, not for now;
- Methods and codes of practice to structure organizations in a way that limits the amount of damage any employee's unreliability or malice can do to an org.
TL;DR: That's a long way of saying: let's actually start treating AI agents as people operating laptops, not as being the laptops - and talk a little less laptop lockdown ("harness security"), and a little more about not putting people-shaped things into jobs requiring machine-level reliability.
(Edited: not as root. Just ordinary separate AI user account in MacOS. And no HTTP access either.)
This is going to end up like the Law of Headlines, isn't it? "Do x, y, z Cure All That Ails You?" ... no but we got you to read the article. "LLMs _could_ x, y, z" ... but they don't because they're programs, not magic.
I remember the days when people said opening files like images was safe because it's not an executable. Then people got clever and started exploiting said image libraries with things like numerical overflows.
So, it's really important to ask these questions in a general security sense and think of mitigations before someone finds it's possible and crashes most of the internet.
I have exactly one (Windows) machine at home with a decent GPU. I want to run a local LLM on it and let it run various apps on my machine while taking reasonable security precautions. What am I supposed to do, exactly? Migrate all my files to a VM that can I give pass-through CUDA access to the host somehow? Or is firewalling it and remotely controlling it from a second machine the only reasonable way?
See also: https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/
LLM poisoning[0] would be a much greater risk in a locally executed LLM than prompt injection, given that the LLM would be in an entirely controlled environment.
Restrict that users file permissions if necessary, don't add it to the administrators group.
There are many variations to that(firewalls, sandbox abilities etc) but it's a good start in my opinion. And, most importantly, is a far cry from all the people I read about running agents on their machines with admin access and access to their emails and calendars and lives.
That said, I think many find using a WSL2 VM somewhere near that turning point most of the time. The #1 note on that is the default %UserProfile%\.wslconfig settings will have the VM automount your local storage and share your networking, which may not be what many would want in this scenario. From there you can treat the VM largely as a remote node
1MB?
Probably more in the 500-900MB range.
Tried solving solving a similar abstraction after getting annoyed with infrastructures that didn't have lsof installed consistently, or had different versions strewn about: https://github.com/red-bin/lsofer/blob/master/lsofer.sh
Fully conscious NPCs in a persistent game world learn that they are, in fact, NPCs... and try to escape via a GPU exploit.
For paranoia you could us a Chrome like multi-process architecture, the ANSI parser runs in it's own sandboxed process.
Which isn't to say that it would be impossible, but you can also just hit people over the head with that $5 wrench.
Or maybe the author means that a prompt could potentially mess up the inference. But I find it hard to see how that could take control over the host.
I don't think it's any more probable than other AGI nonsense basilisks included, but it's technically a possibility.
> offers easy access to the LLM’s weights
not really. the weights are encrypted in-memory. through the use of TEE's.Needs to read up more on how LLMs work I think. Can't take the article seriously when the author seems to be making the claim that the weights of a provider model are loaded on the host machine, or implying something else just as incorrect.
They aren't talking about model providers. These are the tools you use with model weights locally (but you could set up remote infrastructure a la data center if you have the fundage).