The most obvious example here is `read` or `view_image`. If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Does your prototype overcome the limitations mentioned in the article?
> If a multimodal model needs to read an image, it cannot use cat for that because the harness needs to inject the actual image payload into the protocol of the LLM.
Sorry, I'm just not familiar, but it sounds like there's a protocol the model expects when receiving images that bash does not support?
(This is Bonteq, I was just logged into the wrong account.)
> While in theory the agent could provide a CLI tool that talks to the outer harness via environment variables and Unix sockets, it’s a rather crude process
But I don't know why it's crude to be honest. I'm running pi in tmux and have a CLI to prompt it from any shell session / neovim and it works good. So such way of communication is already needed besides codemode.
JS was chosen probably because (1) it's easy to sandbox (there's QuickJS) and (2) (my guess) some models are probably post-trained on JS Codemode.
It becomes much crummier when hands and brain are on different machines.
It is not from my perspective because Codemode runs in the brain, and bash necessarily runs where the hands are. So if the hands need to reach into the brain, I need to set up a communication layer from the hands to the brain.
Because the models are trained on JavaScript for code mode. You get away with way fewer instructions. They also want to be able to express concurrency and that works very well with the Promise global.
But a big reason is that code mode runs on the harness side so bash is a tricky target in particular.
1. bash can also express concurrency easily, with sync and async (using standard & syntax) support for each command, and standard cancellation (kill, though crude). 2. "way fewer instructions": It requires no instruction for agents to use bash either, except for merely listing the custom commands (view_image, apply_patch, etc.). Also, bash has standard progressive disclosure mechanism (--help) that models will automatically use with no instruction.
In a world where brain and hand are on different machines, getting the bash hands to reach back into the harness brain is something that requires a) putting tools in its hands that it does not know about b) are tricky to set up, usually involving some sort of socket based back channel.
I tried this quite a bit, by having pi be always there on the hands side, but it causes a lot of complexity and the LLMs really do not understand it well at all.
Infact harder to sandbox bash (just-bash or brush based) than it is to js or lua, which has fantastic embedded tooling.
Also there are things like subagents, etc (which may be considered tools).