Docker Agent
github.com
github.com
I've been working on Pullboard and just open sourced it: https://github.com/pullboard-dev/pullboard
To me, orchestration is not really the core problem. The bigger challenge is handling agent coherency over long periods of time, (drift) so my solution is a living forum-like system where agents shout to one another, document work in items, and follow doctrine handed down by the developer and work against a known specification that's checkable. I have used this workflow in several large projects that MUST be correct, and it works for me. Currently cleaning it up for other users.
The vanilla out of the box product is actually quite close to old Trello, in simplicity. But there are so many ways to fuck it up and overcomplicate it, and so many companies proceed to fuck it up and overcomplicate it.
I have learned recently that I don't inherently hate JIRA, I hate our JIRA and byzantine process flows.
I think because other task trackers like Trello are simpler and have less config options (per my last use, long ago, pre-Atlassian-acquisition), it's impossible knock on wood to get into the antipatterns I've seen in those tools.
Ironically I think one of the main goals of management I have seen consistently over time - ensuring tickets are accurate, up to date, and reflect reality - would end up much more true if the task tracking tools were just much simpler, with fewer fancy features.
Need a specific feature? Plug that in later when you actually need it.
Chat is great for novel open-ended work, but is terrible at concurrency - i.e. managing > 1 agent runs. As agents get cheaper, I suspect we'll see more more bounded processes where you want to run lots of items through them at once, and thus Jira-like things... (we ended up having this problem and building something like this, albeit more medieval themed [0])
We also found the other piece that Jira-style solves (and that chat makes super hard) is multi-player. I've yet to see a good implementation of a multi-player chat solution to agents working concurrently; and we found that an async model is the only sane solution to this.
> gate: failing
Gave me a good chuckle. Cool idea though, and I like the records. Reminds me a bit of the OpenAI hacking incidents.
Thank you for coding something that I intended to code but have been too lazy to do.
Work is first defined in new or updated specifications, then the change gets made. You have to resist making natural language Todo lists and have the agent write runnable unit tests. My project is a CLI tool so it's fairly easy.
Verifying if it's done is a matter of running the specifications test suite.
In a totally not original way, I'm testing this workflow in my own agent orchestration tool. I know, there's just so many already. I'm building it for myself.
So far, the system is holding up but it's way too early to declare it a success.
You can check it out here:
https://github.com/egzo-ai/egzo/tree/main/specs
I think the same idea could be applied to other projects in different domains.
With this workflow I can read and update the specs, and run it against a built binary and be much more confident that the code works. Claude code has picked up the system without complaining for the most part.
We also open-sourced a similar system called https://github.com/madeinorbit/podium
We do have what you call Items and Shouts in the form of an agent communication system and a Linear style issue tracker that agents and humans use together.
In our case everything around orchestration is discussed with the agent itself and they take action for the user via our CLI. We found any UI to prescribe how teams of agents should coordinate too clunky. So in our case you just tell the agent: "every codex luna max implementer gets an opencode muse spark reviewer. when a subtree is done, review with opus high.". It then sets up the graph and the system enforces the rules.
It works really well for us.
How is "no code required" a selling point when we all have agents that can almost instantly puke out all the mostly-correct code you could want?
All the cool kids have one!
I say this as a gopher who also has a custom harness written in Go, it's a real challenge compared to the TS setups, but TS plugins are also a security concern
I haven't given this much thought (I also just started with golang last year) but I'm assuming the only real path here is to provide extension SDKs in Lua or other languages whose interpreters have been implemented in go.
I think that's what grafana's k6 [1] does with a JS interpreter
Go has the plugin package, but I haven't actually seen anyone use it (outside of toys and blog posts)
https://oneuptime.com/blog/post/2026-01-25-plugin-system-go-...
why not just use an SDK where you are using a plugin?
It may work for some things, but you really want to be wired into that system to do the more interesting things.
jk its trash
Vue and svelte are easy to use and learn/adopt. But look what won. React.js.
So whatever the ai agent with the vast user basae will win regardless.
---
I know it's a slippery slope, but given you brought up JS framwork analogy, I had to fall into it
Now Angular v1, man… if I ever see another digest loop error in my life, I will scream
Cursor is easily the worst experience I've had with an AI harness, yet it was acquired for billions of dollars in spite of being a middleman tied to a ripoff of VS Code. None of the colleagues of mine who were touting it last year are still using it. But if you can make yourself look like the next big thing, you'll get money thrown at you. Just look at Omarchy.
A good experience necessarily takes time and careful thought. Nobody these days can do that while also remaining relevant, or even surviving.
I kinda like cursor, mostly because it provides a nice review UI and I can use whatever model I want.
Getting away from all the Claude-speak has been a revelation, and fortunately Anthropic helped me here by hyping up GLM 5.3. Like, if there's an open weight Fable competitor, then why wouldn't I use it?
"Uh, OK, so how do I get this thing working? Where are the docs?"
"Well, first you install its unique framework..."
If, like me, you couldn't find any security-related info on the linked page.
If you don't want to use their harness then you wouldn't use docker agent and instead use their `sbx` cli to run the harness of your choice (claude, codex, pi, etc).
I think it will be great if Docker can get people used to using secure VMs. I am developing a similar project (still a work in progress): https://github.com/gregwebs/agent-vm
But currently uv is good enough to create isolated, repeatable environments. And it isn't limited to Linux like standard Docker. So when I want to test small samples on my local PC's GPU (Windows or mac) before running on beefier hardware, I can do that easily.
I tried not to do anything too crazy, but even use using an API key in the docker agent natively was getting me sessions that would hang and couldn't be recovered.
I work at another big tech company, and up until recently that would’ve been true for us as well (but it’s changed now, we partnered with OpenAI)
Sorry, what do you mean there?
You can drop any GitHub repo and the agent builds and deploys the app to your Tarvis workspace along with taking care of ssl domain and any other configuration and making sure the app is healthy.
looks like a weird abstractions. why do i need this if i have modal.com, e2b, cloudflare that use original docker + some toolings around + way to run + isolations on network level.
Having it docker branded, I could understand. It’s confusing but the docker brand is strong.
But exposing it as a docker subcommand is very confusing to me.
But it evolved differently and although it runs very well in docker containers and docker sandboxes, it also runs very well anywhere else.
It is called 'Docker' as the company, and it has nothing to do with the container technology.
Great tech. Not a business.
Does anyone remember all these web portals from 2000? What people mostly wanted was a search engine but they added everything else and lost focus.
Yep, open ai sol model wrote that.
There was a brief point in time at the beginning of the internet where if you saw an image online you could be kind of confident that it hadn't been manipulated significantly... But somewhere around the mid to late 2000s it became prudent to just assume any image you've seen online got at least a blemish filter pass through something like Photoshop.
Same thing now with written text. Sometimes you'll be a little bit more or less sure than an llm produced the output but never able to say 100% confidently.
That's a benified "not a drawback".
Something like a containerized sudo ACL for commands and for network credential , scoped by oauth scope
That way you could truly establish boundaries for an agent before it’s launched , and it could request additional permissions during execution .
At the end, the link with docker compose is just the yaml form. Which by the way can be replaced with hcl files or Go code.
just use and customise Pi.
Nothing else even comes close.
I'm not sure why anyone would want this. Is there something this actually does better than any other harness? Do current harnesses suck that bad at orchestrating things like Docker? If I do all my work inside a Linux VM and tell an agent inside of it what I need done, it figures out everything, including container orchestration.
The very job that AI gents are meant to do should mean that most of the documentation for this Docker Agent is obsolete/unnecessary.
Nowadays there are only two justifiable public languages you should be using for anything a) C++ b) Rust
Rust itself is dubious due to abysmal compile times and the fact that the only guarantees it gives you are of security (aka a skill issue). Modern LLMs write C++ that is as safe if not safer than Rust. Any other language choice is objectively wrong.
* For those inquisitive enough the keyword "public" is doing all the heavy lifting here. Pretty much everyone should be developing and using an in-house DSL at this point.