Towards a harness that can do anything
eardatasci.github.io
eardatasci.github.io
Think of a typical loop we may ask of Claude Code today (assume we are not using TDD): run some test suite with fail fast mode, diagnose if the failure is due to recent feature changes (pass reference to backend/frontend, github issues, PRD,...). Ask CC to decide if test failed due to feature change and then update the test. Perhaps ask CC to use sub-agent to investigate and fix (if deemed so). Commit each fix, move on to next.
I know, this has so many ways to make blunder but I am talking about the agent here, not our error-prone test maintenance. What if we had an agent that had context of your codebase, deterministically ran test suite, linter, hooks, etc. The "English" prompt would become a code loop with the LLM only brought in to decide if a test has failed because of feature change. Also, we can extract git log, JIRA and what not.
Each tool here is real code. Executable code that calls others and only prompts when they meet edge cases. Edge cases are defined but we can now accelerate the maintenance of these tools using agents themselves. But the system is built on "programs that do one thing and do it well" and then reach out to an LLM for its specific edge case. The agent is how these executables work with each other.
There is this ACM blog post called "Manual Work is a Bug" [0] that was originally written to help humans automate processes using code. I find it just as applicable today as when it was written. You and the LLM look at what has to be done and then figure out the scripts/tools to make it happen. You then tie those tools into a system.
The more I use the above the more it makes sense and the worse the whole "just commit the prompt" seems like nonsense.
By trade I am a .Net software developer so as a lot of people would imagine — I was not able to accept a script that wouldn’t be reusable and flexible, basically over engineered.
I do quite some devops so I finally had to accept the fact that I can write simple script with hardcoded values that will live on a server (where I can copy paste and change values to meet other server) and most likely I will not have to look at that script for years as it will be running with cron doing its job without an issue.
Over engineered scripts designed from get go always required debugging from time to time so lots of time I was just doing stuff manually to make it quicker.
So I started winning when I accepted first script can be really simple and when needed I can move it to be parametrized but if not it will just keep doing it's job there on the server.
The upshot of this is you actually have a much better understanding of the different way your script needs to work if you're adapting it for a third use instance.
i dont think that really holds for a large amount, if not nearly all, of the use-cases for AI where it is either failing and shouldnt be in the loop at all or it is capable of developing some code to fix the problem permanently and its okay if that code is not perfect as long as it works.
refactoring with AI can always be a future use-case when the AI improves
What I am saying is the opposite - use Claude Code or whatever else - generate actual "programs". Basically scripts. We have tons of ways for "programs" to interact with each other. Then have clearly defined edge case handlers - think "try/catch". How far do you want to go down the rabbit hole in the "catch"? Do you want to re-write a new version of the "program" itself? I do not know, but this type of a system is what Unix already is, with the addition of programs themselves reaching out to LLMs in well defined edge case handlers.
https://www.langchain.com/blog/tuning-the-harness-not-the-mo...
The API is basically what you see as a user of Claude Code or Pi or whatever. You can make new sessions, send messages to sessions, configure which MCPs get started, etc.
I’ve been poking at something similar to what you’re talking about via that route. My client prompts the agent to do a thing, and then afterwards launches deterministic things to check it which can either re-prompt the original session or start a new session.
Eg it automatically runs the tests afterwards, and will send a new prompt in the original chat to fix them if they fail. I also briefly poked at a security analyzer that gets changed files via git and makes a new session to check whether there are security issues and propose a fix that then gets sent to the original session.
If you want a circular loop where the LLM can adjust its own workflow while keeping it deterministic, you can let the agent modify the ACP client that drives it.
File edits are just tool calls under the hood. If you’re using a decent agent then you should be able to override or extend the filesystem tools. If you’re on ACP, file reads/writes get proxied to your ACP client and you can inject your hooks there.
It’s pretty trivial to implement, this is well within the bounds of things most agents implement (for open source agents anyways, no idea how to extend Claude Code or Codex these days).
As coding agents have accelerated my work, I just build tons of tooling around existing software. Or in rare cases build new ones. If we zoom out of software engineering, we will still be in the realm of files - text or binary. That does not change.
The question is - do we let agents run the tools or the "programs" call the LLMs. The OS is the new agent, but not the same sense of "agent". I want LLMs to be lightly sprinkled in a future "agent" OS, not the other way around.
OP's idea "everything is a text file" is good and I use it too. My plans are saved as task.md files, numbered and named. Work items are checkboxes inside the file, closed work items are checked and a comment is added on the same line to provide feedback about the implementation.
I also keep a current-state-of-the-world document, it should be <20KB of text, keep the essential decisions and intents. Loading it allows resuming in <30s.
Something I never saw anyone else do - I save all user messages in a chat_log.md file which is referenced for intent alignment and state recovery. I consider the chat log on the one hand, and coded tests on the other hand as the two walls, the agent works in the mid section between them.
https://horiacristescu.github.io/claude-playbook-plugin/docs...
Gherkin style tests also come to mind
One of the meta-processes designed in is pushing automated processes, both defined and discovered, down as far as possible. "Down" here means as far towards the metal as reasonable. So automate the automatable stuff, and leave the LLMs to do stuff LLMs are actually good at.
A trivial example is 'handle this bugfix ticket'. Many actions in a bugfix are pre-defined, for example a git commit at the end of the ticket. So Maelstrom will, at the end of a bugfix workflow, will force a git commit from the LLM that did the implementation. The LLM never even sees the git command, it just fills in a JSON field with a commit summary, and the workflow handles the commit.
There are some inroads into this vision - but I haven't seen anything build directly for this (beside my own experiment).
I have some 'vibe noted' notes on this: https://zby.github.io/commonplace/notes/unified-calling-conv..., https://zby.github.io/commonplace/notes/rlm-tendril-and-llm-...
You may want to read earlier discussions https://news.ycombinator.com/item?id=48881112 And https://news.ycombinator.com/item?id=48051562
I
Thanks for dynamic workflow pointer. I don’t know if I like JavaScript for workflow definitions tho. IMHO sshwarts has the right idea on severely constrained workflow definition language. I plan to look closer into ADK2 workflows as well.
I’m not saying anyone did what you are doing, I’m saying multiple pieces are converging on that, at least in my imagination.
If you think about it, the transformers architecture was created to solve language translation. It works well for human language to code and other way around, already!
What we need is better tooling for this translation on either side. I started working on https://github.com/brainless/nocodo/blob/feature/praxis_agen... for this reason - how can we go from human language to code representing it.
I just did a complex (for me) task: I needed to wrap a 2015 build of Dosbox Daum, a 32 bit binary, in an AppImage. Claude kept finding incremental bugs, and I went through two cycles of depletion of my token rate with Claude. It kept getting close, but..... something was off each time.
So I took the Claude output and Chatgippity polished it off with a few more rounds. I then wondered how much Claude was "just showing enough" to try to hook me into subscribing.
That said, LLMs were quite useful, and I learned a lot about ELF binaries, and extracting dependencies. It's the ideal task: a breadth/obscure task that is documented but poorly explained, that I wouldn't have easily been able to do without LLMs.
Anyway, back to the article, do we really want arbitrary-billing silent tasks running? Like AWS billing spikes are bad enough to lose sleep over.
Also, if you want quiet rebellion against AI, developers should shove as much busywork on AI to overwhelm the AI budgets for your orgs, because it is very apparent to me that you can keep the LLMs doing lots of hardening, testing, redundacy, and optimization tasks with larger and larger and larger token windows and burn those tokens baby.
Once I actually have my plan/spec, it's the same process every time, it needs to be as deterministic as possible, using agents as tools throughout the process.
Yada yada yada introduction done so I can drop the link to what I'm building which is exactly that
I keep pushing back open sourcing it but it's truly close to ready and will be fully free to use.
It's the most advanced deterministic agentic orchestrator on the planet.
Right now we mostly YOLO prompts with some docs/skills in the mix but I think it will start to look more like internal MCPs, with tools the LLM can string together. I think the reality will be most tasks end up serviceable by Haiku-level LLM, not Fable.
It’s a DSL I’ve been working on to encode mixed deterministic/probabilisitic agent behavior.
Current tools. Opencode and whatever cli i can't avoid (like claude code for my first month which I don't thinking I'll renew) usually accessed using Paseo for it's excellen mobile client.
Why force the LLM to use files over vector database or key-value stores, just because it's a design principal for UNIX (which is designed for human users, not LLMs.)
> When in doubt, simplify. Remove, trim and minimize. Reproduce issues in as small cases as possible, understand the full design completely, there is no shortcuts for this.
Something I am convinced of though, there probably isn't a single `best` harness for all tasks. Different workloads will likely perform better with certain combinations of model + harness, especially when we are talking about token budgeting and cost tracking.
Ambiance feels like a great base “kernel” to build those variants on top of, rather than the one true harness.
LLM's are language models, you can absolutely control them with bash scripts and deterministic code, there are plenty of frameworks that already do that, and a great engineer will use them, but LLMS are at their most powerful when a user can give the agent an input and the model can run its ReAct loop. Wanting to free an LLM from a chat pane is like wanting to free email from the thread model, or closer, removing the chat window to DM a friend or colleague.
>what can we learn from the before-fore times, when people used to actually write code?
Treat an agent like a human writing code. Give them the best context, give them the best tools. This is why harnesses are overly complicated, because they need to guide the model through the context and tools it has available in a way that is efficient. A good harness is not incompatible with the Unix philosphy, it can do one thing well (interfacing with LLMs and giving them access to filesystems and compute), it will heavily use bash, stringing commands together with the cli tools that it knows (it's context) that it has, and and LLM will naturally handle text streams because that is what it does best.
>Everything is a File
If you want things to be deterministic why resort to plaintext? Wouldn't we want as much as possible to be typed? A computer can parse json which is what you want if you are trying to make your harness as deterministic as possible.
>It watches our FS for changes with cursors on textfiles,
Wow. What is your monthly token bill? I don't know how that would use less tokens than a 30 minute heartbeat, which as you mention will already use a lot of tokens. Why not have it notify your agent after a certain amount of files have been changed, or certain files you deem important?
It seems like this user works at a 12 week programmer retreat and seems to post their cohort's blog posts about the projects they work on.
Effective people managers (of whom I would not specifically consider myself) have known these tenets for as long as history. “Be concise”, “state your intent clearly”, funny how these are touted as novel “strategies” with which to expertly direct AI.
I don’t agree that “everything is a file”. Files are arrays of bytes. For an LLM, everything is a vector of tokens/embeddings.
An aphorism I recently heard: "All sufficiently advanced technology eventually becomes a web browser".
… seems apt especially in the context of the progression from chat-windows to harnesses and onwards to “harnesses that can do anything”.
And yet, as of now, LLMs have a hammer (Bash / command line utilities) and every problem they have looks like a nail.
If there are people around you who are non-techies that are using Claude Code or similar, you'll hear them ask "what the heck is cron?" and "why is it talking to me about Bash again?".
At some point we may have LLMs working on their own output using a set of tools that'd be the equivalent of Bash (or any other terminal prompt) + command line utilities manipulating not files but tokens/embedded vectors but as of now, it sure looks like everything is a file, especially to LLMs.
LLMs are good at traversing traditional file systems because file/tree-like structures very likely encompass an overwhelming supermajority of their training set. By contrast, other approaches like graph-based vector/data discovery via sql queries seem equally or more promising (to me), not to mention the ability to run curl/http queries. Either way, it’s still all prompt engineering approaches to discover/manage context. In this way, files seem like a “local minimum” ie a medium that dually optimizes interpretability and accessibility for both LLMs AND humans.
Just riffing here, but this could even broadly be considered a discussion of memory-vs-storage (somewhat analogous to fluid vs crystallized memory in humans). In this way, one could imagine the models performing context compaction/backup by periodically dumping their context to files (or some other non-volatile storage) WITHOUT coming back to feature space.. just dump/load the tokens directly.
Personally, I am much more interested in even other approaches like VLMs (using visual tokens), architectures like auto encoders and its variations, jepa architectures and other approaches that emphasize operating primarily within the latent space.
To be fair, my particular interests have always been more in the computer vision area, and have been amused to watch the attention mechanism rise to prominence even over CNNs (given the contrast |similarity in their mechanics). Then again, I began my journey in the ML field when GANs were still the hotness, but I digress……
for LLMs, it’s tokens all the way down, and the name of the game is how discoverable and accessible can you make them?
My harness is a Claude Code plugin with its own brainstorming, adr, and planning skills with associated review and interview skills. Behavioral testing related to acceptance criteria is built in. Everything in my harness is gated to prevent ratholes.
I recently inflated a docker container to execute a set of work with Claude in unsafe mode and immediately saw problems with everything it was doing…and then I realized I had not installed my harness.
Running Claude without an engineering harness is like driving a car without brakes or a steering wheel.
My own experience is that Claude will use your skills but will ignore your agents or custom search tools.
Current project: https://sxp.studio/apps/subjectivezero
I feel like Docker Compose / K8S / VM / Dagger.io layers are close but can't quite always recurse flexibly and aren't always simple to run with. Networking / devices / auth are often awkward choke points
It parses the LLM output for tool calls, executes the command, and puts the output back into the LLM input. That's all there is to it.
If it had a lossless, massive context window (100m-1b tokens), then it will squash everything. Give it bash + r/w and it can in theory /goal anything.
I think there's something to be gained in a production environment be siloing agents for reproducebility/auditability, but I suspect that will go away in the future.
There's that video of a silly demo someone made of an OS that was just nested copilot instances that generated the HTML of each window, which allowed you to do whatever you could imagine. It was seen as silly because it was, but that seems truly transformative.
It hit me at the end, when I read this:
> The idea behind Ambiance is simple: the model's priors [...] Everything else here is just in service of that.
I've noticed "in service of" take off similar to "load-bearing" with LLMs, and the whole structure just pattern matched to Claude for me.
I went back and scanned it over again, and noticed several other tells:
* "Think of U/L [Unix / Linux] as a motivating analogy rather than a direct comparison." * "Priors" _and_ "a priori" used in the same article. * "A real kernel […]. The Ambiance Kernel […]. The Kernel […]." — LLMs love this pattern.
To be abundantly clear, *I'm not calling this AI SLOP*; it's obvious that a human put a lot of thought into this, gives a shit about the topic, and shared interesting ideas leading to a productive discussion. I really did like it.
There's just something empty or hollow about the LLM style; paraphrasing Joni Mitchell "there's something lost and some thing gained" [1] when using AI. Your writing, your code is more consistent, more structured, more planned out — generally better but in a way that loses the character behind human writing.
[1]: From "Both Sides Now", a great song about looking something from two perspectives: youthful innocence, and jaded cynicism. Listen to the original 1969 version first, and then the 2000 remake as you can tell she's singing from the respective perspectives. Deeply meaningful song!
“Are you AI-ing me? Em dash giving you away”
It’s a pretty load-bearing (lol) example, yet it was clear in context it wasn’t serious. I’m beginning to notice people getting pretty good at detecting AI content, which I find reassuring.
It makes sense: AI-written prose is basically the homogenization of all styles of writing into a single voice, so it’ll stand out against the unique personalities we are used to seeing. It is just taking a minute for the average person to gain literacy — but it seems to be happening rather quickly, again unsurprising since we’re such social monkeys with a lot of our 15 watt brains dedicated to socialization and identity recognition.
Use your em-dash proudly!
isn't that a tall claim?
But
I don't want the AI to summarize my email for me.
I don't want the AI to summarize my calendar for me.
I don't want the AI to summarize my Zoom call for me.
Thanks