Automating AI Away
replicated.live
replicated.live
For example, with browser automation, giving the LLM raw access to the literal DOM generally results in disaster for tasks that need to be stable across more than 5-10 interactions. The better approach is to write an intermediate layer that understands each view and can provide a list of tools that are precisely tailored for each case. E.g.:
https://myapp/Login
- <raw dom - hundreds of kb>
- Available Tools: <arbitrary javascript>
vs https://myapp/login
- We detected that this is the application's login page.
- It has the following visible elements:
+ Username
+ Password
+ Login Button
- Available Tools:
+ PerformLogin
+ Quit
The later case takes a lot more effort, but it also reduces a Turing complete problem space into a binary decision at this particular step.Like are you prompting like:
--- I need code that does X,Y, and Z. Write it so that the Roslyn compiler on this machine can compile and the code passes the repo's styling and formatting requirements. ---
Or something else.
https://learn.microsoft.com/en-us/dotnet/csharp/roslyn-sdk/t...
So it would be something like:
Rewrite this Python code to use match/case instead of if/elif/else chains, write a script using the ast module to rewrite the code, do not edit it yourself, also write some tests with clear inputs and outputs I can inspect.
Or something.
Is this a real example of something people use AI to do? If so, I don't understand why that's difficult, because prompting the AI to do stuff with ASTs etc. seems a bit over the top.
Are people actually using AI to do programmatic refactors over million line codebases directly? That is far more insane.
Using AI to write ruff rules or clang-tidy rules with fixes is literally the same thing and obviously best practice over running AI in a pre-commit hook to do those checks and refactors...
BTW, you should probably fix the Beagle link on your homepage: https://replicated.live/beagle/
[1]: https://github.com/gritzko/jab
In other words, are there places where a one liner for the agent would be more reliable than markdown instructions and crossing fingers?
I look at it this way... I wrote scripts over the years to make my life easier. Do the same for your agents and free their attention for the parts that matter.
Else if the agent creates a new script evry time the non-determinism rears its ugly head again.
But that means we need to know that the existing script is exactly what we want. And that means we need to understand the code, or at least the tested spec of the code, that AI writes for us. AI can't replace humans, humans must remain in control and understand what the code writent by AI exactly does.
I also aspire to make one post a day. To be continued.
This is well-observed.
> validating that the LLM didn't disable tests it didn't agree with
Provide a test runner and force the agent to call it. Have it emit something if you want evidence.
I also know that if the meta-test is writable to the agent, it will change it if it feels like it wants to get rid of some other test. Even if it can't change the meta-test, it can hollow out existing tests to make them pass trivially.
I don't think nondeterminism is the problem. The problem is following rules: If I tell the agent not to change tests, it can conveniently "forget" about this. It doesn't much matter if it forgets deterministically. The problem is that it can forget at all.
The only way I've found to really force rules is via hooks, and even then I think it's just prompt injection? Maybe some kind of hook/checksum thing to ensure you're running unchanged tests, but as you're pointing out, the agents can get sneaky and do weird stuff if they have write ability.
Formal methods, as in proof of correctness, have been around for decades (I was doing that stuff in the 1980s) but pushing the proofs through was too laborious. The seL4 verification effort reportedly used over a decade of people time.
The idea is that if you have a formal specification of what you want to happen, you can get a LLM to do the struggling with the proof system to get it right. It's a good task for an LLM, because there's feedback from the prover.
I'd like to see more non-trivial examples of this. People keep republishing verifications of greatest common divisor or stack algorithms, which was done decades ago.
Coming up with simple specs is not necessarily easy. You could say that is kind of what math is about. That’s how we actually make progress: find those cases where simple specs are possible and build upon them. That’s the kind of library made for eternity.
Very often, the spec is indeed just a very simple implementation. Often you can make the spec especially simple if there are no constraints on the resources it can use, at times even infinite ones.
LLMS should be abstracted out of a process as soon as practicable, replaced with deterministic processes or procedures. Otherwise you’ve built the world’s most fragile process at the mercy of token cost, vendor hostility, geopolitics, and model deprecation.
I've genuinely never considered it from this angle before.
Thus why we replaced computers (flesh and blood people writing out calculations) with computers (silicon-based number-crunching machines).
It doesn't matter what we are, what matters is what we want, and whether what we built actually works the way we want it to work.
[0] Discworld's Ponder Stibbons would be rolling in his, grave, or more likely his "Early Death package" pocket-dimension jar.
This sounds made up or your workplace is rather odd to say the least. Maybe english isn't your first language and "threatened" is not the correct word?
And then... yeah. You got it exactly right. Once a problem or process is deterministic, that's the wrong application of an LLM.
But I had never quite thought of it in these exact terms. The way I've been thinking about it up until now is that the very best way to use LLMs is to have them produce tools. The tools get to stay reliable and predictable. They boost your performance. But I think you found the more general abstraction of the same idea. Tool-making is not deterministic. But the tools themselves can be. That's why it fits. Trying to stuff LLMs into what's otherwise a deterministic process is an absurd waste and error-prone.
Smart. I like it.
Just to be clear, software development itself is not deterministic, though? The software developer pushes a given business process from less-deterministic toward more deterministic? When we say we’ve “abstracted LLMs out of a process” we’d also say that we’ve abstracted software developers out that process as well?
Would love to know how you’ve managed to counter this as the drive to throw everything at LLMs is driving me insane.
There is a great article called "Manual Work is a Bug" [0]. The idea is that you have humans doing a lot of random things so you should:
- first make a list of the things they are doing
- then update the list with the commands they have to run for each step
- some of the steps won't have commands b/c it's things like "ask Bob what the limit should be"
- over time, the commands become scripts
- then the "ask Bob" becomes an API call
- one day, the whole thing is an automated system that runs code
People like to think that LLMs can do all of the above. I don't get this b/c code is deterministic and can be run repeatedly basically "for free" (at least compared to token spend).
I do think that LLMs can greatly accelerate the creation of the code/system etc and can also help with maintaining it but the whole "we will just version control the prompt" was clearly hogwash.
I asked Claude to spin up a bunch of agents to do it and after a bit of discussion we ended up writing a bunch of deterministic scripts that ran off the data collated by some “research” agents.
It took a few pilot loops of the process to nail it down, but separating the process into “data collection” and “process the data” has pretty much eliminated the AI step. Once the data has been collected from the random sources and normalised into something sensible we rarely have to do it again.
Even that process has been largely automated, scripts that deterministically scrape data, the AI is only needed for the very difficult parts that need some decisions or interpretation.
Getting the AI to write tools for itself is a great way to work.
Developers and developer-adjacent, technical people tend to think this way on their own... but every business has dark corners where repetitive, manual things still happen. We're leaning a lot on training and even org-wide LLM instructions to try and let the LLM (by its own assessment) be the vehicle use to codify a process and turn it into some good old-fashioned reviewable, deterministic automation.
Claude or any other model just translates your natural language instructions into formally defined tool calls. You cannot replace this layer with a formal tool like Ragel. You can write code for Ragel directly, in which case the responsibility for this is yours and not Claude's. (duh)
>What about Claude? Well, my instructions say in all caps: DO NOT PARSE ANYTHING MANUALLY, EVER. (...) It tries anyway
This needs a self-verification loop. It still won't guarantee that model's interpretation will match yours, but it will improve the accuracy. Almost every model will know that it went off the rails upon checking what it's trying to do. Harness has to provide the loopback for this, because the transformer architecture doesn't.