My current strategy is to not read any of the code written by my agents
twitter.com
twitter.com
I like optionals, but I see his point too.
This does not require a new language feature every time there is a bug, as he asserts. It requires language features to let you be descriptive in your type definitions so that invalid program states don't exist by definition. You literally cannot write one down. It's the same idea as saying you can't assign a Monkey to an int64. Scala's ZIO also shows that in fact you can type-infer whether a given path will produce errors or nulls, so you don't need to have perfect knowledge up front or go back and change tons of code if that changes. Errors can automatically propagate, and you need to handle them once, somewhere. Checked exceptions were a fantastic idea; you just need to let the compiler infer them everywhere.
Or you just rely on the type checker to say you're not allowed to pass a Book to a variable of type Account, and you never use casting (incidentally, I don't remember where, but I remember Clean Code having some example of "good" code that relies on casts, which shows the mindset). Then you literally can't even write a test case for this, because it's impossible to write the illogic at all.
If the compiler produces wrong code, all bets are off. Your tests can also miscompile. Your unsafe code could modify another thread's stack memory between instructions and act as an evil gremlin so that literally no line is trustworthy, so even your null check is pointless. Or you could... not do that, and treat your programs as logical reasoning.
Charles Babbage even addressed this:
> On two occasions I have been asked, 'Pray, Mr. Babbage, if you put into the machine wrong figures, will the right answers come out?' I am not able rightly to apprehend the kind of confusion of ideas that could provoke such a question.
But at the end of the day, he's a guy with a long career in programming, which makes him significantly better than the median, but it doesn't make him three sigma above.
In this case, he's got a pretty good short- to medium-term argument that if you have a robust testing and verification suite, AI code that passes it all is a terrific outcome.
I'm really not convinced about long-term. Possibly, AI ends up writing even better in the future and we never have to worry about human maintainability.
But it's also possible AI is at/near its limit, and we'll always need human maintainers. In which case, incomprehensible vibe coding might be a problem.
The really interesting thing to me is that formal verification wouldn't have helped there either -- it would have just written a mathematical proof of the correctness of the backwards feature.
The point was that if your bullet hits someone, you would report it to the server, and then the server would check if you got the hit or not. (So you couldn't just send fake hit reports.) So the game would still be fair and CPU usage would be greatly reduced.
The way the agent implemented it was the other way around. It had players send reports when they had been hit. So they could simply not send them and make themselves invincible. Also the game still checked collisions anyway making the whole feature pointless (on top of broken).
(Actually it made the game slower than it was originally!)
Hope that clears things up.
--
An AI agent actually pointed out the absurdity of the new code to me. So I'm assuming it wasn't that one which wrote it. Probably one of the smaller ones I was testing. But unfortunately I didn't get them to sign their names so I'll never know!
I haven’t seen it do that in quite a while, but it was an interesting failure mode!
Also, what were your mechanisms for reviewing the plan against the described business outcome? Anything else you could have done there to catch the backwardness?
And finally, did you run any of it across models from different foundation labs? I'll often run important stuff generated by codex across Anthropic, grok, and Gemini with an opus or fable judge...
In this particular instance I'm not sure what happened. I don't know where it got the idea from to do it backwards.
My best guess is that the to-do items were too vague. I built up a lot of context in my head from back and forth with the agents, and so my mental model was pretty solid, but probably an insufficient amount of that was encoded in the repo itself.
I'm not sure what you mean by reviewing the plan — do you mean making a detailed implementation plan before beginning the work?
Most of the changes were pretty straightforward, or at least we'd already worked out most of the details and put them into to-dos. So for the most part my prompt was just "alright go ahead and implement the next thing on the list."
If I had to guess I'd say the main reason it went wrong was because the why was missing. The to-do item specified what work remained to be done, but it did not explain the reason for each item. I think that was the core of the issue.
So it probably ended up seeing a bunch of individual changes out of context and not understanding what the point was supposed to be.
---
>did you run any of it across models from different foundation labs?
Working on a different part of the same repo a week later, I used Fable and Sol to find issues. Then I let them run cross critique on each other's reports. Then I had each of them generate a plan, and then I had them do cross critique on the plans. And then I repeated that until they were both satisfied with the results, i.e. until the two plans merged into one coherent plan.
It was interesting because they're both had different strengths and weaknesses in different parts of the process. (One of them sound more issues in the initial phase, but the other came up with a more comprehensive fix for each one.)
I've seen very promising results for model alloys with security research (it was posted here about a year ago[0]), so I wanted to give it a try myself.
There doesn't seem to be a very convenient way to do it (maybe one of the new harnesses which runs the proprietary harnesses as subprocesses?), so I just passed markdown files between the two models manually.
LLM code past a few pages is really unreadable. A workflow of LLM writes code, human reads/checks/corrects is IME very unproductive and frustrating, and you're usually better off writing by hand (YMMV).
If you want to reap some productivity improvements, it can't rely on you reading, grokkong and accepting the LLM code line by line. You just kind of need to suck it in.
So the best you can do is to indeed, constrain it heavily, maybe write the interface, then have the LLM write tests, then have it write the code.
I'm not saying I like the approach, but I agree manually verifying LLM code is very frustrating and unproductive.
I find it plausible that an extra agentic review pass and more testing can bring this number up to the point that one never needs to review code again. AI writes pretty good code nowadays.
(You still need to be diligent and decide the architecture during the planning, and read the gotchas and "things to note" that the agent will spit out at the end of implementation if it had to diverge from the plan.)
Giving up reviewing code is not strategy, but because:
Review AI code line by line is like watch movies frame by frame, and is impossible, very difficult, terribly boring, or abandoned sooner or later.
I think a lot of the "AI Addicts" are masking illiteracy, which is why they cling so passionately to the technology.
"What kind of software do you have agents writing without human supervision?"
Which, is fine, but not what most people are hoping for with these tools.
Do you think that we're going to use less RAM with AI produced code?
and suddenly things work again, it's magic!
How shallow does one need be to feel compelled to identify with this?
this is definitely a better outcome since all the human effort now shifts to spec, design, test scenarios tailored to the domain - as it should be.
having seen my share of human slop through decades (mine included), finally it's a relief to not have to deal with devs that don't have high standards, and the endless arguments/politics that ensue.
with llm it's just a text file away from compliance.
Any opinion that starts with such a blatant appeal to authority can safely be ignored.
> Started programming in 1983. Old?
I took that sentence as asking if he was old because he couldn’t trust AI and uncle bob brought up that he was much older and trusted the AI output because he trusted his constraints and test harnesses
Robert Martin started coding in the late 60s and then became an author and a software design consultant - his statement defines the start date of his relevant experience but doesn't speak to the end date of that experience or the density of his experience. He has a wealth of experience but not in a field that's relevant to his comment.
I think an appeal to authority is a good basis for extending a bit of extra trust to statements - but you should always verify things yourself.
When you start getting good results from agents you soon realize you are the bottleneck.
Automating the verification of the code is the way to go, otherwise it's just not worth it. It takes longer to read and understand code than to write code, so why bother with agents if you are going to manually review it all anyway?
One thing I miss after ditching Copilot (it got too expensive) was how I could trivially ask for features to be written by one model and verified by another. Have Opus write it and GPT or Gemini verify it.
I figured they were entirely separate models and therefore unlikely to hallucinate in the same way, so it gave me a quick sense of confidence.
Currently I use claude code (different models but all variants of the same) so I have them do planning, review of the plan, implementation, review of the implementation, and unit tests. It's fine, but copilot felt easier.
If you want even greater fun, launch claude and codex in the same working tree and make them fight it out in real time.