Why are coding agents so dumb?
mtlynch.io
mtlynch.io
> Agents can’t manage tasks
Say "use subagents".
> Agents can’t delegate
Say "use subagents that are Haiku/Sonnet"
> Agents have never heard of agents
This isn't even a problem: the author is just confused why Anthropic didn't bake all of Claude Code documentation into the harness, but of course it makes no sense to pollute context like that.
> Agents suck at communicating plans
True, communication could be better.
> Agents take any excuse to stop working
Just use /loop.
> Agents are only useful when they take unnecessary risks
Just use a sandbox. OK, technically, this one isn't literally solved by Anthropic/OpenAI, but it is solved by a million agent sandbox startups.
> True, communication could be better.
Nitpick, but bit of a contradiction.
And of course, if the agent has access to the network (which is probably required in order to talk to the AI server), then you have to think about what other things on your local network it could get into.
Lately I've been asking a lot of these sorts of questions while setting up a sandboxed environment for my agents (generally, claude and pi). I was quite surprised to learn how many different ways there are to configure git run a script out of the .git directory. And the various AIs know all about this.
At this point my agents get read-only access to my .git directories.
Also, "esoteric abandonware" is a great way to put it. I keep looking for solutions to problems, finding dead projects that haven't been committed to in months, they have all the indicators of AI slopification (the primary one for me being emoji-filled READMEs). I don't think these types of projects will be around for the long-haul. These projects are just as much about community and people as they are about code. And people are much less likely to care about a short-lived project that doesn't have the backing of people that intuitively understand the internals of it. It's like building on sand.
Personally, I've decided that allowing the "agents" (these buzzwords are quickly becoming pet peeves of mine) to write the code for me is a bad idea, because the faster the LLM constructs architecture than I can, the less I'm absorbing the details of it, and the less likely I am to intuit obvious shortcomings. They do help me work much faster than before though, because I still use them like a research/learning assistant chatbot. Learning what I'm doing wrong faster, while still being the one putting the pieces together, is very great.
* Hold on to that isolating feeling. Those of us who feel AI is an incredible accelerator also feel just as isolated as you do. We wonder why you aren't seeing what we are seeing - and that isn't criticism, we just genuinely don't understand where the gap is and what would help other people jump across.
* I wouldn't bother looking at GitHub, or using that as any kind of measure. Personally getting away from all dependencies is my goal, and avoids a lot of supply chain issues.
* I don't use any plugins or external skills. I do use a ton of self built local MCP tools though, and it drives much of my workflow. That's me though.
* You sometimes need to work up to it. I've had periods of hitting a wall this year with LLMs. Once I end up building a framework, and the framework is in place, everything flows. The models can learn from looking at your adjacent projects, and they just fly. That's for coding, but it's similar for solving other problems... record everything, give it a ton of context, etc.
* If the software "kinda works kinda doesn't", then someone isn't doing enough testing. You should be dogfooding your software every day. Use your AIs to help you find edge cases, sure, but they're alien minds. Your human mind will sometimes think in ways they don't, but that your customers will. Sometimes we write software for AI to use, sure, but if you're writing for humans, you still should add some human mind thinking just to double check the model's work. "That's right, it goes in the square hole."
It's okay to struggle, some of us have been building toolkits over the last two years that make AI work so much easier. But don't get sucked into thinking "it doesn't work", especially in the Opus 5.5 era. On my side of the fence, these last two weeks have been the most intense and productive in a long time.
Lastly: some of us have tips we will not share. Some is slight selfishness - if others haven't discovered what the models can do, that is a small moat. But there are also some tips you just have to find on your own. I could tell you playing The Witness is some of the best training you could have for working with AI, and many will think that makes no sense at all. But I think a handful will nod and quietly understand, the lesson can't be taught, it must be learned.
Sorry for the long ramble, but I genuinely hope some of that helps.
What I was thinking of here - if you remember the statue near the mountain, and it reaches out. There's also other statues in the area. But somewhere around that area, there are also trees. And if you look at the trees from the right place... the statues, the trees, they're doing the same thing. Why are trees doing the same thing?! Pondering that led in interesting directions.
The Witness was so multilayered, in ways that unexpectedly rewired my mind. My suggestion is that the rewiring that can occur from The Witness, is also a useful mental shift when interacting with AI. I think many of the people I see going fastest and most fluidly with AI have had that mental shift.
On a different note, where I can be more direct. Prompting is not just "the prompt". The entire environment is the prompt. Everything about the folder the AI finds itself in, adjacent files it's allowed to look at, everything it sees and encounters as it progresses, they each steer the model slightly. So where there used to be "prompt engineering" and "context engineering", there's also "environment engineering". The environment is the context.
And that perspective may bring you back to The Witness again.
I'm not blaming you, but it's incredible how much open source garbage there is and people spend years making actually good stuff that actually works and nobody's publishing it. So to me, and I'm trying, but to me your point is hearsay at best. You may as well have told me it took 2 seconds and you've always had a perfect time with it. It's all meaningless to me.
It still leverages agents (people are used to them already) but can deal with executing it on a different vm, dealing with git worktrees, etc (not that coding agents themselves cannot be improved, but I think the sw development infrastructure around it has a lot to improve too).
When you asked “how do I use Claude code to…”, you invoked the system prompt’s Academy Skill [1], pretty much verbatim.
“When a user asks a question about Claude, a Claude product, or a general "how do I use AI for X" question, check the Academy catalog (see "The catalog" below) for a strong match”
The catalog includes a Claude code documentation hub, and the thinking traces in your screenshot show those docs being fetched. So. Working as intended I guess.
Also, if you turn off web access or simply ask it to answer from internal knowledge about its “own freaking features”, as you put it, you’ll find that Claude does indeed have such knowledge.
[1] https://github.com/anthropics/skills/blob/main/skills/academ...
PS There’s a master .json page of that catalogue that’s actually kind of cool to skim through (ugly as json is).
I started to gloss over hard after a few paragraphs because I don't really have any the problems you describe anymor; at least not to the point i'd be blogging about them. Instead, I've just been iterating with an agent on various pi extensions that solve the issues.
As with operating systems - you can hold out in the hope somebody solves the right mix of issues in general, and they match your situation well-enough.
That's your choice.
I would note though that being the passive consumer gets you either Mac or Windows UX & prices.
Isn't this the point of Pi to implement all the features you need yourself? Complaining about lack of features in Pi doesn't make sense to me, it's their whole identity.
Although I must say that describing specific problems is valuable on its own. It's just (a) why stop there and (b) if stop there, why shape it as a complaint and not as a list of features that are worth discussing.
I'm confused how you see this article as a complaint about Pi. I only mention Pi at the very end to say it's one of the agents I've tried. I checked it out, found it too barebones for me and moved on, but I understand why some people prefer it. The existence of Pi isn't a counterargument to, "Why don't agents support development workflows that should be commonplace?"
But I'm still kind of puzzled by this critique. Isn't it kind of like saying, "Why are you complaining about the agent? The C programming language exists, so you can use that to create any agent you want."
If I wrote a blog post complaining that "Programming languages don't let me do memory management", would you not think it's fair to respond with "What are you talking about? Have you heard of C?"
My take is that agents don't do the expected thing out of the box. I acknowledge that you can extend some agents and customize them, but I'm saying that it's sort of silly that the features I think most people want don't exist as default, out of the box standards.
> If I wrote a blog post complaining that "Programming languages don't let me do memory management", would you not think it's fair to respond with "What are you talking about? Have you heard of C?"
I don't think that's a fair comparison because C does let you do memory management out of the box. I think the better comparison is like, "Go doesn't let you do pattern matching," and the response is like, "Why don't you just patch Go to let you do pattern matching and maintain your own fork of Go?"
I still occasionally have issues with open-weight models, but the frontier labs have solved the above for most use cases.
It works pretty well with Opus/Sol or better types of models. But with Sonnet/Terra level, it chases its tail and is pretty crappy.
I'm just waiting until Chinese models get to the required level, its going to be great for the cost.
A lot can be improved, but this is already so much speed.
This is a pretty big failing, which is compounded by the fact that most humans don't know which model to pick, or make assumptions based on Anthropic's hierarchy or "effort" involved.
Like Fable: your toughest challenges. You mean, like Fields Medal toughest challenges? Or analyzing and updating three monster spreadsheet toughest challenges? Or writing a new novel in the style of William Gibson toughest challenges?
It's often nearly impossible to know in advance how hard the task will be, so the model has to guess, and will inevitably miss. Just because you know how hard the task is, doesn't mean it's obvious. Often it's only obvious to you, because you have additional context that models are missing.
A bad guess can lead to either context rot (if subagent wasn't used), or slower execution (if subagent was used but it turned out to be a bad idea). Both are bad UX.
Now, we could improve accuracy by having the harness do a research on each task before executing... which leads to even slower execution, also bad UX.
The descriptions are near-useless and tend to flip around, as model families are not released in sync anymore, that's true, but fortunately, thanks in a big way to subscription pricing, the choice is simple: start with the best model on offer, and when you run out of quota, downgrade to the next best (or briefly switch providers).
I didn't want to complain about OpenCode too much because it's open-source and kind of the scrappy underdog to Claude/Codex, but I do find its multitasking support pretty limited. I have to specifically remind it to spin up subagents, and when it does, it'll do 2-3 tasks in parallel and then wait until all of them finish. It can't seem to inspect tasks in progress, and if a subagent dies or hangs, the main agent can't seem to retrieve any information from it, even though I as the human can look at the session and see what happened.
Are you doing anything special to make OpenCode multitask well?
I'm confused about opencode's v1/v2 split. The docs say v2 is here, but their Github is only v1.x.[0] And the docs don't really explain why v2 is better, but there's a long list of things I have to do to migrate,[1] so I haven't been that eager to pull the trigger, especially since I can't tell if they consider it alpha/beta/release.
https://github.com/anomalyco/opencode/tags
If you use Nix it's a flake so super easy to run too.
Seconded that v2 has much better background shell + subagent management: it's now much more eager to appropriately spawn either of them on its own, and does so in a way that is asynchronous, i.e it doesn't block the conversation and you can examine them by pressing down arrow at the prompt box.
I let most agents work asynchronously and don't pay attention so I don't care that much about the sequential nature. But if it's a problem for you then fix your harness. This is a bit like saying "Why are shoes so shit? There's a stone in one and it just gets stuck there and your foot steps on it and it hurts". Take off the shoe, and shake out the rock. Put the shoe back on. You have the power.
They even have split my decisions to human decisions. Proposed and approved work. They can iterate on approved work without me just fine.
And I only just started with agentic coding in last few weeks before that I was mostly a copy paste chat person.
Edit: I have access to codex, vscode, GitHub co pilot cli and all anthropic and openai models (excluding mythos).
But maybe that’s on us, AI doesn’t care about all these special cases, it’s not debt to it as it will simply read them all when making changes. We’re obsessed with quality and what code is supposed to look like but those are human standards, AIs evolve to look at this complexity as a single picture, they can simply see through it so what is spaghetti code to us is merely some code to them that works as it should and is efficient. It’s interesting we can see how the two things drift apart, you would think at some point AI generated code should explode but it hold together unreasonably well in most cases…
Each time you encounter a shitty thing you hate, add a new 'review type' / 'thing to watch out for' and just ask your agent to add it to your hooks for you. This works well with Claude at least.
I have about a dozen or so hooks that run on every integration branch my agents write that review for all sorts of things from correctness to spec, performance improvement opportunities, modularity, analysis of any dependencies added, 'definition of done', UI/UX, etc.
I recently told Claude it should run the whole suite of reviews twice. I will probably go on and proceed to having it run like 5 times eventually idfk.
But the more you start asking your agents to modify their own behavior, using the native solutions offered by Cursor, or Claude, or Codex, the sooner you'll start to feel better about the results.
In your workflow, who implements the review feedback - the review subagent or the code-writing-subagent?
Do have a baseline styleguide (like Google's Go style guide) for the review subagents, or is it entirely the subjective things and specific corrections? I remember 6 months ago it seemed like piling general "good taste" code advice into AGENTS.md was considered bad.
Do you move between harnesses or have you gone all in on claude? I've bounced between claude/codex/omp, maybe to my detriment.
The biggest things I've struggled with are models having taste. For spec writing I was having a lot of issues with them making statements that were interpretable in a superposition of ways, eg "we'll do XYZ with entities that support and need it" when there's 3 possible entities and the model hand-waved at exactly the wrong tokens.
I added AGENTS.md guidance + memories to be unambiguous (with short but good examples) and "no coined shorthand". Now I'm getting a marked increase in specificity, but it's places that don't matter (claude explaining existing code to itself). I'm having difficulty controlling the spew of new text, but feature writing/research still gets fuzzy and lazy around the difficult underspecified aspects of the problem/feature.
Do I just keep dumping examples into reviewer subagent context and have them rewrite and simplify the research subagent spew?
I repeatedly have "Risks and gotchas" sections have a whole paragraph dedicated to things that don't matter, and then a single sentence bullet point that's actually a huge problem when I dig into it.
Do I have a generic hook to make subagents to review bullet points and size them commensurate with impact? If I tell them to "have taste", do they have taste?
I'm just trying to get good code done. I hate these things.
Then have a separate review agent review the work against your codified working philosophy. THEN, give it a review yourself(just don't tell anyone you look at code or you'll get called a boomer).
You said you've tried this but that it hasn't helped massively. I have found this to help massively just not without fault.
The struggle is real though.
You're not doing anything wrong. The cold hard reality is that this tech doesn't work half as well as its supporters claim it does. You've seen it with your own eyes, as have I. LLMs are not good at programming.
Your shoe analogy also breaks down because really the shoe is the issue, not the stone. And expecting everybody to make their own shoes is, well, I mean we just don't do it that way anymore for good reason. Let the cobblers make the shoes, and the runners wear them.
Codex Astra can do a great-(ish) job as a project coordinator dispatching tasks to a pool of 6.1 Sol sub agents. You can even give it an explicit goal and ownership over ensuring the work is carried out efficiently.
However the OOTB harness(and prompt) configuration may not do this for you. You'll have to provide guidance over how you want it to operate through your prompt, a skill, or etc.
And I'll say even though it's really good at this.. Having even more layers than 2 can help; a single agent given too many responsibilities will start to become fixated on a number of them while neglected others. You can check in occasionally to "nudge" it or you might need to split out responsibilities more..
I will say it's crazy Codex doesn't have more built-in task and sub agent management features. I almost wish that it had some stock orchestration patterns that worked OOTB, and then you could opt-in to a leaner setup where you provide more of the instruction.
> You can do multiple things in parallel and context switch millions of times faster than humans. Why are you doing these embarrassingly parallel tasks one at a time?
Tasks are "add e2e coverage"and "run final verification"
Context switch into embarrassing yourself running final verification in parallel with adding pasphrase type
> what coding agents should be able to do out of the box My dream agent
Yes, keep dreaming!
One thing that seems to help for me is to do the docs before plans (collaboratively edit with agent). Then I understand what this change is going to look like from the user's perspective before we start implementation. This seems to help keep things on track.
While I don't use this plugin a lot anymore, I think doc-driven development is one of the most effective ways to do development in any paradigm, I should probably refresh this plugin and use it more:
https://github.com/tmpdir-org/tmpdir-claude-code-marketplace...
I run an LLM server with Qwen 3.6 in the office, and OpenCode, which the OP mentioned, usually defaults to sequential TODO lists, and it works fine with our little LLM server with 3-4 parallel users. But I noticed that once in a while the LLM got overloaded with requests in the queue, and you couldn't do anything for 20-30 minutes. My investigation led me to an employee who used QwenCode. I tried it myself then, and indeed, it immediately launched something like 6 parallel subagents, where OpenCode would have sequential TODOs with the same model by default.
So in the end, I had to detect QwenCode on the server side and serialize all its parallel requests into a single request queue, because it made life miserable for other OpenCode users :)
Have you tried telling models about your dream agent environment?
They can build it.
Looks like this is a layer on top of OpenCode. Seems interesting. I'll check it out. Thanks for the tip!
The dev is very responsive on Discord and I'm sure he'd be happy to hear your thoughts and suggestions!
Makes me wonder about doing some archaeology and trying out really old harnesses on modern models...
I don't expect the model to be trained on the latest features of the harness, but I think the harness should ship with its own docs so that the model can quickly search local docs rather than search online page by page for docs that don't necessarily match the local harness.
As expected. A model's knowledge is what was it ingested a creation.t
Unfortunately what we get is worse - for the same reason. Model version thinks it is its previous version.
This is a great list for future Agent / Harness software engineers to read!
Oh sure, there may be some Agents/Harnesses that already accomplish some of these things -- but there doesn't seem to be one (as of the present day that I write this) that accomplish all of them...
As someone that watches the Agent/AI Harness (and related software) space, I will definitely be referring back to, and re-reading this list in the future!
An excellent post!
I asked DeepSeek to translate a page to five languages and it opened five subagents each one working independently on the translation, once they all finished the main agent informed me of the job completion with a bell. Fantastic!
Sooo, which agent?
I have no idea what stops that person from just making it, with an agent of course.