We don't need to tell the LLM where the User model is. The LLM can find it.
However, IME we need to tell the LLM preferences like, "the User model should not have dependencies on other models" (you know, to avoid the God Model problem...)
To discover that preference, the LLM would (best case scenario) need to look at every model, and notice that the User model has no dependencies, but all the other models do, and infer that this is an architecture requirement. Except, some of our other leaf models might not have dependencies anyway, so it wouldn't necessarily be apparent. Or maybe it's a new project and none of the models have dependencies in which case our preference could not be deduced. Etc. All of this is token-intensive and unreliable.
(Perhaps even more ideally than placing this instruction in AGENTS.md or similar, it could be an inline comment in the User model itself. However, I have found LLMs to be very hit-or-miss when it comes to recognizing directives inside inline comments. And some architecture rules can't be easily expressed via inline comments in a particular file...)
Like you said, an inline comment probably makes more sense. Sure, maybe it misses it, then you can just tell it. Better than putting such granular details in a CLAUDE/AGENTS.md and better than having disparate markdown files or decisions documents that get unwieldy.
You know the contents of AGENTS.md/CLAUDE.md just get injected into your context window by the harness? And that it's functionally equivalent to typing the same thing into every session?
There's no difference between typing "keep dependencies out of User" at the beginning of every session, and putting that into AGENTS.md, other than wasting a bunch of time.
Sure, maybe it misses it, then you can just tell it.
Re-doing that work after it gets things wrong is going to take roughly 1-3 orders of magnitude more tokens than including it in AGENTS.md/CLAUDE.md.Here's a good concrete example. I've caught Claude writing python scripts to parse JSON instead of simply using `jq`. I initially attempted to correct this by adding `jq` to my allow list, but it still always reached for Python scripting first, before eventually figuring out it could use `jq`. You think that burned a few extra tokens vs. just telling it `use jq to parse JSON` in AGENTS.md?
I don't really care how claude parses json. I'd say you're just wasting your time caring about that detail. I definitely wouldn't put that kind of useless info into my claude.md/agents.md.
If I DID care, yeah, I think it makes more sense to just correct the action. I'm not parsing json in every session so why would I want it in context for every session AND every sub-agent's session?
I eagerly await your misunderstanding!
I don't mind SOME documentation. I'm growing extremely frustrated with the absolute MOUNTAIN of docs, comments, decision documents, etc being created. I'm so fucking tired of the LLM rube goldberg machines being built.
- it feels productive and has illusion of value
- it seems harmless/free, refusing means that at best you'll be at net zero
- if you do refuse, now you actually need to think about what should be documented
So yeah, LLM-generated garbage will inevitably spread.
The good news is that teams that do have the discipline are way more competitive today.
I've got one example that's at the top of my mind driving me crazy right now.
We have to move some git repos. It requires making sure there are no secrets in git history or committed, in a lot of cases we're just archiving, creating new repos, copying the code over. We just have to mostly mirror the config of the prior repo. There are probably about 100 repos that we need to move. I farmed this out to 3 of my ICs, a Sr and 2 jr/mid level guys. I thought one of them would figure out a way to mass move them and worst case they just move them real quick... shouldn't be too hard.
Well my Sr has spent the past 6-8 weeks building claude skills/plugins to move the repos. Every week he finds an edge case or the tool doesn't do something perfect, so every sprint he has "fixes" for it.
So far he's moved 5 repos. The jr/mid level guys, one has moved 30 the other has moved 10. I grabbed one of these tasks, threw claude at it (again, no tools or any bullshit), gave it a quick prompt, knocked out one of these in 20 minutes (while multi-tasking).
This same IC has built an MR review tool. It takes 20-60 minutes to run and costs $20-$50+ each run. Wants to put it into our CI/CD so that it runs on MRs. If you include all the pipelines that would run because of a dev updating a branch in an MR we're talking hundreds to thousands of pipelines a day.
I thought this dev was token maxing or just trying to look busy and I've had to sit down and talk to him, pull him aside for quick 1 on 1s, and to be honest, I don't get the feeling that he knows what he's doing is a gigantic waste of time. The LLMs are sycophants that glaze you. You think everything coming out of it is GREAT work because it tells you that. And I think this is happening across all layers of every org out there in various degrees.
Review tool is an interesting one - I find LLMs to be quite useful for code reviews, but people should really stop associating code reviews with Github/Gitlab and running them as a part of CI. Those platforms were built for humans because reviewing patches over email was a terrible experience. Clankers have no problem with raw patches, in fact they prefer them over fancy web UI. Code hosting platforms are not a good place for agentic code reviews, and they never will be.
Most LLM reviews should only run locally on developer's machine, in your regular coding harness.
* vanilla claude with our MCPs + a decent prompt (like a paragraph tops) * an agent that was previously made, much more concise skills * the mr tool this IC created
The MR tool didn't even catch some of the stuff in our domain that the others caught and it took 45minutes, cost $20 vs 4mins for vanilla (<$1) and 7mins for the agent ($3ish)
And yeah, code-reviews are supposed to be by humans, but we have a ton of reviews come in where there's a ton of low hanging fruit. This tool is meant to be a gate between getting an actual human code review and some of the slop coming in these days.
I'll go through a bunch more MR reviews, see how things go. But I've gotta start figuring out how to test the documentation, decision docs, and other plugins/skills.