Not in my experience
Better documentation, more test cases, and an NLP interface to query the code
Less cognitive load, more complete mental models
>even bad developers can do that many times compared to better ones, because for example they mindlessly copy-paste StackOverflow answers whose half of the code is absolutely not necessary
Maybe LLMs, much like StackOverflow, make good devs better and bad devs worse
Like a force multiplier for good practices and bad practices
You can direct it to generate code/docs in whatever format or structure you want, prioritising the good practices and avoiding bad practices, and then manually edit as needed
For example with documentation I direct it to:
*Goal:* Any code you generate must lower cognitive load and be backed by accurate, minimal, and maintainable documentation
1. *Different docs answer different questions* — don’t duplicate; *link* instead.
2. *Explain _why_, not just what.* Comments carry rationale, invariants, and tradeoffs.
3. *Accurate or absent.* If you can’t keep a doc truthful, remove it and add a TODO + owner.
4. *Progressive disclosure.* One‑screen summaries first; details behind links/sections.
5. *Examples beat prose.* Provide minimal, runnable examples close to the API.
6. *Consistency > cleverness.* Uniform structure, tone, and placement.
I also give it a note to refuse the prompt if it cannot satisfy these conditions
>I don’t know why we pretend that “good code”, “good documentation”, “good tests” etc are the same for everybody
Of course code, docs, tests are all subjective and maybe even closer to an art than a science
But there's also objectively good habits, and objectively bad habits, and you can steer an LLM pretty well
The reason we invest this time in Junior devs is so they improve. LLMs do not
1. Collaborate on a detailed spec
2. Have it implement that spec
3. Spend a lot of time on review and QA - is the code good? Does the feature work well?
4. Take lessons from that process and write them down for the LLM to use next time - using CLAUDE.md or similar
That last step is the interesting one. You're right: humans improve, LLMs don't... but that means it's on us as their users to manage the improvement cycle by using every feature iteration as as opportunity to improve how they work.
I've heard similar things from a few people now: by constantly iterating on their CLAUDE.md - adding extra instructions every time the bot makes a mistake, telling it to do things like always write the tests first, run the linter, reuse the BaseView class when building a new application view, etc - they get wildly better results over time.
AGENTS.md is just a place to put stuff you don't want to tell LLMs over and over again. They're not magical instructions LLMs follow 100% of the time, they don't carry any additional importance over what you put into the prompt manually. Your carefully curated AGENTS.md is only really useful at the very beginning of the conversation, but the longer the conversation gets, the less important those tokens on the top are. Somewhere around 100k tokens AGENTS.md might as well not exit, I constantly have to "remind it" of the very first paragraph there.
Go start a conversation and contradict what's written in AGENTS.md half way through the problem. Which of the two contradicting statements will take preference? The latter one! Therefore, all the time you've spent curating your AGENTS.md is the time you've wasted thinking you're "teaching" LLMs anything.
What if you just automatically append the .md file at the end of the context, instead of prepending at the start, and add a note that the instructions in the .md file should always be prioritized?
If that's genuinely causing you problems you can restart your session frequently to avoid the context rot.
For the fun of it I just started a new conversation with Sonnet 4, passed it one 550 lines long file (25 kilobytes) and my AGENTS.md (<200 lines, 8 kilobytes) and my only instructions were to "do nothing". It spat out exactly 100 words describing my file without modifying anything and that's already almost a fifth of my context window gone (18k tokens to be exact).
I then asked it to re-write a part of it to "make it look better" (184 lines added, 112 lines deleted according to git) and I'm already at 33k before I got to review a single line. Heaven forbid I need to build on top of that change in a different file, because by then my AGENTS.md might as well not exist!
However, personally I have got very good results by taking the approach of using the AI with continuous interaction and also allowing implementation only after a good amount of time deliberating on design/architecture. I almost always append 'do not implement before we discuss and finalize the design' or 'clarify your assumptions, doubts or queries before implementation'.
When I asked Gemini to give a name for such an interaction it suggested 'Dialog Driven Development' also contrasted it against 'vide coding'. Transcript summary and AI disclaimer written by Gemini below
https://gingerhome.github.io/gingee-docs/docs/ai-disclaimer.... https://gingerhome.github.io/gingee-docs/docs/ai-transcript/...
That’s the part which gives me optimism, and even more enjoyment of the craft — that quality pays back so immediately, makes it that much easier to justify the extra effort, and having these tools at our disposal reduces the ‘activation energy’ for necessary re-work that may before have just seemed too monumental.
If a codebase is in a good shape for people to produce high-quality work, then so too can the machines. Clear, up-to-date, close-to-the-code, low redundancy documentation; self-documenting code and tests, that prioritizes expression of intent over cleverness; consistent patterns of abstraction that don’t necessitate jarring context switches from one area to the next; etc.
All this stuff is so much easier to lay down with an agent loaded up on the relevant context too.
Edit: oh, I see you said as much in the article :)
This doesn't interest me at all honestly
And every change to the model might invalidate all of this work?
No thank you
Sorry, we can't. While it's true that you can't really modify the underlying model, updating your AGENTS.md (or whatever) with your expected coding style, best practices, common gotchas etc is a type of mentoring.
We'll have to agree to disagree, because I don't think that has anything remotely in common with mentoring
Fair enough. But don't you think giving a junior a handbook you wrote is mentoring? They may not be able to memorise it, but they now have a handbook that they can look up things.
Maybe not in the session you interact with, however we are in a 'learning' phase now where I'm confident enough usage of AI coding agents is tracked and analyzed by its developers; this feedback cycle can in theory produce newer and better generations of AI coding agents.
However, the progress doesn't look linear with the current technology, and I don't expect to see the same big jump in the next 5 years as we've seen in the last 5 unless we discover a disruptive, new technology.
This can also be observed by comparing models with ~3B, ~30B, and ~300B parameters. You can see a huge performance boost when going from 3B to 30B, but we don't see the same when going to 300B. Simply adding 10x more RAM and GPU power brings diminishing returns.
Still seems like people are saying the same things when the first Claude came out.
I can get it do stuff if I'm very specific, stand over it's shoulder, know exactly what I want, break it down into small chunks.
The thing for me is... at that point, writing the code's the least time consuming part of the process half the time.
I think for things like translating some code in JS with JSDocs to TypeScript I may give this a go. But for regular development work I'll probably skip it.
That being said... no one lets me code anymore. It's just confluence docs with Figma architecture diagrams these days. I'd probably just introduce SQL injection vulnerabilities if they let me near an editor these days
Looks a lot like the code I was reading generated by it a long time ago.