I think this line around "context tuning" is super interesting - I see a future where, for every model release, devs go and update their CLAUDE.md / skills to adapt to new model behavior.
62 karma · joined September 17, 2025
I think this line around "context tuning" is super interesting - I see a future where, for every model release, devs go and update their CLAUDE.md / skills to adapt to new model behavior.
also I keep hearing complaints that opus is nerfed, but IMO it's nice to have objective data to back that. I feel like half of the nerfing complaints are people getting past honeymoon phase...
- find N tasks from your repo that serve as good representation of what you want the agent to do with the task - run agent with old skill/new skill against those tasks - measure test pass rate / other quality metrics that you care about with skill - token usage, speed, alignment, ... - tests aren't a great measure alone - I've found them to be almost bimodal (most models either pass/fail) and not a good differentiator - use this to make decisions about what to do with the skill - keep skill A, promote skill B, or keep tweaking
I've also had success with an "autoresearch" variant of this, where I have my agent run these tests in a loop and optimize for the scores I'm grading o
more general point being that we need to be methodical about the way we manage agent context. if lat.md shows a 10% broad improvement in agent perf in my repo, then I would certainly push for adoption. until then, vibes aren't enough
I wrote a short blog about this phenomenon here if you're interested https://www.stet.sh/blog/both-pass
also +1 on placing heavy emphasis on the plan. if you have a good plan, then the code becomes trivial. I have started doing a 70/30 or even 80/20 split of time spent on plan / time implementing & reviewing
Evals are broken - OpenAI showed that SWE Bench Verified was in the training data - models were able to reconstruct the changes from memory (https://openai.com/index/why-we-no-longer-evaluate-swe-bench...)
However, this doesn't mean we should completely give up on benchmarking. In fact, as models get more intelligent, and we give them more autonomy, I believe that tracking agent alignment to your coding standards becomes even more important.
What I've been exploring is making a benchmark that is unique per-repo - answering the question of how does the coding agent perform in my repo doing my tasks with my context. No longer do we have to trust general benchmarks.
Of course there will still be difficulties and limitations, but it's a step towards giving devs more information about agent performance, and allowing them to use that information to tweak and optimize the agent further
My theory is that at least some of this is solvable with prompting / orchestration - the question is how to measure and improve that metric. i.e. how do we know which of Claude/Codex/Cursor/Whoever is going to produce the best, most maintainable code *in our codebase*? And how do we measure how that changes over time, with model/harness updates?
How good is the human at using the agent, and how good is the agent itself?
I agree with the thesis here that the traditional DORA metrics don't have as much signal in an agentic world. I like the metrics mentioned in the article - another one I would propose is "number of turns" - the idea being that, if the agent goes off course, the human has to spend more turns course correcting the agent, whereas if the agent is aligned, there are just a few turns in the conversation.
For the "measuring the agent itself" part, I'm convinced that traditional benchmarks are broken, and that we need a way to measure our coding agents on our tasks, and anything else is irrelevant/noise.
having linters is super important IMO - I never try to make the AI do a linter's job. let the AI focus on the hard stuff - architecture, maintainability, cleanliness, and the linter can handle the boring pieces.
I also definitely see the AI making changes that are way larger than necessary. I try to capture that in the eval by comparing a "footprint risk" which is essentially how many unnecessary changes did the AI make vs the original PR.
I would certainly like to move beyond using PRs as a sole source of truth, since humans don't always write great code either. Maybe having LLM-as-a-judge looking for scope creep/bloat would be a decent band-aid?
Interestingly, I had a similar finding where, on the 3 open-source repos I ran evals on, the models (5.1-codex-mini, 5.3-codex, 5.4) all had relatively similar test scores, but when looking at other metrics, such as code quality, or equivalence to the original PR the task was based on, they had massive differences. posted results here if anyone is curious https://www.stet.sh/leaderboard
- AI to evaluate itself (eg ask claude to test out its own skill) - custom built platform (I see interest in this space)
I've actually been thinking about this problem a lot and am working on making a custom eval runner for your codebase. What would your usecase be for this?
By delegating to sub agents (eg for brainstorming or review), you can break out of local maxima while not using quite as many more tokens.
Additionally, when doing any sort of complex task, I do research -> plan -> implement -> review, clearing context after each stage. In that case, would I want to make 7x research docs, 7x plans, etc.? probably not. Instead, a more prudent use of tokens might be to have Claude do research+planning, and have Codex do a review of that plan prior to implementation.
would be cool to see/extend the ttl on the transcripts
https://agentexports.com/v/g92c990f0cfb9962a#lClk4hHKdmv52Nx... https://agentexports.com/v/ga07365c8abedbd2a#5boQrM0ZUz78LIF...
At least for me, there's large value in consuming bigger volumes of Chinese to get me used to pattern-matching on the characters, as opposed to only reading a smaller amount of harder characters that I'm less likely to actually encounter
I haven't dug into the github repo but I'm curious if by "guided decoding" you're referring to logit bias (which I use), or actual token blocking? Interested to know how this works technically.
(shameless self plug) I've actually been solving a similar problem for Mandarin learning - but from the comprehensible input side rather than the dictionary side:
https://koucai.chat - basically AI Mandarin penpals that write at your level
My approach uses logit bias to generate n+1 comprehensible input (essentially artificially raising the probability of the tokens that correspond to the user's vocabulary). Notably I didn't add the concept of a "regeneration loop" (otherwise there would be no +1 in N+1) but think it's a good idea.
Really curious about the grammar issues you mentioned - I also experimented with the idea of an AI-enhanced dictionary (given that the free chinese-english dictionary I have is lacking good examples) but determined that the generated output didn't meet my quality standards. Have you found any models that handle measure words reliably?
I have repeatable workflows that harness the benefits of multiple agents. Repeatable workflows drive consistent results for single agents. Using multiple agents allows you to fully explore the problem space.
An example of using these concepts harmoniously would be creating a custom slash command that spawns sub-agents that each have custom prompts, causing them to do more exploration. The commands + agent prompts make the flow repeatable + improvable
1. Effectively infinite engaging comprehensible input at your level 2. Fantastic way to practice new vocabulary and grammar patterns (AI can provide correction for mistakes) 3. Somewhat fun - if you view chat as a choose your own adventure, the experience becomes more interesting
However, due to the more user-driven approach to this learning method (output-focused, user has to put in effort to chat with the AI and get feedback), there is more friction with using the tool. This isn't necessarily a bad thing - in fact, more friction can lead to more meaningful experiences. That being said, I believe the market will push tools to be low friction and low effort (i.e. gamified apps) that are focused on consumption rather than tools that require more user effort.
just my 2c from a fellow builder. if curious, check it out here! would love any feedback