Even if in the absolute the ceiling remains low, it’s interesting the degree to which good context engineering raises it
1,188 karma · joined March 2, 2018
Even if in the absolute the ceiling remains low, it’s interesting the degree to which good context engineering raises it
If you're going to write about something that's been true and discussed widely online for a year+, at least have the awareness/integrity to not brand it as "this new thing is happening".
a tale as old as time - my second job out of college back in like 2016, I landed at the tail end of a 3-month feature-freeze refactor project. was pitched to the CEO as 1-month, sprawled out to 3 months, still wasn't finished. Non-technical teams were pissed, technical teams were exhausted, all hope was lost. Ended up cutting a bunch of scope and slopping out a bunch of bugs anyway.
sure, readme.md is a great place to put content. But there's things I'd put in a readme that I'd never put in a claude.md if we want to squeeze the most out of these models.
Further, claude/agents.md have special quality-of-life mechanics with the coding agent harnesses like e.g. `injecting this file into the context window whenever an agent touches this directory, no matter whether the model wants to read it or not`
> What people often forget about LLMs is that they are largely trained on public information which means that nothing new needs to be invented.
I don't think this is relevant at all - when you're working with coding agents, the more you can finesse and manage every token that goes into your model and how its presented, the better results you can get. And the public data that goes into the models is near useless if you're working in a complex codebase, compared to the results you can get if you invest time into how context is collected and presented to your agent.
Totally matches my experience- the act of planning the work, defining what you want and what you don’t, ordering the steps and declaring the verification workflows—-whether I write it or another engineer writes it, it makes the review step so much easier from a cognitive load perspective.
You want control over and visibility into what’s being compacted, and /compact doesn’t do great on either
not exactly valuable as guidance since programming languages are very easy to verify, but the https://ghuntley.com/ralph post is an example of whats possible on the very extreme end of the spectrum
All because they have been forced to master technical communication at scale.
but the reason I wrote this (and maybe a side effect of the SF bubble) is MOST of the people I have talked to, from 3-person startups to 1000+ employee public companies, are in a state where this feels novel and valuable, not a foregone conclusion or something happening automatically
Codex and Claude code are neck and neck, but we made the decision to go all in on opus 4, as there are compounding returns in optimizing prompts and building intuition for a specific model
That said I have tested these prompts on codex, amp, opencode, even grok 4 fast via codebuff, and they still work decently well
But they are heavily optimized from our work with opus in particular
"what happens if we end up owning this codebase but don't know how it works / don't know how to steer a model on how to make progress"
There are two common problems w/ primarily-AI-written code
1. Unfamiliar codebase -> research lets you get up to speed quickly on flows and functionality
2. Giant PR Reviews Suck -> plans give you ordered context on what's changing and why
Mitchell has praised ampcode for the thread sharing, another good solution to #2 - https://x.com/mitchellh/status/1963277478795026484
> While the cancelation PR required a little more love to take things over the line, we got incredible progress in just a day.