For the current state of frontier models, you need to break the steps down so that the LLM understands a process like what you might go through as you expect it (which is often different for everyone).
i.e., get it to agree to a spec, then get it to agree to a build plan, agree on unit test signatures, UI etc as needed, then let it build, ...
"Prompt engineering"
I can take all of those steps, turn them into separate skills, then give them to a product manager or business analyst who makes half your salary, but has far more knowledge about the customers needs than you do.
Why does that require an engineer? I think a business analyst or product manager could do that just fine.
The business analyst use to hand off details to another person who would define detail technical specs like database fields names, type and size. Then the programmer would implement it.
> it's probably an issue with your usage of if
> I've rarely seen a repo and a problem that claude can't chew through with the right prompt
> a skill/PEBKAC issue
But then I remember how Anthropic couldn't fix the flickering issue for many months. It just does not compute.
Is it that people working at Anthropic can't prompt and it's a "skill issue" too? I mean, the terminal does not flicker in a lot of other complex TUI apps that I use every day - Midnight Commander, Emacs, tmux, etc. These are open source, Claude could be prompted to "just do what Midnight Commander does". So what is it?
A good terminal program, like Emacs, uses escape sequences to create viewports, scroll text, etc. These sequences tell the terminal to do these things internally.
All the modern AI tools ignore these escape sequences (which they wouldn't even have to know if they used ncurses) and just do frame-by-frame animation. So when OpenCode wants to scroll text, first it sends the text to the screen, then it sends ^[[1;1H and sends the whole text again, minus the first line or so, adding a new line to the bottom. Then it sends ^[[1;1H again, and sends the whole text AGAIN, cutting another line from the top, and adding another new one to the bottom. As if the terminal was a graphics device and you had to draw each frame.
The flickering was probably due to sending ^[[2J to clear the screen before "drawing" each "frame."
I oversimplified, though. OpenCode sends hundreds and hundreds of cursor-positioning escape sequences per "frame" of "animation".
Even though it doesn't flicker anymore, "drawing frames" makes the scrolling noticeably slow and choppy if you run OpenCode (or any other terminal-based AI agent) on a different machine than the one your monitor is plugged into, even on a wired LAN.
No AI model can give a believable explanation of why these apps (or their TUI libraries) are written this way.
Docker also does this, and it drives me mad.
> Every time I get past the green field stage, I just end up throwing out what it writes half the time since its trash.
Is a skill/PEBKAC issue. You still need to exercise engineering best-practices like decomposing work to the smallest unit before taking a task on, brainstorming design first and implementation last, clearly defining your success criteria and requirements before beginning any work, etc.
I'm on a >10yr old codebase and have been able to get my org to orchestrate entire features, fully unit tested, e2e tested, storybooked, from scratch without touching an IDE. Refactorings and the endless mountain of 80% completed migrations from one pattern to another are now trivially able to offload.
Point your SOTA de jeur at the original docs, a few of the original examples/PRs and have it draft a skill describing the work, the scope, and the success metrics. Iterate on the skill with the main agent by subagenting to test the skill until you are happy with the result and it mostly gets it right with the guardrails you've defined. Again - keep the scope extremely small. It gives much less rope for the agents to hang themselves with and it is less cognitive load when you have to review/test the PR.
Then set up a reasonable cadence for it to execute an autonomous thread on and review when you get comfortable.
----
The issue I've been running into lately is simply that we've got so many PRs coming in that actually doing thorough human reviews on them is not sustainable relative to the rate the team is creating agents to open them and people (especially juniors and mid level) are getting burned out by essentially having entire days where they are just doing code reviews.