Here's how the article starts: "Agentic engineering has become so good that it now writes pretty much 100% of my code. And yet I see so many folks trying to solve issues and generating these elaborated charades instead of getting sh*t done."
Here's how it continues:
- I run between 3-8 in parallel
- My agents do git atomic commits, I iterated a lot on the agents file: https://gist.github.com/steipete/d3b9db3fa8eb1d1a692b7656217...
- I currently have 4 OpenAI subs and 1 Anthropic sub, so my overall costs are around 1k/month for basically unlimited tokens.
- My current approach is usually that I start a discussion with codex, I paste in some websites, some ideas, ask it to read code, and we flesh out a new feature together.
- If you do a bigger refactor, codex often stops with a mid-work reply. Queue up continue messages if you wanna go away and just see it done
- When things get hard, prompting and adding some trigger words like “take your time” “comprehensive” “read all code that could be related” “create possible hypothesis” makes codex solve even the trickiest problems.
- My Agent file is currently ~800 lines long and feels like a collection of organizational scar tissue. I didn’t write it, codex did.
It's the same magical incantations and elaborated charades as everyone does. The "the no-bs Way of Agentic Engineering" is full of bs and has nothing concrete except a single link to a bunch of incantations for agents. No idea what his actual "website + tauri app + mobile app" is that he build 100% with AI, but depending on actual functionality, after burning $1000 a month on tokens you may actually have a fully functioning app in React + Typescript with little human supervision.
Yeah at this point you could hire a software developer.
Though I'm aligned that I don't (yet) believe in this "AI writes all my code for me" statements.
I've scrolled bit more. I think in the past 50-100 tweets you only wrote thee talking about this, one of them proudly showing a mistake (invalid tweets containing the same text): https://x.com/steipete/status/1978229441802162548
So, I have to follow you on twitter and sift through garbage indistinguishable from all such "look how great is codex" and "this is my shamanic ritual that works I promise" to maybe see something you work on.
No thank you. I will make my judgement from the long-form article you posted.
And, as I said: depending on actual functionality, after burning $1000 a month on tokens you may actually have a fully functioning app in React + Typescript with little human supervision. I might do the same for anything Twitter-related because I couldn't be arsed to work with Twitter or Twitter APIs.
I think it's good to keep up with what early adopters are doing, but I'm not too fussed about missing something. The plugins is a good example: A few weeks ago there was a post on HN where someone said they are using 18 or 25 or whatever plugins and it's the future, now this person says they are using none. I'm still waiting for the dust to settle, I'm not in a rush.
The trick is to create deterministic hurdles the LLM has to jump over. Tests, linting, benchmarks, etc. You can even do this with diff size to enforce simpler code, tell an agent to develop a feature and keep the character count of the diff below some threshold, and it'll iterate on pruning the solution.
I've started getting desperate to the point of saying 1) "never. never, ever add or remove features without consulting me first and getting approval." Then eventually, 2) appended to the previous "The last rule is the most important rule, because you keep doing it and I need you to stop doing it." Then finally 3), "THE LAST RULE IS THE MOST IMPORTANT RULE, BECAUSE YOU KEEP DOING IT AND I NEED YOU TO STOP DOING IT."
3/4 of my AI bugs are the AI making changes to the functionality of the code when I'm not looking, or repeatedly reinserting bugs that had been previously removed. The most valuable thing I'm getting it to do is to refactor the code it already wrote into shorter well-named functions (during which it still inevitably adds and removes behavior), because it means that I can just debug by hand and stop demanding over and over again that it not ignore what I said.
But, of course, it's not ignoring me, it's not thinking at all. Trying to look for the magic words to keep it from ignoring me and lying about it is just an illusion of control. The thing that will knock it off it's dumb track is likely just a lucky random seed during the 13th attempt. Then, like a sports fan, I add the lucky underwear to my instructions.
edit: the "I'll get AI to write the AI prompt so it will be perfect" stuff is so much voodoo. LLMs have no special insight into what will make LLMs work correctly. I probably should have stopped that last sentence after the word "insight." Feed them a sample prompt that you say doesn't work, and it will explain to you exactly why it's so bad, and could never work. Feed them the same prompt and ask why it works so well, and it will tell you how perfectly crafted it is and why. Then it will offer to tell you how it could be improved.