At first I thought that was pretty compelling, since it includes more edge cases and examples that you otherwise miss.
In the end all that planning still results in a lot of pretty mediocre code that I ended up throwing away most of the time.
Maybe there is a learning curve and I need to tweak the requirements more tho.
For me personally, the most successful approach has been a fast iteration loop with small and focused problems. Being able to generate prototypes based on your actual code and exploring different solutions has been very productive. Interestingly, I kind of have a similar workflow where I use Copilot in ask mode for exploration, before switching to agent mode for implementation, sounds similar to Kiro, but somehow it’s more successful.
Anyways, trying to generate lots of code at once has almost always been a disaster and even the most detailed prompt doesn’t really help much. I’d love to see how the code and projects of people claiming to run more than 5 LLMs concurrently look like, because with the tools I’m using, that would be a mess pretty fast.
So, is it good for writing requirements, and creating design, if not for coding?
In my experience with workflows that let humans and humans (let alone AIs) collaborate effectively, they are NP-hard problems.
I believe people are being honest when they say these things speed them up, because I'm sure it does seem that way to them. But reality doesn't line up with the perception.
A greenfield startup however with agentic coding in it's DNA will be able to run loops around a big company with lots of human bottlenecks.
The question becomes, will greenfield startups, doing agentic coding from the ground up, replace big companies with these human bottlenecks like you describe?
What does a startup, built using agentic coding with proper engineering practices, look like when it becomes a big corporation & succeeds?
I can believe a single developer with one agent doing some small stuff and using some other LLM tools can get a modest productivity boost. But having 5 or 10 of these things doing shit all at once? No way. Any gains are offset by having to merge and quality check all that work.
Every feature I’ve asked Claude Code to write was one I could’ve written myself.
And I’m quite certain it’s faster for my use case.
I won’t be bothered if you choose to ignore agents but the “it’s just useful for the inept” argument is condescending.
You spend a few minutes generating a spec, then agents go off and do their coding, often lasting 10-30 minutes, including running and fixing lints, adding and running tests, ...
Then you come back and review.
But you had 10 of these running at the same time!
You become a manager of AI agents.
For many, this will be a shitty way to spend their time.... But it is very likely the future of this profession.
You want to do that, but Ill bet money you arent doing it.
Thats the problem: this is speculative; maybe it scales sometimes, but mostly people do not work on ten things at once.
“Fix the landing page”
“I’ll make you ten new ones!”
“No. Calm down. Fix this one, and do it now, not when youre finished playing with your prompts”
There are legitimate times when complex pieces of work decompose into parallel tasks, but its the exception not the norm.
Most complex work has linked dependencies that need to be done in order.
Remember the mythical man month? Anyone? Anyone???!!??
You can't just add “more parallel” to get things done faster.
Codex / Jules etc make this pretty easy.
It's often not a sustainable pace with where the current tooling is at, though.
Especially because you still need to do manual fixes and cleanups quite often.
Mhm. Money -> to the dealer.
Anyway… watch the videos the OP has of the coding live streams. Thats the most interesting part of this post: actual real examples of people really using these tools in a way that is transferable and specifically detailed enough to copy and do yourself.
You can’t do 10 of these processes at once, because there’s 8 minutes of human administration which can’t be parallelised for every ~20min block of parallelisable work undertaken by Claude. You can have two, and intermittently three, parallel process at once under the regime described here.
That coupled with the fact that you have to meticulously review every single thing the AI does is going to obliterate any perceived gains you get from going through all the trouble to set this up. And on top of that it's going to be expensive as fuck quick on a non trivial code base.
And before someone says "well you don't have to be that thorough with reviews", in a professional settings absolutely you do. Every single AI policy in every single company out there makes the employee using the tool solely responsible for the output of the AI. Maybe you can speed run when you're fucking around on your own, but you would have to be a total moron to risk your job by not being thorough. And the more mission critical the software the more thorough you have to be.
At the end of the day a human with some degree of expertise is the bottleneck. And we are decades away from these things being able to replace a human.
Joke's on you (and me? and I guess on us as a profession?).
This can all be done autonomously without user interaction. Now many bugs can be few lines of code and might be relatively easy to review. Some of these bug fixes may fail, may be wrong etc. but even if half of them were good, this is absolutely worth it. In my specific experience the success rate was around 70%, and the rest of the fixes were not all worthless but provided some more insight into the bug.
We are years into this, and while the models have gotten better, the guard rails that have to be put on these things to keep the outputs even semi useful are crazy. Look into the system prompts for Claude sometime. And then we have to layer all these additional workflows on top... Despite the hype I don't see any way we get to this actually being a more productive way to work anytime soon.
And not only are we paying money for the privilege to work slower (in some cases people are shelling out for multiple services) but we're paying with our time. There is no way working this way doesn't degrade your fundamental skills, and (maybe) worse the understanding of how things actually work.
Although I suppose we can all take solice in the fact that our jobs aren't going anywhere soon. If this is what it takes to make these things work.
I don't blame people who think this. I've stopped visiting Ai Subreddits because the average comment and post is just terrible, with some straight up delusional.
But broadly speaking - in my experience - either you have your documentation set up correctly and cleanly such that a new junior hire could come in and build or fix something in a few days without too many questions. Or you don't. That same distinction seems to cut between teams who get the most out of AI and those that insist everybody must be losing more time than it costs.
---
I suspect we could even flip it around: the cost it takes to get an AI functioning in your code base is a good proxy for technical debt.