I am talking about product quality and maintainability. Both are more than adequate.
I know this because I have worked on it for an estimated 300 hours. Has the author practiced a similar approach for even a week? I doubt it.
If this works for you - awesome. Until it doesn't.
As always there is 0 code or link. All talk.
I will not publish my app on GitHub for free. It's a paid app, and I am putting in the hours not for your approval, but for commercial gain.
I also do not think it wise to link my HN account to my real name and expose my opinions and comments to my employer and colleagues.
We do not require links to your app. What people are expecting is a description of your approach and sample outputs. So that someone else can try it and have the same standard of output. That's how you make a point that your approach is good.
When we buy books like "The Practice of Programming" or "The Pragmatic Programmer", it's because we are hoping to learn useful and productive behaviors. It isn't to hear boasts about how good the authors are good at using tools.
Even self-help books follow this pattern: Do this, expect that. They're not "Have you tried this too" or "I don't know about you, but I've got good results myself".
If I had any special approach, I would be reluctant to share it with my potential competitors.
That said, I do not. It just works.
Meanwhile people here are posting the thesis that agentic development without careful code review results in an unmaintainable application.
I theorize that this is not something they experienced in practice, because it did not happen for me.
> I theorize that this is not something they experienced in practice, because it did not happen for me.
Are you currently maintaining the application? Like it's in production with paying users? You've only been on the app for 4 months. Compare that to something like Emacs that has been going for 40+ years. You can make a better case when you've been on prod for a few years.
No, my app is not published yet. It will probably take another month, with hopefully no complications arising out of the AppStore review process.
Then, I hope the ad campaign financials work out to compete with old apps of a lower quality that already boast no less than a million reviews.
I get it, you want me to make a case that can objectively convince you of the usefulness of agentic development without code review.
From my perspective, I have no interest in doing so, and I can only share my experience so far. In a few months time we will know more objectively whether my ambitions paid off.
Until then, you will either have to take my word for the quality of the product, or spend tens to hundreds of hours of effort in trying the approach for yourself. OP's article does not contain any specifics for where and how supposedly agentic development failed him either.
If anything there are clear counter example to your claim, such as the major provider agent harnesses which are all almost always fully vibe coded, and riddled with bugs and regressions that make using them painful for users. The only reason people put up with it is because competition in the space is still limited.
That's why I can say with full confidence that it works.
I doubt it will magically all fall apart the moment an external user touches it, or that I will expand the scope dramatically in the near future.
The agent harnesses are an interesting topic. I believe they have large teams shipping a ton of changes weekly. In that environment, is it realistic to expect rock solid software with such a feature set to be developed in a few months and shipped to 10M users?
Yes, the whole argument is that LLMs make shipping quality software at scale in months possible. You're arguing out of two sides of your mouth now. On the one hand, agents have enabled you to build a bullet-proof high quality app in a few months as a solo dev. On the other hand it's supposedly unreasonable to expect a team of engineers with lots of funding and lots of expertise to use those same LLMs to build high quality software in a few months.
The main difference is you are still working in a vacuum and the harness teams have actually shipped to users. Once you do that, you face significantly more challenges than you do tinkering in isolation. Don't claim a methodology works until you've actually proven so. "My personal closed source pet project that no one but me has ever seen works" is not convincing evidence. My pet dragon who is definitely real but that I can't show anyone else agrees.
I think you understand well that there is a major difference between developing a mobile app solo and 10-20 people working on a coding agent harness of vastly larger scope.
I don't intent to prove anything to you. I am telling you that it works for my mobile app development. Your arguments to the contrary are rather weak.
You're acting like everyone here doesn't have hundreds of hours of experience with LLM coding. We all do, we all know what it's like.
You're simply either lying or wrong. If it's the former, I don't care, you're just an asshole on the internet. If it's the latter, you'll learn eventually and it will be quite painful for you.
I believe many of you use it at work, where you need to get code through review, or voluntarily review the generated code.
Otherwise I have no explanation as to how agentic development is failing for you, while it continues to work on my side in a large code base.
But the opinions expressed are basically polar opposites, so there's clearly something to this.
The easy explanation IMO is that it takes some time to learn how to use LLM's effectively for system development. It's still skilled work, just different skills.
Some put in that effort and see results, others are annoyed that the reality doesn't match the hype and bail.
>You're simply either lying or wrong.
Surely it's possible that he was able to make it work even though you didn't?
If humans aren't needed why have you had to spend 300 hours?
I upvoted your comment BTW because I think you might be right, but I'm not sure. I still see people doing a lot of work, despite LLMs.
I have to prompt for the features, test them, then iterate until the UX is acceptable before I merge.
Usually I would juggle 2-5 topics in parallel, unless one of them demands more of my attention.
While there is no more code review involved, it is a lot of QA work and testing on device.
The bottleneck is that the AI does not have taste, does not know what the product should be, and would happily ship horrible slop without my intervention.
Other than the implementation itself there are other things that need to be handled: researching competitors, keywords, pricing, AppStore preparation, TOS and privacy policy, when do you show the rating prompt, translations, etc.
AI still helps with a lot of it, but it takes time and effort.