1. write a sketch of a spec by hand
2. have the llm review the document and question me until it can generate a spec
3. review the spec and revise where needed
4. have it write an implementation plan
5. another round or revision/review
6. executing the plan step by step through the plan, plausing between each step to see if we are still on course and if the decisions it made track with my understanding of what we are doing.
I've been working for a couple of hours tonight, the total cost of the session is €0.6.
it's not the build this thing end to end, but also not quite write function x for me. It is still a lot of manual review, but I find I really need it to even discover what I actually want to build. I just cannot imagine building something in a single shot and getting something that actually has value (unless it is basically a clone of an existing thing). To me the whole value of ai right now is that it's now very cheap to build custom software that exactly matches your preferences.
The one-shot capabilities of frontier models are nice for demonstration purposes, but not actually all that useful to me - the result is often an amalgamation of ad-hoc ideas the agent came up with, poor UI and a hodge-podge of a data model, but at least demonstrating clearly where the specs are lacking.
I agree that in practice, iterative design with lots of review and hand-holding are needed to get quality results. Generated code (I mostly program C++) is often bloated and not succinct or "simple" enough to my tastes, so it requires multiple rounds of cleanup.
For core logic, I find having the agent review code is often faster and more valuable than having it implement it in the first place - it can find and fill in the things I missed.
For my own sanity it is very important that I remain in control and understand the generated code when quality and maintainability are goals of the project.
For myself, with ChatGPT for example, it regularly gets extremely slow if there is a lot of text in a single conversation. Especially if I try to scroll up.
Some pages have been strangely just broken for a while now as well. Usage analytics just renders lots of these duplicate "Usage history" components, where the data just never loads: https://imgur.com/a/vpcQIiw.png
I can totally imagine a scenario where some agent built it, another tested and approved it, and nobody at OpenAI even looked at it once or knows that it's like this.
Now this I have definitely experienced. I chalk that up to it being an electron app, but who knows.
https://www.youtube.com/watch?v=WAeHgE94rVo
The system performs quotation attribution on my local hardware for my near-future, hard sci-fi novel (having nearly 500 quotations) with over 97% accuracy.
The initial prototype was developed quite quickly, but numerous successive iterations were required to fix numerous gaffs by Opus 5 (because it doesn't actually _understand_ what it takes to make general-purpose audiobook narration software).