Trying to fix syntax errors in strong interpolation on a 5-minute-delay loop is hell.
So my agent just listens for green checks and no PR comments and loops until those conditions are met.
I disbelieve this works in anything other than a toy codebase (or an incredibly fine-grained microservice).
The 70% is amazing! But a 30% failure rate requires intense supervision.
Might tend to deviate and waste time, needs guiding once in a while, and to check what is it spewing out, point it in the correct direction.
If I had to output the code myself, would take around 8 hours of constant writing to get around 1k LoC of code. For FUSE level tricky stuff, I might need to spend 3 weeks for 10 LoC. Very easy to burnout and build pain.