How is this part tackled when all that you have is GH issues? Doesn’t this work only for the most trivial issues?
How is this part tackled when all that you have is GH issues? Doesn’t this work only for the most trivial issues?
You can afford a lot of extra guardrails and process to ensure sufficient quality when the result is a system that gets improved autonomously 24/7.
I'm on my way home from a client, and meanwhile another project has spent the last 10 hours improving with no involvement from me. I spent a few minutes reviewing things this morning, after it's spent the whole night improving unattended.
I am all for delegating everything to AI agents, but it just becomes a mess over time if you don’t steer things often enough.
EDIT: I'll add that you can't expect it to guess what you want, but you can let it manage how it delivers it. We don't expect e.g. a product manager to dictate how developers deliver the code, just what the acceptance criteria is, and that's where I'm headed.
From the project: "The plugin enqueues the input and a daemon picks it up - planning, building, reviewing, and validating autonomously."
The part that is not clear to me (and causes most problems for me) is the "validating". It makes a mistake, or decides mocking an interface is fine, etc. declares success and moves on to the next. The bigger the project the more small mistakes compound. It sounds like the agent is doing the validation. What's the approach here for validation?