Five months treating bugs like patients and coding agents like a medical team
cockroachlabs.com
cockroachlabs.com
That type of encoding (starting with the statement: you're at a teaching hospital) is profoundly more efficient than having to deeply and reliably articulate what each of the various roles your agents embody are, let alone their interactions.
Perhaps there is more room for any of us to consider what existing systems or identities are described abundantly in training sets and can be leveraged to ritualistically encode these social/civic/cultural dynamics saying: "act as though you're ___." Not exactly a new point, but one this write-up certainly is pulling me towards.
Thank you for sharing and the care you put into writing this!
I will confess however that some people internally find the model confusing. For example, one user couldn't remember that to get an issue actioned, they needed to put it into the "waiting room". They wanted to be able to action their issues without learning how complex medical systems operate. There seems to be more work on our plates to make this resonate with all users.
I compare this to OpenAI's symphony prompt which, at a high level, does exactly the same thing outlined here except model selection and doesn't really have the need for roles or hospital metaphors.
The comparison with OpenAI's Symphony is valid (the two systems do broadly the same thing). One thing that's different about Sinai is that it keeps the code author and reviewer in separate agents with separate context. At the time we built it, we suspected that this would lead to better outcomes, but again, we haven't validated that it does.
I confess that there's a lot more validation we could do with this model but we just haven't found the time. One thing we don't cover in the post is that this is a side project for both of us, so we have less time than we'd like to perform experimentation and validation.
In the end, medical students learn to become doctors. Cheaper models never learn to become more capable models, so the metaphor is not apt.
It was definitely human driven, but I believe that the agents did the actual trimming.
Once we started to review the code in more detail (and put the hospital into "human approval mode" to stop the momentum) we found one major thing that we corrected.
When issues were decomposed into a DAG, sibling issues would often defer work that they thought that their sibling(s) would handle, and that work sometimes just got dropped. This was because there was no way for sibling issues to communicate with each other. We've since added a deferred scope ledger and policy. If any issue is going to defer scope it must comment about the deferral in the code, and write it to a special ledger with the rationale, and how/when the deferral should be actioned. This has prevented several items from falling through the cracks, but we're constantly refining what is a valid deferral and what should be fixed immediately.
There are several other things that we've corrected over time, which makes me think that another blog post is in order. The full list would be too extensive to try and address in the comments section.
One thing I've been experimenting with in my own agent orchestration system is an agent that hangs around and does post-merge acceptance testing after the work ships. Any plans to add a follow-up phase? Travel nurse?