Most teams I work with hardcode a single “golden path” for agents, then rely on dashboards, alerts, and tribal knowledge to notice when behavior degrades. By the time someone debugs model choice, tool params, or prompt drift, the environment has already changed again. The feedback loop is slow and brittle.
What’s interesting here is the explicit shift from observability to outcome-driven control. Routing based on actual production success rather than static benchmarks or offline evals aligns with how reliability engineering evolved in other domains. We moved from “what happened?” to “what should the system do next?” years ago.
A couple of questions I’m curious about:
- How do you define and normalize “success” across heterogeneous tasks without overfitting to short-term signals?
- How do you prevent oscillation or path thrashing when outcomes are noisy or sparse?
- Is there a notion of confidence or regret baked into the routing decisions over time?
Overall, this feels less like a router and more like an autonomous control plane for agents. If it holds up under real-world variance, this is a meaningful step toward agents that are self-healing rather than constantly babysat.