The metric we actually care about is closer to "cost per accepted outcome" which would roughly be "total inference + retries + failures + the human effort required to verify, correct, or resume the work."
The first few are easy to instrument.
But "human effort verification" and “accepted” are much harder to define consistently across different agent workflows.
For our code review product we have a decent proxy because we can measure user reactions and whether findings are actionable and if it's merged/accepted.
For more open-ended agent tasks, I don’t think the industry has a particularly good answer yet... I would love a point if there is one?
The cheaper model isn’t actually cheaper if a human has to spend the savings babysitting it...