This implies you didn’t run any sort of evaluations? It is not realistic for any sort of production use case to do this.
Though honestly I'm also surprised by the hype around Jev from a POV of "wait, are so many people just building on these by using them for classification tasks vs something more multi-step or generative?"
Jev and this decisions api are mostly useful for inference at scale in a workload where cost and latency matter… and that’s where evals become crucial. Could coding tools use it? Sure, but that’s probably a special case.
This same suite can simply be run very handoff to switch to a new prod model.