I've been using a prompt management service (cloud) for 2 months and am pretty happy that I can define prompts and check quality outside the code, and it helps me to do the routing manually.
Having benchmarks (I assume this is ongoing?) would make it even more interesting, as I wouldn't have to manually manage the routing each time I add a new model.
You mentioned you don't have a margin right now, but how about keeping logs, the dashboard for cost, and benchmarking?