My basics are I have a test suite that: 1. Finds newest models of my versions 2. Inferences every single model and a few providers for each with a short problem 3. Analyze latency and if a model failed the stupid simple questions drop it and the provider 4. Run larger context haystack kinds of problems.
It's cheap and fast less than 5$ so I can do this daily, hourly, whatever depending on how much I care. If it's mission critical I would say you need a two day study running once an hour to know the STD of model variance.
Then lock a top 3 contenders via latency dropping routing.
Is this easy? No. Is it cheap? Also no. Is it better than just using a trusted labs api? Also not really.
But it does give you exponentially more flexibility. Being able to run 10 unique models at the flick of a switch on a problem for pareto front analysis is amazing. And giving a dropdown for customers for multiple model options is powerful.