I'm skeptical of the quality of probability calibration for models if you aren't giving them training data. The issue is that there is a prior distribution that's unique to your specific data and the calibration is really sensitive to that.
Before now you had to train a model on your specific classification problem, now these new models don't require any specific training at all to do pretty well on novel problems.
There are classes of problem where this shifts the economics from "paying a human to do this is cheaper than AI" to "the software is now cheaper than the human"
I've asked twice now about what I'm missing and for a specific use case where you can't just do this with a regular LLM call and nobody has replied that so if you have the answer that would be great. Looking for something specific instead of just it's faster or cheaper which is definitely nice but I'm just not seeing what this opens up that was not previously possible
For instance if you have a predictive model for market prices that is not calibrated that's... nice. If you have a calibrated model you can add a Kelly better and you have a trading strategy that makes money. Similarly if you are classifying articles or images or other contents to make a feed you might believe that 70% or 95% or some other level of precision is "good enough" and you can set the knob and turn on the cruise control.
And theoretically will give you better answers statistically as it's calibrated.