Brings to mind the quote attributed to Samuelson and Keynes: “When events change, I change my mind. What do you do?”
Brings to mind the quote attributed to Samuelson and Keynes: “When events change, I change my mind. What do you do?”
Assume a model states there is a 99% likelihood something will occur.
Now the data changes, and the likelihood drops to 1%.
Was the original model "correct" insofar that there was a 99% likelihood of something occurring (given the information it had at the time)? Or should it have "priced in" the fact that data may change substantially, and 99% was far too overconfident?
How are we supposed to interpret variability in model estimates? Do we throw up our hands and say "the data changed"? Or do we hold the models somewhat accountable, saying - no, you weren't "right at the time, given your data". If your estimates are changing so strongly, you are wrong. A 99% estimate that drops to 1% is simply, undeniably "unreliable."
In this case, we somewhat care about "model robustness", but how does this extrapolate to situations where the data changing _should in fact_ impact the model substantially?
I suspect the answer necessitates a deeper look into the nature of probability, risk and uncertainty.
Formally what you're describing is the bias-variance tradeoff.[1] You can assess this by looking at the conditioning of your model, which measures how sensitive it is to changes.[2] Roughly speaking, the condition number of an estimator (or generally, function) measures how large the change f(x) -> f(y) in the range for the change x -> y in the domain, where x and y are close. That will give you variance. If you try to minimize bias too much, you may overfit your model and it would exhibit high variance in cross validation. If you try to minimize variance, you may fail to capture relations in your underlying data, which would exhibit higher bias.
Practically speaking, for your specific example: if a relatively small change in the sample data resulted in the model adjusting its prediction to 99% from 1%, I would assume your model is severely overfit (I can't quantify exactly how small without more context, but let's agree it's small). From a meta Bayesian perspective, it would take quite a lot of further cross validation for me to drop that belief ;)
_____________________