I don’t agree with the “fairly small changes” part -- it looked like they were talking about 1% change in mean estimate and something like 25% change in variance. Those sound small, but they’re really not!
Or I guess another way to put it is, given the highly polarised two party system in the US, elections tend to be fairly balanced and small differences have an inflated impact on the results.
But whether you see the difference as small (1% difference in popular vote) or large (3x difference in chance of victory), the models are different, and one of them must be better, although it’s probably impossible to tell which it is!
Comparing models across multiple elections and calculating the Bayesian regret is one way to do it. The models get tweaked each election so this isn’t exact, but it could give a sense of the skill of each forecaster.
Does anyone have links for previous election forecasts from Morris and/or Gelman?