A good way to make them a bit more concrete to put some money on it, but even then it's hard to be sure. You can however use this to measure how good someone is in predicting. Take sports for example, by betting according to someone's prediction you can see how much money you earn (if their probabilities are 'true' you are almost guaranteed to earn money, eventually). But really in that case you're hoping that someone's ability to predict one sports match is indicative of their ability to predict the next, which is a good bet, but obviously doesn't hold for these kinds of miscellaneous rare events.
In this case you can imagine a betting system, and you can imagine some people might turn out to be better at it than others (though the relevant measure is not the one used in this article), so I suppose you can talk about someone's ability to predict these events. Though in the end the bookie always wins, so make of that what you will.
Of course this being probability theory all of these methods only work with high probability, so they may not work at all.
This is mine on Metaculus (mentioned in the article):
There's still a bit of noise at 205 predictions but you can see a pattern emerging!
Basically this is an exercise of garbage in and garbage out.
If one capital gets nuked the odds of another one suffering the same fate skyrocket.
I predict there's an 8% chance of one drive in my four-hard-drive array failing in the next year.
I predict there's an 8% chance the next UK election ends without any party holding a clear majority, leading to a labour-lib dem coalition government.
For the first prediction to be accurate, I just need to read the backblaze hard drive stats and multiply.
For the second prediction to be accurate, though? That depends on a lot of factors that are a lot harder to know. For example, would the lib dems be likely to enter a coalition, given how the last one went for them in 2015?
If you just want to use a different prediction method that fixes the problem and makes better predictions, looking at aggregate performance over many unrelated predictions does help. Compare how well different methods do, go with the best.
If you predict 100s of events, then use a method like Brier scores, you can get a good idea of how good you are. Even then it's probabilistic, but with enough samples it becomes incredibly unlikely you are just the worlds luckiest guesser.
> it becomes incredibly unlikely you are just the worlds luckiest guesser
And yet that is still more likely than the idea that you can predict the future.In any case, we are not talking about _you_ becoming the world's best guesser, we are talking about _somebody_ becoming the world's best guesser. That is going to be someone, so no need to be surprised when they emerge.
so I think this comes down to how the predictions were arrived. One way to do this is to ask individuals who both have similarly high predictive scores and bet on the same types of events to explain some of their past predictions, and if their methods are similar then you've learned a new predictive tool.
But yeah it's not possible to do this after just one event, you need some track record to be able to say that someone is statistically significantly better.
so its more about what price between 0 and 100 would you bet on that belief
at 20 you have a 500% gain if it resolves at 100
and this willingness to bet is telegraphed to the crowd, it doesn't really mean conviction but maps to the crowds aggregate understanding of the probability occurring or not
Most likely an application of the law of large numbers
The downside of bayesian probability that you point out is that isn't not straightforward to evaluate whether odds were correct or not. An example of this would be the 2016 US presidential elections; even the highest estimations of the odds of Trump beating Clinton were 30%, which after the fact was often cited as the predictions being "wrong". It's not easy to falsify this though, because a 30% chance doesn't mean something is _guaranteed_ not to happen; maybe it really was 70% likely not to happen, but we happened to end up in the 30%!
Frequentism doesn't suffer from this issue, but it also makes it impossible to talk about certain types of events (like presidential elections). In practice, bayesian probability gets used a lot when people want to talk about those sorts of events because there's not really any obvious alternative.
I have no problem with 538 saying that politician A has a 30% chance of winning because it’s a blended weighted average of several independent measurements of the ground truth (polling). There’s no independent measurements of the ground truth happening here. Indeed, things regarding war plans would be classified documents random people wouldn’t have to come up with a better estimate.
When someone has a track record of consistently predicting these "unique" events better than average, then clearly there must be some pattern there that they're picking up on and the events aren't as unique as one would think.
At the end of the day, someone who financially invests based on these predictions will eventually end up richer than someone who doesn't, whether you believe that should be possible or not.