A ragtag band of internet friends became the best at forecasting world events
vox.com
vox.com
I particularly liked the idea of base rates and averaging them. [1] I can see here that the idea of them is not to be on their own accurate (some of the methodologies individually are dumb), but do get an idea for the scales of numbers to talk about (the difference between a 5% event and a 10% event is very hard to notice as a human)
[1] "One was the rate at which provinces claimed by China (like Hong Kong, Macau, and Tibet) have eventually been absorbed, peacefully or by force; another was how often control of Taiwan has changed over the last few hundred years (twice; once when Japan took over from the Qing Empire in 1895 and once when the Chinese Nationalists did in 1945); the third base rate used Laplace’s rule. Laplace’s rule states that the probability of something that hasn’t happened before happening is 1 divided by N+2, where N is the number of times it hasn’t happened in the past. So the odds of the People’s Republic of China invading Taiwan this year is 1 divided by 75 (the number of years since 1949 when this has not happened) plus 2, or 1/77, or 1.3 percent.
Sempere averaged his three base rates to get his initial prediction: 8 percent"
Fwiw, when I've tried it, it has tended to suck in a whole lot of my time.
For me, other projects seem more meaningful. E.g., right now I'm setting up this: https://alert-team.org, another member is behind https://theaidigest.org/progress-and-dangers, a third works at https://forecastingresearch.org/, etc. Maybe we're an idealistic bunch, but also maybe we're making a bad judgment call, and still yet it's possible we couldn't cut it.
I mean, it makes me quite happy that y'all apparently use your skills for wholesome, humanitarian things instead of just focusing on making $$$.
fwiw, it appears you'd have a decent chance. You guys seem to have an extreme level of objectivity and rationality. There is a lot of domain knowledge required to bet on say, healthcare stocks or whatever, but maybe too much domain knowledge harms rather than hurts.
Naively, if any of the predictions could be made +EV accounting for time preference, the boffins' predictions should have been mirrored into the market to someone's profit.
I remain unconvinced that they're not merely very lucky, the same way that some hedge fund operators or day traders are lucky.
For climate change, there's already a lot of existing models out there calculated by much more rigorous metrics than "the Laplace rule" or just naively adding together averages.
We know that every species consumes all available resources until an external force, whether it's predators or environmental collapse checks its growth. Humans are no different - any time a poor person uses a paper straw, a rich one will fly their private jet. Muh backyard arguments on nuclear means we will continue to burn coal and gas as fast as possible, even while renewables help drive up more demand.
So I think it's safe to assume that emissions will continue growing.
[0]: https://cei.org/blog/wrong-again-50-years-of-failed-eco-poca...
FWIW, I'm very much NOT endorsing cei.org, nor saying climate change is not a big problem. Just saying that the track record of climate scientists is not great. Climate change is important enough to warrant the attention of people who can actually predict things.
Forecasting in general is a silly affair. As a data scientist I get all kinds of requests for all kinds of predictions, and my only real job is to part those requesting parties from their money.
Or in other words, if you want to say "this time is different" and deviate from the base rate, you actually need to look at the previous times to find out whether it really is different this time, and if so, by how much.
The party relies on divide and conquer of uniformed fools. But the fools are in the minority by now - and the divide and conquer is one sneaky trustfall decentralized communication platform away from never working again.
If the information is entirely new (ie, what impact does aliens taking over the Moon have on China/Taiwan) then there'll always be difficulty estimating that impact
Also, as the article notes, scholars like Tetlock have studied the the track record of scholars and experts and found them to be less accurate than this type of approach.
It’s a literary sleight of hand but useful to note as it undermines the entire premise of the article. That’s because these predictions are bunk because these techniques don’t work with the stock market which uses far more rigorous statistical methods for pricing (which happened 60s-90s with the rise of quants).
It can be hard to convince people with math alone.
This is mine on Metaculus (mentioned in the article):
There's still a bit of noise at 205 predictions but you can see a pattern emerging!
Basically this is an exercise of garbage in and garbage out.
If one capital gets nuked the odds of another one suffering the same fate skyrocket.
I predict there's an 8% chance of one drive in my four-hard-drive array failing in the next year.
I predict there's an 8% chance the next UK election ends without any party holding a clear majority, leading to a labour-lib dem coalition government.
For the first prediction to be accurate, I just need to read the backblaze hard drive stats and multiply.
For the second prediction to be accurate, though? That depends on a lot of factors that are a lot harder to know. For example, would the lib dems be likely to enter a coalition, given how the last one went for them in 2015?
If you just want to use a different prediction method that fixes the problem and makes better predictions, looking at aggregate performance over many unrelated predictions does help. Compare how well different methods do, go with the best.
A good way to make them a bit more concrete to put some money on it, but even then it's hard to be sure. You can however use this to measure how good someone is in predicting. Take sports for example, by betting according to someone's prediction you can see how much money you earn (if their probabilities are 'true' you are almost guaranteed to earn money, eventually). But really in that case you're hoping that someone's ability to predict one sports match is indicative of their ability to predict the next, which is a good bet, but obviously doesn't hold for these kinds of miscellaneous rare events.
In this case you can imagine a betting system, and you can imagine some people might turn out to be better at it than others (though the relevant measure is not the one used in this article), so I suppose you can talk about someone's ability to predict these events. Though in the end the bookie always wins, so make of that what you will.
Of course this being probability theory all of these methods only work with high probability, so they may not work at all.
so its more about what price between 0 and 100 would you bet on that belief
at 20 you have a 500% gain if it resolves at 100
and this willingness to bet is telegraphed to the crowd, it doesn't really mean conviction but maps to the crowds aggregate understanding of the probability occurring or not
If you predict 100s of events, then use a method like Brier scores, you can get a good idea of how good you are. Even then it's probabilistic, but with enough samples it becomes incredibly unlikely you are just the worlds luckiest guesser.
> it becomes incredibly unlikely you are just the worlds luckiest guesser
And yet that is still more likely than the idea that you can predict the future.In any case, we are not talking about _you_ becoming the world's best guesser, we are talking about _somebody_ becoming the world's best guesser. That is going to be someone, so no need to be surprised when they emerge.
so I think this comes down to how the predictions were arrived. One way to do this is to ask individuals who both have similarly high predictive scores and bet on the same types of events to explain some of their past predictions, and if their methods are similar then you've learned a new predictive tool.
Most likely an application of the law of large numbers
But yeah it's not possible to do this after just one event, you need some track record to be able to say that someone is statistically significantly better.
When someone has a track record of consistently predicting these "unique" events better than average, then clearly there must be some pattern there that they're picking up on and the events aren't as unique as one would think.
At the end of the day, someone who financially invests based on these predictions will eventually end up richer than someone who doesn't, whether you believe that should be possible or not.
The downside of bayesian probability that you point out is that isn't not straightforward to evaluate whether odds were correct or not. An example of this would be the 2016 US presidential elections; even the highest estimations of the odds of Trump beating Clinton were 30%, which after the fact was often cited as the predictions being "wrong". It's not easy to falsify this though, because a 30% chance doesn't mean something is _guaranteed_ not to happen; maybe it really was 70% likely not to happen, but we happened to end up in the 30%!
Frequentism doesn't suffer from this issue, but it also makes it impossible to talk about certain types of events (like presidential elections). In practice, bayesian probability gets used a lot when people want to talk about those sorts of events because there's not really any obvious alternative.
I have no problem with 538 saying that politician A has a 30% chance of winning because it’s a blended weighted average of several independent measurements of the ground truth (polling). There’s no independent measurements of the ground truth happening here. Indeed, things regarding war plans would be classified documents random people wouldn’t have to come up with a better estimate.
> The prediction got some press attention and earned rejoinders from nuclear experts like Peter Scoblic, who argued it significantly understated the risk of a nuclear exchange. It was a big moment for the group — but also an example of a prediction that’s very, very difficult to get right. The further you’re straying from the ordinary course of history (and a nuclear bomb going off in London would be straying very far), the harder this is.
Yup, the group got it right but predicting a rare event doesn't happen isn't that difficult, it's just notable because everyone was overly freaked out, particularly in the media due to self-repeated sensationalism. Peter Scoblic is correct that the risk is significantly understated because it's not correctly adjusting for the impact of the black swan event happening (e.g. if a nuclear explosion were to occur, you'd expect nuclear retaliations).
One day, 30 years later, the volcano's caldera began to spit magma and bubble with gas, so that morning the old hermit walked outside and changed his sign: "99.99% accurate"...
Looking at edge cases is good for sanity checking, so it's a good habit, and I commend you.
In your example, though, we can also consider the base rate of an event which hasn't happened in 99 years as 1/101, per Laplace's rule of succession. https://en.m.wikipedia.org/wiki/Rule_of_succession
Like many things, if you do it badly it doesn’t work. If you had that little data, you’d look at the rate across many similar volcanoes.
Focusing on base rate makes you more effective than others because people tend to only focus on the delta from the base rate. Tensions are “elevated”. Ok, elevated from what? People don’t actually ask that. They pull a number out of thin air and double it.
If you intentionally consider “what is the base rate?” and “how is this different from the base case?” you empirically end up with better results.
You’re also not married to the base rate. If you think a factor makes the odds 100x higher, go for it. You just have to say that explicitly.
- if you estimate the probability based on bad data, you get a bad answer.
- the base rate is a very simple model for the chance something happens – count similar events and divide by the number of potential similar events. One might describe it as an early prior before considering other information
- whether or not something (the eruption) happens in a year
- the probability one predicts for an event when considering more information. For example with the ‘pressure building’ model of the eruption you might decrease the probability immediately after an eruption and increase if it’s been a while (and increase a lot if smoke is coming out)
- sure it’s easy to be right predicting that unlikely events won’t happen. I think the claim of the OP is more that one may hope that the good prediction record transfers to the prediction of unlikely events
The sites which host these forecasting competitions correct for the bias against rare events through what's called "proper scoring" rules -- there's some specific maths to it, but the short version is that you're exponentially rewarded for being a correct contrarian and exponentially punished for being confidently wrong.
There are limits to that too, of course -- the folks in the article will "only" have made on the order of mid hundreds to low thousands of predictions, so roughly speaking, you can expect these people to be calibrated for 1% or 0.5% odds but probably not 0.1% odds.
Sparsity is a problem whenever you use data to predict or model something. Your example here is essentially subsampling only zeros from a sparse time series. The existence of sparsity isn't a new insight that invalidates everything. It's a challenge to be overcome by careful practice.
> when in reality your base rate in the year following an eruption is 0 with every passing year your probability of an eruption would increase & increase past 10% for every year past 100 that goes without explosion.
Sure, with large amounts of prior knowledge, you can do better than a naive base rate starting point. I'm sure that practitioners know this. Even in this contrived example, the base rate would have been a good first guess.
> predicting a rare event doesn't happen isn't that difficult
Doing it accurately (in the sense of having a low Brier score) is apparently difficult.
It seems unlikely that the U.K. is getting nuked alone, they have a bunch of nuclear armed allies. Even if you don’t believe in quantum immortality, if you are predicting ‘yes,’ you only get to collect your points in the ‘yes, and I survive, and so does the betting market’ case.
> something like war where pressures build up and war becomes more likely rather than less.
The likelihood of war between the UK and US, for example, has not steadily increased.
This is pretty much normal. It's just that it carries on under the surface, most people are unaware where "sources state that there is a 20% chance of .." comments from the military-strategic orbit come from.
These % figures aren't very meaningful. I personally think they are mis-read more than they are understood. I would not cross the road on a 1 in 10 chance of death. I undertake activity which at least formally has a 1 in 100 chance of negative outcome frequently, and most of us do 1 in 1000 or upward without thinking about it.
The Anti-Vax Phobics were rioting on a 1 in 100,000 risk issue.
I tend to think China-Taiwan by 2030 is in the 5% bucket as well for a number of reasons. Principally, the join over news stating the Chinese Military is massively corrupt, and there are generals weeding out over-claimed capacity, the lack of visible improvement in their naval fleet beyond one Aircraft carrier, and the massive bright red wave of blood which will stem from an opposed landing from sea, in modern warfare. Blood is a remarkable thing, in terms of it's influence on the people at large. China is not Russia, they are nothing like as fatalist about their sons and daughters. An occupation of Taiwan will incur massive loss in the generation which is the one-child policy outcome: Many chinese families will lose their investment in the future. Not to mention the missiles coming back the other way will make this invasion a weapon with two sides, one of which faces coastal investment on the mainland.
That and the highly internally facing quality of the rhetoric: Talking up Taiwan is seasonal politics as people jockey for control of the structures. It's not preparatory to invasion.
What is your evidence for this claim? How many people needlessly died due to Chinese communism? Maybe 50,000,000?
All sacrificed to an ideology. What has happened before can happen again.
Their fleet is expanding rapidly. Usually the discussion from experts is about how quickly China's naval power is growing. China happens to have two carriers now, and one more under construction.
Regarding aircraft carriers:
China doesn't need an aircraft carrier for Taiwan, which is ~100 miles off the mainland; air force bases on the mainland work much better.
Also, the relevance of aircraft carriers to high-end warfare is now in doubt: Accurate anti-ship missiles can hit carriers at much greater range than carrier planes can attack. China in particular has spent decades building a military that keeps US carriers too far from Taiwan to join the fight.
Generally, carriers may only be good for symbols of power (perhaps China's purpose for building a couple) and against less capable enemies - the Houthi's, for example, not China. They also enable you to have an airbase without local permission - if the US wants to bomb Somalia (as they currently are doing), and nobody nearby wants to be involved and let the US use their airbases, the US can put a carrier in international waters. It's a floating, mobile bit of domestic territory.
> That and the highly internally facing quality of the rhetoric: Talking up Taiwan is seasonal politics as people jockey for control of the structures. It's not preparatory to invasion.
Such talk can create a movement that takes on a life of its own; wars have been fought because someone inflamed the population and the leaders had no choice (or, being human, were inflamed themselves). One way to look at it is that such talk increases public support for war.
> The team had an answer, and it’s an answer that goes some way toward explaining why this group has managed to get so good at predicting the future.
Then a whole section and a half of stuff I don't care about to get to the answer:
> there was broad agreement that the base rate of war — between China and Taiwan or just between countries in general — is not very high
This is something I've been thinking deeply about for a while now. Curious to hear other people's thoughts.
... but I guess I'm not following how this question is related to the article? They're not claiming to use some kind of supernatural means to predict the future.
And even the latter only if you mainly wanted to predict events where said intervention will occur and be a major factor (e.g. predicting an earthquake doesn't really need to take into account that someone will be spared from dying there by divine intervention).
Probably because they don't think it exists.
No. Where does this idea come from? And certainly not in all denominations, including not in some major ones.
In fact (assuming you go but their stories) He often makes his present very clear and detectable.
That is about being superior and incomprehensible.
Something appears to have changed after all the stories were written down.
And given that each year sees ~50[1] new armed conflicts (for a total of 195 countries total) I'd posit that the base rate significantly exceeds their estimate of 22%.
Yes, of course, a war between China & Taiwan is not just like any war, and there are significant characteristics that influence the likelihood that aren't present in other wars, but at that point, talking of a base rate is just misleading.
And when you then "average" the "base rate" of hand picked special cases like "how often did control of Taiwan change", you're waving your hands. Wildly.
Which means the point isn't the idea of base rates per se, but choosing the right base rates - or answering the question "what is relevant here", and then analyzing from there.
The other success factor for forecasting orgs is choosing what you forecast. As such, I have questions about the actual forecasts made being locked behind a paywall: https://samotsvety.org/track-record/ (It doesn't mean they're not good at forecasting, but it's hard to judge how good)
[1] https://ourworldindata.org/conflict-measures-how-do-research...
> One was the rate at which provinces claimed by China (like Hong Kong, Macau, and Tibet) have eventually been absorbed, peacefully or by force; another was how often control of Taiwan has changed over the last few hundred years (twice; once when Japan took over from the Qing Empire in 1895 and once when the Chinese Nationalists did in 1945); the third base rate used Laplace’s rule. Laplace’s rule states that the probability of something that hasn’t happened before happening is 1 divided by N+2, where N is the number of times it hasn’t happened in the past. So the odds of the People’s Republic of China invading Taiwan this year is 1 divided by 75 (the number of years since 1949 when this has not happened) plus 2, or 1/77, or 1.3 percent.
The first rate is maybe since 1949 (or maybe 1841 or 1557?), the second is over an arbitrary number of centuries, and the last is also since 1949, but could be calculated wildly differently if rephrased as "years mainland Chinese government has control over Taiwan" and extended back to the Qing dynasty. Somehow these are compatible to do an average of.
Then "He nudged up to 12 percent." Not that estimated percent chances can ever really be "wrong".
We might think that this seems rather obvious, and that's probably true. Samotsvety are being compared to people like IC analysts, politicians and pundits. Not people exactly famous for their firm commitment to rationalism and accountability. If you compared their track record vs top hedge funds it might look a little different, but of course hedge funds don't try to estimate the probabilities of things like a nuke hitting London.
Also, if this all basically sounds like the economy is run like a mob casino... yes
Why assume that those would be people's highest priorities? If I had some superpower, I would not spend my time using it for either of those things.
Basically, if you have many people predicting something, just from randomness some will be more successful and some won't be.
The only questions are when and how and what will the rest of the world will do about it.
Forecasting your own life: https://fatebook.io/
Training your calibration: https://programs.clearerthinking.org/calibrate_your_judgment... https://www.quantifiedintuitions.org/calibration
Books to read: Superforecasting, Scout Mindset
Maybe a prediction market-based approach can weigh the beliefs of people about the event, but that doesn't mean it's a real probability.
Probability theory has given people the mistaken belief that everything is like a card or dice game. It isn't.
https://en.wikipedia.org/wiki/Von_Neumann%E2%80%93Morgenster...
https://en.wikipedia.org/wiki/Bayesian_inference
Maybe everything's not a card game, but you really don't want to find yourself in a situation that isn't.
Bayes'
Ukraine for example, prepared itself despite forecasts of low probability of an attack. Israel's military intelligence on the other hand knew the details of Hamas attack to the point and chose to ignore it despite high probability of the event.
As a forecaster: It's fun! It's an interesting way to learn about the world -- rather than gathering inert facts, you're forced to integrate them into a mental model, and then your model is tested for its validity empirically. It's also difficult and competitive, if you like that sort of thing.
As a consumer of forecasts: There's good research that prediction markets and other forecast aggregators are the best technology we have as a society for quantifying uncertainty. Not everyone will listen to them, just like not everyone eats their veggies, but c'est la vie.
I don't think there's too much risk of self-defeating forecasts (where a low forecast lulls decision makers into a false sense of security) - at least not yet. They're still pretty niche.
I know pretty surely that for example people in respective areas of finance are pretty interested in forecasts and will often act on them.
I know the US government has depended on forecasting in intelligence and defense for a long time built on the RAND legacy. I find it hard to believe most other sophisticated nations don’t also do similar forecasting and planning around forecasts given how influential these efforts in the US are.