A piece of it:
https://twitter.com/NateSilver538/status/1294263127668924416...
A piece of it:
https://twitter.com/NateSilver538/status/1294263127668924416...
I'm with Nate Silver on this one, 72% and 89% are as close as 75% and 50% in terms of being double the likelihood for a Republican win.
We will likely never know which model is more accurate but given Nate Silvers history on these predictions I trust him a bit more than others.
In a postscript (P.S.) Andrew Gelman clarified that the forecast distributions are close and not the probabilities of winning [0].
> "I’m not saying that 72% and 89% are close in probabilistic terms. I agree that there’s a big difference between 8-1 odds and 3-1 odds. What I’m saying is that the large difference in these probabilities arises from small differences in the forecast distributions. Changing a forecast from 0.54 +/- 0.02 to 0.53 +/- 0.025 is highly consequential to the win probability, but as uncertainty distributions they are similar and they are hard to distinguish directly. You can get much different headline probabilities from similar underlying distributions."
[0]: https://statmodeling.stat.columbia.edu/2020/08/14/new-fiveth...
>In that sense, the best tip-off that the forecasts are different is they have Biden at 97% to win the popular vote, and have sometimes been as high as 98% (and 99% before they revised their model) whereas we are at 82%. You're getting up to 97-98%, you're getting VERY confident.
Silver made the same point about probability before the 2016 election, and got into a similar Twitter argument (https://www.politico.com/story/2016/11/nate-silver-huffingto...). One guess on who proved to be correct.
This isn't actually true. The probability of winning is based on aggregation of regional probabilities. The same national vote probability distribution can lead to very, very, different election probability distributions due to regional variation. The electoral college essentially guarantees that the national probability distribution is worthless for actually predicting who will be President.
If candidate B then wins, it does not mean that our analysis was "proved to be [in]correct". By itself, it doesn't actually say anything about the quality of our analysis. After all, we explicitly pointed out that this was a possibility, and it would be strange to argue "your analysis said this might happen, and then it did, so your analysis was incorrect". There's just not enough information to draw any conclusions.
How about 100 times? 10 times? 2 times?
At what point does evidence cease to 'say anything about the quality of our analysis'? The answer is never. Every datapoint can be used to update your priors according to bayesian statistics.
Especially for someone who titles them-self "Chief Scientist"
Aren't regional probabilities just conditions on the national distribution? Building the national Presidential election distribution from the conditioned distributions versus just using the sampled national distribution is bayesian statistics.
>The probability of winning is based on the popular support for a candidate
This was the argument you made. And I don't see how this is Bayesian vs frequency when all I am saying is that the same national popular vote distribution can have large variance in regional conditional distributions which leads to large variance in election outcome distribution due to the electoral college.
Sounds pretty bayesian to me.
But in 538's case they use the similar methods to forecast many individual contests like presidential primaries, presidential elections, senate elections, house elections, governor elections, club soccer, college football, March Madness men and women, MLB-NBA-NFL games and playoffs, etc.
Multiple-year track records over many events so you can compare their forecasts to actual results over time.
Get the data and compute your own error rate here: https://github.com/fivethirtyeight/checking-our-work-data
Most of the linked thread isn't about his own work or the Economist model, but about the false description of their forecasts as being pretty close, made as part of the (true) description that fairly small changes in intermediate results in the Economist model would lead to the same bottom line forecast as Silver's model.
Or I guess another way to put it is, given the highly polarised two party system in the US, elections tend to be fairly balanced and small differences have an inflated impact on the results.
But whether you see the difference as small (1% difference in popular vote) or large (3x difference in chance of victory), the models are different, and one of them must be better, although it’s probably impossible to tell which it is!
Comparing models across multiple elections and calculating the Bayesian regret is one way to do it. The models get tweaked each election so this isn’t exact, but it could give a sense of the skill of each forecaster.
Does anyone have links for previous election forecasts from Morris and/or Gelman?
As long as the same inputs aren’t fundamentally unavailable (and even then if the model has systematic handling of missing data, though the validity of that comparison is less clear) you can run the tweaked model on past elections (the main problem there is since those are probably the data used to generate them, it rewards overfitting; you might do better by running them against past elections with differing sets of random dropout of input data points [individual instances of particular polls, etc.] to mitigate that.)
Well, once he rose to popularity, 538 has been part of the mainstream media, affiliated with the NYTimes, then owned by ESPN, then transferred within the Disney Empire to ABC News. But, sure, it's pretty consistently been one of the better (not just election) statistical forecasting shops in the mainstream media.
We're saying, "he's better at predicting things than the mainstream media!" The issue is that the mainstream media's bar is so low, there is no bar. And the reason this guy is somewhat respected is because his bar is only slightly higher than rat shit, which seems amazing compared to what we're used to.
The guy sucks at actually modelling outcomes. Because everyone does. He just sucks very slightly less than everyone else. Maybe the issue isn't that Nate Silver is slightly better than average, maybe the issue is that we shouldn't be sharing this garbage on the news, since it's not reflective of reality.
It's like saying, "sure, Cuba is a human rights nightmare, but it's okay to trade and interact with them, cause at least it ain't the Congo!"
I think he does a better job at emphasizing the uncertainty while still showing that polls can be pretty reliable.
Why was it a fundamental failure? A 5% chance is one in twenty, it happens.
95% is “Obama vs a dog”. Maybe in another country. Every 20 elections, Obama is bitten by the dog and doesn’t make it.
That's what it was, don't forget how impossible it seemed at the time, it was a giant upset and shock.
> You think if Hillary went up against Trump, America would choose her 19 times out of 20?
If you are using "Hillary" and "Trump" figuratively about future elections with similarly matched candidates, then yes, see above.
Only in your filter bubble.
I wouldn't take his not being prepared as evidence of anything.
I'm not saying she ran a bad campaign (though she did). I'm talking about her "negatives". Decades of scandals. (Yeah, Trump had them too, but at a minimum it meant that Hillary couldn't use Trump's scandals against him. Also, Hillary's scandals got a lot more national coverage when they happened than Trump's did.) Benghazi. The email server (and with it, the impression that she thought that rules were for other people). The impression that she thought that she was owed the presidency, rather than having to earn it. The way the DNC chose her over Sanders, overruling the will of many of the primary voters. And on and on.
It wasn't obvious at the time, because much of the press was pro-Hillary. But she was a terrible candidate. I think if the Democrats had run anyone else, they probably would have won against Trump.
Biden isn't disliked as much as Hillary was. In that sense, he's a better candidate. (Trump is disliked as much as Hillary was, and then some. Perhaps that was your point.) But Biden doesn't generate any enthusiasm, except that he's "not Trump". (In fairness, some of the support for Trump was that he wasn't Hillary.)
1. Trump is no longer an unknown.
2. Biden is far more popular than Clinton was at her peak.
3. Biden is male (which is sadly very relevant in US elections).
4. We are no longer in the biggest economic expansion in US history. We may be in a historic recession.
5. There is a pandemic that has killed 60x the number of Americans that died on 9/11, and it is no under control in the US at all.
6. Kamala Harris is not like Tim Kaine.
7. Lots of anti-Trump voters are afraid to be complacent this time.
8. Suburban women, true independent voters (those who dislike both candidates), and older Americans are either breaking for Biden or have substantially switched to Democratic support after going for Trump in 2016.
A better question would be: is anything the same this time around?
4. But the biggest (longest, anyway) economic expansion was largely under Obama. That wasn't a reason to vote for Trump in 2016.
The rest of your list I agree with, with the possible exception of 8. I'm not sure that we can tell. I think the climate is so polarized, and the Democrats have the microphone to such a degree, that people who support Trump aren't willing to say in public that they do so. But you could in fact be right. We'll see.
For #4 the economic growth was under Obama but it meant that at the time of the election people were (mostly) fat and happy; the cost of taking a chance seemed to be lower and it was easier to tell yourself that 'this guy is a businessman, he could not fuck up this strong economy and might even do better.' Now people can see what a catastrophic mistake it was to make this assumption, but I can almost understand the reasoning. I think the point of claim #4 was basically that the most that any candidate could claim about the economy is that they would have recovered it better or done it faster -- another important claim that could be made about it would be that you would make sure that the 'right people' got more benefits from the rising economy and not those elites; the basic populist economic playbook.
So, 538 did actually change their error estimate during the 2016 campaign to better account for problems of correlation. From a purely mathematical standpoint that is kinda the wrong thing to do, but it is arguably better in line with what the readers expect an error estimate to be.
Either they spin it that they were perfectly correct in those 95% of cases where they get it right, or they spin it that they were the least wrong (and therefore the most correct) in the other 5% of cases where everyone gets it wrong.
There's a joke that you should always express 60% confidence in your predictions, since if the prediction pans out, you can claim to be right, but if it fails, you can bring attention to the "two times out of five wrong" part.
There is not enough data to back-test their model on a single election which happens every 4 years, so the claim that a 60% prediction is fundamentally very different from a 95% prediction is statistically dubious.
[ADDED: Or maybe something like that really is a few percent probability and you can end up with a 95% probability anyway. It just feels as if there's some upper limit to what you can measure using polls.]
Was it? What's the likely hood that someone who was polling as well as Clinton losing? 15% like The Upshot at the NY Times had? ~7% like with Sam Wang's model? Even 7% is around 1 in 14, not something shockingly improbable.
People often act like Silver's prediction was good because it gave Trump a higher probability of winning than many of the others, but that's not how probability works. If you say there's a 1/6 chance of rolling a die and getting a 2, and I say there's a 5/6 chance of rolling a 2, and we roll and get a 2, that doesn't mean that I'm correct. I don't think we've had enough rolls of elections resembling 2016 to really have a good grasp of where the percentages should be.
In general I question the value of assigning probabilities to election outcomes. This is especially true when you look at the probabilities a few months earlier - for instance, 538 had Clinton going from 49.9% on July 30, 2016 to 88.1% on October 18, 2016. Look at the probabilities they gave during the recent Democratic primaries, and they're also very bouncy. These probabilities lead people to believe that there's a much better understanding of the state of the race than what actually exists, to the point where I'd argue it edges up against pseudoscience.
Therefore I think it's fair to say that the weight an election forecast assigns to the actual winner is a direct indicator of the accuracy of its model. We aren't trying to guess at how a set of dice are weighted, knowing they'll only be thrown once—we're trying to get as close as we can to knowing who is going to vote and who they are going to vote for, and (absent some large disaster or upheaval) a misforecast will be largely attributable to systematic errors in our methodology.
who is "he" in your last sentence? ty
The page is still up, you don’t have to pull that number from memory. At the end, it was also nowhere near 95-98%. https://projects.fivethirtyeight.com/2016-election-forecast/
His 2016 model and he himself were much more predictive of the trump EC win, he repeatedly stated it was a possible outcome, something the vast majority of other forecasters completely missed.
These polls also never factor in things like "social acceptability of admitting that one voted for an unpopular candidate" or "groupthink among media organizations which aligns to their side of the isle." The entire polling fiasco should be interpreted as the limitations of quantitative data, as opposed to qualitative data.
Our final forecast, issued early Tuesday evening, had Trump with a 29 percent chance of winning the Electoral College.1 By comparison, other models tracked by The New York Times put Trump’s odds at: 15 percent, 8 percent, 2 percent and less than 1 percent. And betting markets put Trump’s chances at just 18 percent at midnight on Tuesday, when Dixville Notch, New Hampshire, cast its votes.
https://fivethirtyeight.com/features/why-fivethirtyeight-gav...
Good luck measuring those.
Consider this on the individual poller level. Every four years you get a new batch of pundit driven ideas about what "really" driving the voters. Suppose you add a question that's meant to magically reveal hidden voter preferences that are hidden by shame. What do the results of that question mean for the bottom line? Well, you're probably going to need multiple election cycle to find out. And by the time you do, the pundit have a new pile of bullshit for you to implement.
Now take it up a level. You've got a bunch of pollsters of caring predictive quality. Let them figure out what works best and include an estimate of their quality in how you process their results. I have no idea is pollster X's new question about cheese preferences is predictive, and honestly neither do they until a couple more cycles pass. All you can do is wight by past performance.
KISS wins in these situations.
Not sure why you are getting downvotes. This phenomenon has been academically researched and documented under the name “preference falsification” with many past examples. If anyone is interested the 1995 book “Private Truths, Public Lies” is on this research.
IMO LSE is a negative signal of someone's mathematical / statistical skill. I once interviewed someone that was doing a PhD in Math at LSE, and couldn't program or solve simple math / probability questions.
It's quite obvious why - if you want to study math in London, you go to Imperial. LSE isn't probably even a second choice.
..and that's only presidential elections. See https://projects.fivethirtyeight.com/checking-our-work/ for their quantitative reflections on accuracy. Looks pretty good.
Yeah, that’s what you’re supposed to do - use new evidence to update your beliefs.
There are also tools and websites out there that people can use to make predictions and track their own calibration over time. They're pretty fun in terms of encouraging one's own sense of rationality.