How Can We Tell Which Election Forecasts Are True?
quantamagazine.org
quantamagazine.org
>FiveThirtyEight and the PEC both predicted the outcome of the 2012 presidential race with spectacular accuracy.
Markos Moulitsas and Drew Linzer (Daily Kos and Votamatic respectively) were both more accurate than Silver (https://www.dailykos.com/story/2012/11/14/1161465/-More-accu... http://rationality.org/2012/11/09/was-nate-silver-the-most-a...) AND they teamed up this year (http://elections.dailykos.com/app/elections/2016/office/pres...) AND The Upshot is showing their forecast as the most confident in its projections (http://www.nytimes.com/interactive/2016/upshot/presidential-...)
Why does Silver continue to be rewarded with praise of spectacular accuracy while Markos Moulitsas and Drew Linzer don't get mentioned? Hell, Silver and Wong are products OF Daily Kos. How can you write an entire article about forecasting accuracy and not even link to their site? At least The Upshot respects them enough to include them.
Apart from any measures of accuracy, he does deserve the recognition to have invented the genre. Back in '08 there were I believe two people doing it – him and Tanenbaum's http://www.electoral-vote.com
Each group creating predictions builds a mathematical model that simplifies reality into a set of numbers that necessarily have large uncertainty and error built in. The models are, by design, not intended to replicate reality. They are intended to be a smaller, simplified version, retaining as many of the relevant properties as possible, but with no illusions (one hopes) that they actually represent reality. (If they did, you'd expect them to make exact predictions, not percentage chances.)
Presumably these teams do their own math "right", so the only useful question is "which model is _closer_ to reality?" But there's no way to know that, except to test them an infinite number of times, and ooops can't do that either...
It's actually not simple at all! This has been one of the most contentious questions of statistics for over a century.
> All one has to do is imagine the election being repeated an infinity number of times under a stationary-non-changing distribution.
That's using a frequentist interpretation of probability, which is not the only interpretation. (Nor is it, in fact, the interpretation used by e.g. Nate Silver, who is one of the statisticians mentioned in the article).
(The frequentist interpretation, incidentally, has been shown to be self-contradictory, for reasons that involve continuous mathematics that I won't get into. In other words, you can construct situations that have nonsensical interpretations under the frequentist approach. That doesn't mean frequentist statistics aren't useful, but - like any model - the approach has limitations.)
It would be nice if all polls calculated a CI or probability distribution of their results (i know 538 does). But even 538 doesn't display their priors.
But this would confuse the public even more.. e... asking them to interpret a 95% chance to win with low confidence
It all boils down to how this random variable is defined.
It depends on what you mean by "repeated an infinite number of times".
First, there's a problem with the framing of that statement, because you can't literally repeat something an infinite number of times. You can repeat something indefinitely, and observe the behavior as the number of repetitions increases without bound, but that's not quite the same thing. (I realize this sounds pedantic, but if you want to understand the underlying interpretation of probability, it's an important distinction to be aware of).
So, let's say you asked, "if the election were repeated over and over again, without bound, what would the outcomes be?" And the answer to that is... well, it depends on who you ask. Setting aside the "fate vs. free will" philosophical question, different statisticians would answer that question differently, depending on which interpretation of probability they use. There are three main interpretations of probability, though these days, people pretty much only ever use two (frequentist and Bayesian). And you could construct either a frequentist or a Bayesian argument that the same candidate would always win, or that a different candidate could win each time.
Think of the election as a function - some unknown inputs go in, and an outcome goes out. What are the inputs to that function? (That's a rhetorical question - we're not really sure what makes every individual voter vote, and what disruptions could happen to nudge each person to vote a different way, or to not make it to the polls that day, or for their ballot to get lost in the mail, etc.). Congratulations, you have a random variable[0]! Now all you have to do is decompose it to determine the underlying components and the true distribution they draw from[1].
I realize that's not a very satisfying answer, but the underlying issue ("what does probability mean") is not a simple question, and statisticians have been debating it for over a century.
[0] which is, confusingly, neither random nor a variable.
[1] Which is itself controversial way of putting it, because some statisticians would argue that there is no "true" distribution, or that, if there is, we can never actually know it.
Edit: this is what "under a stationary distribution" means. It means holding everything that goes into the model, such as the polls, constant, while allowing everything else to vary, then you run the experiment many times. The easy way to do that is just let Nate Silver keep doing his thing for a few centuries.
[1] https://en.wikipedia.org/wiki/Probability_interpretations#Ph...
So maybe the issue is just in the language we use to talk about it, that we shouldn't even refer to it as forecasting at all.
So if Trump beats Clinton, that really embarrassing for 538, but if she wins the popular vote by less than 6.3%, that's also slightly off. If he wins Florida, that's slightly less off, but still a miss for 538, and so on.
IMO, the only way to really know which model is closer to reality is to wait until November 9 and find out. That's not helpful for deciding who to pay attention to in 2010, though.
I know a guy who correctly predicted the Obama/Romney popular vote outcome within 1% of each's actual take IIRC. As of now, he's got Clinton at 47% and Trump 34%, based on nothing more than personal outlook, hunches, and mulling over various polling data (he had 47% Clinton to 41% Trump just before the Conventions). Point being it might be just as likely an irrational human can point out human behavior (en masse, e.g. voting) on par with a statistical model, given some particular circumstances.
(or maybe, since there are predictions and results for every state and every senate race, it's actually possible to use those for an evaluation. Note that it's not as easy as scoring a "won by trump" in favor of a forecaster who was giving trump a 60% chance. You need a measure that expects your 60% predictions to be wrong 40% of the time, like the Brier score)
Eh, not really. Primarily elections (at all levels) are hard to predict for a number of reasons, but the presidential general election is about as straightforward as it gets.
Different forecasts vary with respect to the exact estimate and the degrees of confidence, but all the scientific, polls based approaches tend to predict the same range of outcomes. I wouldn't really call that "contradictory" for two models to give the same basic prediction but with slight differences in the errors and exact estimate.
However aren't most things in life a total waste of time, like posting on HN complaining about total waste of time?
according to the forecasters the election is already over, who's model is best is the only interesting undecided component left.
right now we have an election between a Blowhard Reality Star and a Crony Crook, where most of the population would prefer neither.
so combining the realities of the choice already made (barring an insane scandal) with an election between a douche and turd sandwich, statistical models are about the only semi-honest thing we can turn to.
do you really think this election is about "the issues" and not who is better at pandering? both of their campaigns are designed to tell people what they want to hear, and figure out the rest later.
democracy in general is a feedback loop between the press and politicians with the voting population as pawns/fodder. journalists and politicians have one shared concern: staying relevant and needed.
They were all statisticians he said. They were brought in to support decisions marketing made because as the Director said "people believe PhDs more than us. They make our decisions believable."
Same for using JD Power. Totally result driven. They got a lot of money to produce a result.
So to me anytime I see polls the only thing that matters is the money trail.
As far as elections go, we can see the accuracy in a month.
Luckily, there are people motivated by such questions. And even it may not make a difference to me except for it's entertainment value here, the impact statistics have had on our quality of life is undeniable.
There was so much invention in statistics and operations research in WW2 that you can make a credible case that it was the ally's decisive advantage. Since then, medical research comes to mind as field where it has been useful to have people who don't answer "dunno – why don't you wait a month and see" but "I think this may be possible."
I recall the event as a young marketer / techie because I always thought I wanted a PhD. But when I saw how far they were away from business decisions being only support I knew a PhD at this company was not for me.
There are people who earn a living with polls. This is a not-so-bad way to spend ones time. There are worse.
Experiment is the arbiter of truth.
"The reason for this is that pundits (the better ones, anyway) don’t just predict election outcomes, they also tell you how confident they are in each of their predictions. For example, Silver gave Obama a 50.3% chance of winning in Florida. That’s pretty damn close to 50/50 or “even odds.” So if Romney had won Florida, Silver would have been wrong, but only a little wrong. In contrast, Silver’s forecast was 92% confident that Rick Berg would win a Senate seat in North Dakota, but Berg lost. For that prediction, Silver was a lot wrong. Still, predictions with 92% confidence should be wrong 8% of the time (otherwise, that 92% confidence is underconfident), and Silver made a lot of correct predictions. So how can we tell who did best? We need a method that accounts not just for the predicted outcomes, but also for the confidence of each prediction. There are many ways you can score predictions to reward accuracy and punish arrogance, but most of them are gameable, for example by overstating one’s true beliefs. The methods which aren’t cheatable — where you score best if you are honest — are all called “proper scoring rules.” One of the most common proper scoring rules is the Brier score. A Brier score is simply a number between 0 and 1, and as with golf, a lower score is better."
http://rationality.org/2012/11/09/was-nate-silver-the-most-a...
"Also note that Wang & Ferguson got a better Brier score than Silver despite getting Florida (barely) wrong while Silver got Florida right. The Atlantic Wire gave Wang & Ferguson only a “Silver Star” for this reason, but our more detailed analysis shows that Wang & Ferguson probably should have gotten a “Gold Star.”"
No, Silver would not have been wrong at all, or at least we'll never know. He doesn't predict a winner; he gives probabilities of victory. It would have been perfectly consistent with Obama having a 50.3% chance of winning (or 63.5% or 32.8% or 99.9% or whatever probability you want to assign) and then for Romney to have won. Even after the election was over, a Florida victory for Romney would not have meant that Silver's pre-election probability assignment of 50.3% to Obama was wrong. That's not how probabilities work.
Think of sporting events. A huge underdog with only, say, a 2% chance of winning may beat the favorite. This does not mean that it had better than a 2% chance of winning before the event. Depending on the reason for the underdog's victory, it may not even mean that the underdog has better than a 2% chance of winning if they play again. One in a million events happen, and it doesn't mean that their chances of happening were greater than one in a million.
This is consistent with Silver's own explanations of what he's doing, and (I'm pretty sure ) he is careful to say he doesn't make "predictions". (Actually, not sure, it may be merely that he often clarifies that a different result than predicted does not make the probability assignment that prompted the prediction "wrong". The prediction and the probability analysis it's based on are different things.) I'm not saying that the Brier Score can't differentiate between better and worse probability assignments, but I am saying it's confusing the issue if it calls probability assignments "right" or "wrong". They may be "better" or "worse" than each other, but not "right" or "wrong". This may be using language in more technical manner, but it seems to me (and to Silver) to be an important distinction.
I think being concerned about this authors usage of the words "right and wrong" misses the point. The media will say "you were wrong" if you were on the wrong side of a 51/49% prediction, even if you shouldn't be penalized for it. Wang & Ferguson got dinged by The Atlantic Wire for being wrong, when they shouldnt have. The last quote in my post, supports your point. The person who is least wrong is the person with the lowest aggregate Brier score. The author was using right and wrong, because the articles they were referencing used those terms. The quote was written with the colloquial language of what it addressed.
The probability analysis that the prediction is based on does admit of degrees; one probability assignment may be better or worse than another. But saying a prediction is "only a little wrong" implicitly refers to the probability assignment it was based on; it confuses the concepts of "prediction" and "probability". The probability assignment was not necessarily "wrong" to assign a higher probability to the event that didn't occur. It may have been a bad probability assignment, but discerning that would involve considering factors other than the outcome of the event.
https://terrytao.wordpress.com/2016/06/01/how-to-assign-part...
> The "leaked memo" is a small test of readers' credulity. Written in apocalyptic terms, with strange capitalization ("the Liberal base is demoralized — but will become more enthusiastic the more Hillary is seen as an iconoclast who will overturn the last foundations of Western Culture"), it was uploaded to Scribd by a pro-Trump website with the confidence-inspiring name of RealTrueNews. For a short while, it was credited to Nate Silver; his name was blacked out, rendering it slightly less odd that someone who did not work for Monmouth was credited on a Monmouth polling memo.
https://www.washingtonpost.com/news/post-politics/wp/2016/09...
http://www.zerohedge.com/news/2016-10-11/first-post-debate-p...
Edit: Correction noted.
Post debate NBC/WSJ poll was conducted by group specifically paid by the Clinton campaign. Not exactly what you'd want for a major poll.
Can you dispute the damning evidence?
ie, more heat, less light.