Human Me: I saw all this stuff in 2016 and Trump still beat the overwhelming odds. (Yes 16% chance of winning is still a chance, I get it... there's just so much deja vu with all these visualizations and articles about his campaign imploding).
Human Me: I saw all this stuff in 2016 and Trump still beat the overwhelming odds. (Yes 16% chance of winning is still a chance, I get it... there's just so much deja vu with all these visualizations and articles about his campaign imploding).
Based on what are you saying this?
I refuse to believe relativity until Einstein gives me a personal audience. Gravity, for that matter, is wholly unbelievable unless Galileo will sit with me under a tree. This is how the scientific method works: you must personally meet the scientist in order to believe anything they say. :)
But seriously, I have no idea what point you're trying to make here. I've met geneticists but have never met a virologist. Does that mean I should have more belief in the existence of DNA than in the germ theory?
Running the actual experiment requires a real election. Polls were more accurate in 2018 than in 2016, but we haven't yet had a presidential election with updated methodology. And presidential elections are different from midterms.
Again, of course you can construct a polling methodology that gives a pre-determined desired result (including the exact results of any particular election). This is called fitting a parameter.
And, again, the fact that you can tweak the weights on 2016 polling data so that the polls predict the right result is literally just a mathematical fact. And an uninteresting fact at that.
But you can't run a real experiment to test the the effectiveness of a new sampling method without actually holding an election.
I'm not even making any particular claim about the accuracy of polls this year.
I'm literally just saying that the fact that you can fit a model to give a historically correct answer on historical data that used one sampling method doesn't necessarily tell you the accuracy of a forecast that uses a different sampling method.
TBH, if we were talking about an advertising campaign, this wouldn't even need to be said out loud (or, if it did, the team member who didn't understand would be PIP'd ASAP).
> but the polls are not conducted by district but are conducted by "statewide random polling" which is about as useful in a country with an electoral system such as the US as the popular vote count. That is, not at all.
That's just false. Nearly every state allocates its electoral college votes based upon the state's popular vote. Hypothetical 100% accurate polling of the state's popular vote will absolutely tell you who gets the EC votes in at least 48 states. No?
Re: the distribution of the popular vote within each state, high-quality state polls DO weight for geography. But, again, you don't need to know
Also, the last person ANYONE should be listening to is 538/Nate Silver. People need to remember that before he was in politics, he used to be in sports betting but because his "models" were so terrible, he got pushed out of the industry. Nothings changed and he is still as wrong as ever.
If HN wants some high quality punditry, look at People's Pundit and their shows, Barnes Law (20 year election better wbo has never failed to make a profit in a cycle), and Cotto Gotfried. Stay away from the Nate Silvers and the Nate Cohns, terrible data and terrible forecasts.
That website's subtitle is "forecasting presidential elections since 1912". What does that even mean? Is Helmut 120+ years old? Or does he mean "fitting" instead of "forecasting"?
BTW: Helmut's model was jut wrong in 2016. Its prediction -- the actual statistical forecast the model actually made -- was that Trump had an 87%-chance of winning the popular vote. That was its forecast -- of the popular vote, not of the electoral college. Its electoral college forecast was then conditional on its popular vote forecast. The actual core quantitative prediction the model made was in fact wrong. Literally, the primary model got the right result by accident.
Of course, Helmut retroactively claims that because he got the EC right it doesn't matter. Which would be somewhat reasonable if:
a) that's how the model actually worked (it isn't -- the model's core prediction was about the popular vote and it was wrong), or
b) he didn't make exactly the OPPOSITE corrective about his model's performance whenever he talks about 2000, where he always stresses that it at least got the popular vote right even though it got the EC wrong.
Why does this matter?
First, Helmut's model isn't as good as he claims it is even on historical data. He moves the goalposts from year to year.
Second, Models that get the right answer for the wrong reason are usually over-fit to historical data. They should be taken with a grain of salt in years where their core feature set might behave differently that in previous years. (E.g., a year in which the model's main predictive feature -- primary results -- were cut short due to a 1-in-100-year pandemic).
Anyways, there are a lot of other contrarian/first-principles models that back-test well and predict a (D) win. Of course, no one cares about those this year. But they might next time around if the polling-based methodologies predict an (R) win.
FWIW, I tend to agree that the polling aggregation models are "meh" and should have way more uncertainty than they currently do. Mostly because a) I think turnout is much more correlated between swing states than those models typically assume and b) I don't think aggregating polls is actually all that useful.
Basically, everyone in election polling is making educated guesses and everyone in election polling over-sells their certainty.
Also, re: betting markets. Through a combination of bets on state-level results, you can construct positions that pay out if EITHER Democrats win the popular vote OR Republicans win the electoral college. The odds of Republicans winning the popular vote but losing the college are basically zero. I guess my point is just that there are lots of idiots betting on elections and you can trivially construct (small) sure-to-win positions.
Hispanic vote numbers are wrong this cycle (too high of support for Biden in the models). Might not matter, I'm still expecting a Biden victory.
This sounds like something a soviet comission would declare after an accident. What went wrong and how was it corrected? You needn't know, it has been corrected.
Opining without doing any research is lazy and undermines good conversation. Your exchange with GP literally amounts to a "uh huh / nuh uh".
Name-calling, FWIW, is even less productive.
>> you needn't know, it was corrected
There were long reports on what went wrong (and what didn't) in 2016 polling.
For example, https://www.aapor.org/Education-Resources/Reports/An-Evaluat...
More importantly, there are explanations from pollsters about what they changed and they speak openly about remaining possible sources of uncertainty/error. See e.g. the quotes in this article: https://www.newsweek.com/how-pollsters-changed-their-game-af...
All of this is a quick google search away, and anyone who has actually paid attention to polling knows exactly how polling has changed in the past 4 years. It's literally impossible to read anything on this topic without knowing that polling firms openly talk about methodology changes. Therefore, your original comment was either intentionally misleading or profoundly uninformed.
NB: You can absolutely agree or disagree with their methodological choices, especially around "shy Trump effect" and whether Trafalgar-like "what do your neighbors think?" questions make sense.
But even if you disagree with their choices, _it's plainly untrue to say their attitude is "you needn't know". To the contrary, they're quite open about how their polling methodology has changed._
Are you willing to defend with evidence and reason your initial claim that pollsters haven't explained what they have/haven't done to try and correct polling methodology? Or are we just going to name-call like we're 12 year olds on a playground/irrelevant partisan zombies in the internet comment section?
If you talked to people you would understand that what they optimize for is not really related to the caricatures of them, its not the opposite of what "your side" cares about just completely different.
When it is obvious that all the news comes from New York City and Los Angeles, and the world resyndicates that, while a large swath of America that actually matters cares about different things, it is easy to tell why the models fail.
You would have bet the farm on the undervalued bet on the prediction markets.
This is what's frustrating about sites like 538. They have Biden as an 87% chance to win. A betting site I frequent has Biden's winning odds as -180 (1.556 for Europeans). If they really believe the accuracy of their models, they should literally be betting the farm on a Biden win.
I mean, maybe they are; how would you know? (Though it would seem vaguely improper; not sure about the journalistic ethics take on this but it's at the very least not a great look.) Presumably the money a few 538 employees could bet wouldn't shift the betting markets very much.
Maybe it's simply not a good idea to analyse this in terms of chances of winning because we will witness just 1 event but the probability says that in 400 years we will have 16 Trump presidencies.
It's very counterintuitive, I don't know how to think of it. It's like saying that %99 of StartUps fail but that stat is meaningless to the individual startups(meaningful to govt & investors that will have to deal with 100s or more startups).
From subjects perspective, their own startup have %100 chance of success and Trump has %100 chance of winning.
Maybe these "chance of winning" statistics should be compiled by examining the roadblocks to the success.
If I told you "there's a 30% chance of rain tomorrow" and it rained, you probably wouldn't accuse me of being wrong (based on that one data point at least).
This year, there are some truly weird things happening. For example, Trump is losing ground along seniors while gaining ground among non-white voters. He is polling 10-20 points better among Hispanic people in Florida, for example, than in 2016. https://fivethirtyeight.com/features/trump-is-losing-ground-.... Trump is significantly outperforming Romney or McCain among non-whites, and significantly underperforming among seniors. Polls do lots of adjustments based on assumptions about e.g. voter turnout by race, and I don't know if pollsters have really gotten their arms around this situation.
The explanation of the "16% chance Trump wins" got a little unscientific in the media last time. To be precise, that's the probability that Trump wins based on the statistical distribution of the different polls and the margins of error within polls. (Roughly speaking, it's "measurement error.") It does not address the possibility that there are systematic methodological shortcomings in the polls. Both things happened in 2016: there was a significant chance of Trump winning anyway, but there was also systemic errors in assumptions regarding low-propensity voters, voters who haven't gone to college, and undecided voters. Pollsters have seemingly fixed those errors, but that doesn't mean there aren't other ones.
Yes, it does. FiveThirtyEight's modeling includes sources of uncertainty beyond just the margin of sampling error of the individual polls.
They've tallied up how much the polling consensus erred for past elections, and have found that while it is not consistent which direction the polling consensus will err for a give election cycle, it is possible to quantify what is a typical size of polling error. Or to put it another way, which methodological shortcomings matter changes from one election to the next, but that doesn't stop you from estimating how big an effect it is likely to have.
Assuming that the polls are likely to be collectively wrong by some amount is why the FiveThirtyEight forecast's probability distribution is so wide, with their 80% confidence interval for Trump's electoral college vote total currently spanning roughly 120 to 280.
(They also include uncertainty factors for how much the true state of the electorate's opinions could shift between now and election day, but those factors are fading out of the forecast model as election day approaches.)
I think 538 does take this into account in terms of the uncertainty, which plays into their probabilities. I know they certainly talk about this.
Not too surprising. Many Hispanic in Florida are Cuban, and republican. There are two Cuban US senators (Marco Rubio and Ted Cruz), they are both republican.
Hispanic people are not a single voting block, due to their diverse cultural backgrounds (black, white, Indigenous South American, mixed, rich, poor)
This is just going to be 2016 all over again where somehow polling = "the correct prediction" when none it is taking into consideration the voting record of the electoral college. Just a general consensus among a biased selection.
Can you clarify what you mean? Wisconsin never went for a Bush, and other than Trump's win in 2016 it's been voting for Democrats in presidential elections since Reagan. What electoral college record and red majority are you referring to?
And even if you reject all polling evidence, why is that one Trump win indicative of a solidifying shift toward Republicans, as opposed to expecting a reversion to the mean and Wisconsin going back to being a generally blue state?
I grant you plenty of people have sufficient trouble understanding probabilities that they would mistake 86% for a lock, though.
Really, Trump won by a hair's breadth. 538 shouldn't have given him a huge chance to win- he just barely scraped through. It's a genuinely hard problem when the popular vote is very close, since the electoral college throws a huge chaos wrench into everything.,
Just because you end up in the 9/100 future doesn't mean that the "real" probability was 50%. It's weak evidence that Bob's win scenarios mostly didn't involve blowouts.
It's like barely winning a sprint against Usain Bolt: maybe he had an injury midway through. Doesn't mean that he and I were equally likely to win.
That's the information we have, and it's real, so it's much more valid than any poll which poorly samples a few thousand people.
Maybe it was a fluke result. But that's unlikely because the other Midwestern states had similar results. But a fluke in which direction? You don't know, so it's even odds in both directions.
Interestingly, in 2016 they thought Trump was more likely than the betting markets did. Betfair had him on about 22%. So if they were betting they'd have made a profit.
This year the situation has reversed. Betfair has 35% for Trump, whereas 538 has him on only 12%. So they'd be betting on Biden.
So yeah, a lot of people seem convinced that since the pre-election polling showed Trump at a disadvantage last time, we'll see an identical surprise victory this time. I personally think that makes as much sense as leaving your umbrella at home on a day with a 75% chance of rain because the last time there was a similar forecast, no rain fell.
I get that 538 is an HN favorite, but they are too damn politically biased for their own good.
1) https://fivethirtyeight.com/features/how-i-acted-like-a-pund...
I don't see where you've said anything that comes close to supporting such a strong assertion.
2016 was FiveThirtyEight's third time forecasting a presidential election. Getting one of those three "wrong" when they gave themselves slightly less than one in three chances of getting the 2016 election "wrong" doesn't seem at all surprising. It sounds like you're demanding that FiveThirtyEight make their forecasts deliberately underconfident but still expecting them to call the correct winner every time.
How was that problematic?
https://projects.fivethirtyeight.com/2016-election-forecast/
A sibling comment brings up a batter, and I think that's a great analogy. A good batter has a batting average of, say 0.273. Nobody is shocked when they hit the ball!
538 gave a similar chance to Trump. He hit the ball.
> While the errors were nationwide, they were spread unevenly. The more whites without college degrees were in a state, the more Trump outperformed his FiveThirtyEight polls-only adjusted polling average,1 suggesting the polls underestimated his support with that group. And the bigger the lead we forecast for Trump, the more he outperformed his polls.2 In the average state won by Trump, the polls missed by an average of 7.4 percentage points (in either direction); in Clinton states, they missed by an average of 3.7 points. It’s typical for polls to miss in states that aren’t close, though. The most important concentration of polling errors was regional: Polls understated Trump’s margin by 4 points or more in a group of Midwestern states that he was expected to mostly lose but mostly won: Iowa, Ohio, Pennsylvania, Michigan, Wisconsin and Minnesota.
If the errors had been random rather than systematic, it would have been statistically quite unlikely for Trump to have swept all of those midwestern states.
Of course, the polls themselves also acknowledge the chance of a systemic error. They're reported similar to +/- 2% 19 times out of 20. They make no claim that the 19/20 of multiple polls are not correlated.
But their key insight is that polling errors are often correlated, and not just statistical sampling errors that can randomly offset in either direction.
Sure, but even the worst poll-based prediction outfit isn't just aggregating poll results and using the uncertainty based on sample size to determine the probabilities of different outcomes.
Pretty much all of them are using some model derived from the past relationships of polls to election results, which will, to a greater or lesser extent, capture systemic polling error that is not unique to the current cycle in the model.
Well, 538 model took into account historical evidence that election differences from polling averages between states are strongly correlated rather than independent; arguably, that's an effect that exists in part because such deviations are in part due to systemic errors (though the particular systemic errors may differ from cycle to cycle) in the polling, rather than pure sampling error.