It’s fairly revealing of society’s general innumeracy, just as it was 4 years ago when Trump won.
It’s fairly revealing of society’s general innumeracy, just as it was 4 years ago when Trump won.
Should Biden ultimately be declared the victor, I’m concerned that the public will never receive the postmortem on what went wrong and how it happened again that it deserves.
Polls can only guess about turnout and the try to work backwards from there. The turnout estimates were wrong but not shockingly so, and as a consequence a lot of polls ended up having the result be at the far end of their margin of error. Nothing went wrong. Polling is hard. Get over this idea that you can have some sort of certainty regarding an election until we actually hold the election.
The DI poll is basically a push poll with questions like "Who do you believe is telling the truth about alleged Biden family corruption?”
https://democracyinstitute.org/poll-donald-trump-set-to-win-...
But, what the pollsters got crazy wrong, was Trumps polling numbers, by wide margins, far outside the margin of error.
I think they got it wrong this time too, and I think it all comes down to their methodology for actually getting a random sample of voters. Many polls are still married to live interviews and also to landline contact, and also tied to live interview polling - I think there is a partisan slant that they are not accounting for that includes a "propensity to answer a polling survey in the first place"
They're well aware of that slant, and try their damndest to correct for it. But apparently they just don't have a good model of how large or how volatile that propensity attribute is, and how it relates to political leanings.
I was polled this year. One of the first questions I was asked is if I was answering the call on a landline or on a cell phone. They weight those differently. Also of course my basic demographic info (wage, race, marital status, education level, home ownership status, etc.)
I think your whole comment is correct, but instead of not accounting for the propensity factor, they do... just badly :)
I have no idea what the solution is to this. Maybe there isn't one, and instead we can get a silver-lining effect: If people trust the results of polls less, they're hopefully less likely to think the results are preordained. More voter turnout?
In particular, Trump was extremely critical of polls and mail-in voting, and we have pretty much confirmed that his statements caused his supporters to avoid mail-in voting, so it's not too much of a leap to suggest the same thing happened with polls. Notably, this is a different effect from 2016, where IIRC Trump appealed to "non-likely" voters whose responses were inappropriately discounted.
It is a solution that has its own problems, but mandatory voting would solve this problem. The polling failures are primarily in trying to weight the sample that you manage to capture in a way that reflects the actual turnout on election day. If everyone is forced to vote then you eliminate the RV/LV problem and pollsters would only need to make sure their sample was representative of the population at large.
https://news.ycombinator.com/item?id=23548802
https://news.ycombinator.com/item?id=20141336
https://news.ycombinator.com/item?id=20117157
https://news.ycombinator.com/item?id=19751970
https://news.ycombinator.com/item?id=19743834
Are the Uncle Police OK with silent protest? http://v6y.net/unclepolice.mp4
When I grew up California and Virginia were red states, and I am sure people then also said money spent there was not going to move the needle and was a waste. When you have a billion+ to spend on races sometimes a little hope now pays off in a decade or two even if you lose this race.
CA was a Republican state from the late 19th century up until the 1990s.
When I hear the conversation about "the polls were wrong" it's never been "the predictions were outside the error margin" so much as "voter sentiment was not accurately measured in regions where voters were most likely to flip and why they would do so."
Sara Gideon was favored to win the Maine race in the polling, because there hadn't been a single poll showing Collins in the lead since July. She lost her race by 9 points.
It's also strange to see the region makes a difference in the poll error. The polls in Minnesota were basically spot-on, but in Wisconsin (demographically very similar), the polling average was Biden +8, with one ABC news poll showing him +17, the kind of outlier result you'd expect with a +8 average. He's gonna win there by ~1 percentage point.
There's something wrong with how a lot of these pollsters determine samples, or how they judge someone's likeliness to vote.
I wouldn't be surprised Trump actually wins.
Trump lost, get over it.
Lawsuits could definitely do something.
If there is fraud or misconduct, that can matter. Hell, the US spent two years investigating some Facebook ads and hackers.
But right now there's no clear indication that will in fact happen.
On the first Monday after the second Wednesday in December (Dec 14) the electors are going to meet in the respective state capitals, whereupon they will cast their votes and attach six copies of their vote to six Certificates of Ascertainment which will go to the president of the Senate (Chuck Grassley), two to the national archivist (David Ferriero), and then one to their secretary of state and one to the chief justice of whatever federal district their state it in. Voila! Now at 12:01 on the 20th of January anyone, even you, could deliver the oath of office to Joe Biden and swear him in as the 46th president.
Yes.
Biden's margins are hair thin.
Exactly, polling is very difficult, and getting even more so.
I was polled a few years ago by Gallup or Pew (or one of the other well known ones). The call was from an unknown number and I took it. No way I'd do that now with all the robocalls.
Then the polling results might ironically be more accurate if people believed in them less.
edit: To be clear, when polls overwhelmingly suggest a landslide, it suppresses votes from both sides. But a much higher proportion of the losing side will choose not to vote, thus inflating the gap.
Predicted vs Actual (FiveThirtyEight's averages vs NYTimes' current tally; Biden's margin is positive)
PA: +4.7% vs +.5% (-4.2%)
FL: +2.5 vs -3.4 (-5.9)
TX: -1.5 vs -5.9 (-4.4)
OH: -0.6 vs -8.1 (-7.5)
PA: +4.7 vs +0.5 (-4.2)
IA: -1.5 vs -8.2 (-6.7)
NC: +1.7 vs -1.4 (-3.1)
WI: +8.3 vs +0.6 (-7.7)
GA: +0.9 vs +0.1 (-.8)
MI: +8.0 vs +2.6 (-5.4)
AZ: +2.6 vs +0.6 (-2.0)
NV: +6.2 vs +2.0 (-4.2)
They pretty consistently overpredicted Biden's margin by about 4-7% in almost all of the swing states. Even if that's within, or close to, the margin of error; there's a systemic issue if it's happening in almost every state.
Edit: Added AZ and NV; and the difference
The polls were at least as wrong in 2016, even if that error was less significant this time (because the polled margin started out bigger).
Pollsters and poll-watchers downplaying this is not a good sign for future improvement.
EDIT: Overall though, I'm not sure that's a loss. People pay way more attention to polls than they should anyway.
At some point, reporting on polls when they have this bad of a track-record serves no journalistic purpose and just confuses the public. Like, the discussions people were having in the days before the election about Biden's strength barely resembles reality.
I respect FiveThirtyEight and their work, and I do think they're intellectually honest and generally speaking good at their jobs. But they can only be as good as their sources, and when the sources are this terrible, they shouldn't be reported on this, and not given the statistical and scientific sheen of authenticity 538 gives them. Like, that site is starting to have a net negative influence on the world, and when that is the case with a journalistic institution, what are you even doing?
But 538 chooses how to weight the polls. For example, they only gave Rasmussen (which was far closer on these) a C rating, preferring less accurate polls.
FWIW, their articles clearly lean left [1], IDK if that affects their analysis/forecasts.
[1] https://www.adfontesmedia.com/interactive-media-bias-chart-2... similar to WaPo
It's not like we do elections every month to test out their probability distribution against empirical data. The distribution collapses into a binary outcome at the end.
I have a dice. I claim the distribution is of equal outcome for each side. Well...we don't get to test the dice more than once. 1 sample size does not prove that the 538's predictions were right (or wrong).
Thanks for assuming I can't do math, no way to argue with someone but I am actually pretty bad at it. :-)
As to the question of why bother, it is because bad polling is better than no polling at all. Campaigns are now multi billion dollar enterprises managing tens of thousands of temporary employees for the creation of a product that will only be sold once and in 18+ months from when they start the process. Any data is better than nothing.
The fact that the public has become obsessed with polls is probably due to the ongoing nationalization of politics.
https://statmodeling.stat.columbia.edu/2020/11/04/dont-kid-y...
https://statmodeling.stat.columbia.edu/2020/11/07/what-would...
I don't know the specifics of their model, but probably they are claiming "with these polls, the probability of this outcome is...". The polls being consistently biased doesn't tell us much about 538s model. They said Biden would almost surely win and despite a massive surprise in favour of Trump, Biden won.
And even if Biden lost, 10% upsets in presidential are expected to happen once every 10 elections like this one.
If per-state error isn't normally distributed, that's evidence of bias, or bad polling.
538 has corrections for bias already. They seem to have worked in this instance - I repeat myself but: massive surprise, Biden still president.
The existence of bias doesn't invalidate their predictions. Everyone knows that polls can be badly off target in a biased way - that isn't a new phenomenon.
When they talk about X% chance of Y being president they should be optimising to the outcome, not the margins.
Occasional flaws in polling is understandable and tolerated. But when those misses repeatedly line up the same way, and are rather sizeable, that's evidence of either systematic flaws, or outright bias.
Perhaps you shouldn't be commenting on polling then.
No!
Assuming the per-state error would be normally distributed in some neutral world is making huge assumptions about the nature of the electorate, polling, and the correlations of errors between states, you can't do that! You would specifically /not/ expect per-state error to be evenly distributed because the nature of the error would have similar impacts on similar populations and there are similar populations of people that live in different states in differing numbers.
You should review the literature about the nature of the (fairly small) polling misses that impacted the swing states and thus disproportionately the outcome in the 2016 election. You will probably find it interesting.
There are unavoidable, expected, sampling errors which are, by definition, random. That's why valid, trusted polls calculate a confidence interval instead of a single discrete result.
Other types of "errors" -- election results that repeatedly fall outside the confidence interval, or are consistently on only one side of the mean -- only arise when the poll is flawed for some reason. Maybe you relied on landlines only, maybe you spoke with too many men, or too many young people, asked bad questions, miscalculated "likely voter," whatever. Accurate, valid, trusted polls don't have these flaws, the ONLY errors are small, random, expected sampling errors.
That is what each of these statistical models did, yes. And the actual outcomes fell into these confidence intervals.
> Other types of "errors" -- election results that repeatedly fall outside the confidence interval, or are consistently on only one side of the mean -- only arise when the poll is flawed for some reason.
Or the model was inaccurate. Perhaps the priors were too specific. Perhaps the data was missing, misrecorded, not tabulated properly, who knows. Again, the results fell within the CI of most models, the problem was simply that the result fell too close to the mean for most statisticians' comfort.
- Texas: 538 said Trump +1.0, actually won by 6
- Ohio: Trump +0.4, won by 8
- Iowa: Trump +1.4, won by 8
- Florida: Trump -2.5, won by 3
- Penn.: Trump -4.6, lost by 0.5
- Nevada: Trump -4.9, lost by 2
- Wisconsin: Trump -7.9, lost by 0.6
https://en.wikipedia.org/wiki/Statewide_opinion_polling_for_...
https://www.nytimes.com/interactive/2020/11/03/us/elections/...
The models and their data are public. The 538 model predicted an 80%CI of electoral votes for Biden as: 267-419, with the CI centered around 348.49 EVs. That means that Biden had an 80% chance of landing in the above confidence interval. Things seem to be shaking out to Biden winning with 297 EVs. Notice that this falls squarely within the CI of the model, but much further from the median of the CI than expected.
So yes, the results fell within the CI.
Drilling into Florida specifically (simply because I've been playing around with Florida's data), the 538 model predicts an 80%CI of Biden winning 47.55%-54.19% of the vote. Biden lost Florida, and received 47.8% of the vote. Again, note that this is on the left side of this CI but still within it. The 538 model was correct, the actual results just resided in its left tail.
Many political handicappers had predicted that the Democrats would pick up three to 15 seats, growing their 232-to-197 majority
https://www.washingtonpost.com/politics/house-races/2020/11/...
Entering Election Day, forecasters projected Democrats would gain House seats and challenge for the Senate majority.
https://www.cnbc.com/2020/11/05/2020-election-results-democr...
Most nonpartisan handicappers had long since predicted that Democrats were very likely to win the majority on November 3. "Democrats remain the clear favorites to take back the Senate with just days to go until Election Day," wrote the Cook Political Report's Senate editor Jessica Taylor on October 29.
https://www.cnn.com/2020/11/04/politics/2020-election-senate...
While I haven't checked each and every individual state, I'm pretty sure they all fell within the CI. Tail end yes, but within the CI.
> (BTW, who uses 80%? and not 90-95%?)
... The left edge of the 80% CI shows a Biden loss. The point was 538's model was not any more confident than that about a Biden win. So yeah, not the highest confidence.
> Deny it all you want, gaslight, cover your eyes, whatever -- but clear, convincing, overwhelming evidence of a systematic flaw or bias in the underlying polls is right there in front of you.
Posting a bunch of media articles doesn't prove anything. I'm not saying there isn't systemic bias here, but your argument is simply that you wanted the polls to be more accurate and you wanted the media to write better articles about uncertainty. There's no rigorous definition of "systemic bias" here that I can even try to prove through data, all you've done is post links. You seem to be more angry at the media coverage than the actual model, but that's not the same as the model being incorrect.
Anyway I think there's no more for us to gain here by talking. Personally, I never trust the media on anything even somewhat mathematical. They can't even get pop science right, how can they get something as important as an election statistical model correct.
The CI is due to sampling error, not model error. If the error of the estimate is due to sampling error, the estimate should be randomly distributed about true value. When the estimate is consistently biased in one direction, that's modelling error, which the CI does not capture.
What does "estimate" mean here? Gelman's model is a Bayesian one, and 538 uses a Markov Chain model. In these instances, what would the "estimate" be? In a frequentist model, yes, you come up with an ML (or MAP or such) estimate, and if the ML estimate is incorrect, then there probably is an issue with the model, but neither of these models use a single estimate. Bayesian methods are all about modelling a posterior, and so the CI is "just" finding which parts of the posterior centered around the median contain the area of your CI.
I'm not saying that there isn't model error or sampling error or both. I'm just saying we don't know what caused it yet.
Yes, they do. Because (among many other reasons) humans have a choice whether or not to respond, you can't do an ideal random sample subject to only sampling error for a poll. All polls have non-sampling error on top of sampling error, it is impossible not to.
Pollsters do that for continuously, and there were definite recalibrations in the wake of 2016.
OTOH, the conditions which produce non-sampling errors aren't static, and it's impossible to reliably even measure the aggregate of non-sampling error in any particular event (because sampling error exists, and while it's statistical distribution can be computed the actual error attributable to it in a by particular event can't be, so you never no how much actual error is due to non-sampling error much less any particular source of non-sampling error.)
This is false, if you think sampling errors are "by definition random" you don't understand polling.
Which is fine, just accept that you don't and dig into the literature or move on.
https://fivethirtyeight.com/features/the-real-story-of-2016/
https://fivethirtyeight.com/features/how-i-acted-like-a-pund...
https://fivethirtyeight.com/features/what-i-got-wrong-in-201...
https://fivethirtyeight.com/videos/what-harry-got-wrong-in-2...
I have noticed them getting a bit more defensive recently which is irritating, but I think they are honest. There's only so much "garbage in garbage out" they can compensate for. If this ends up at +74 EV, which is where it looks to be heading, it'll be "Z=-.6" prediction (not actually normally distributed, but it's the easy calculation) prediction, meaning 27% of the outcomes were more Trump. It's not great, but I appreciate them giving a realistic model.
Nate Silver still makes fun of Betting odds: https://twitter.com/NateSilver538/status/1325186985988935680
I don't see him posting how the Bettfair was way more accurate than his polls.
I also get that Nate Silver doesnt poll himself, but he aggreates polling information, assign grades and mixes it into his formula. I just have zero trust in it. It's not like we get to test his predictions often. 1 sample size every 4 years.
If I tell you that the odds of your dice roll being a 6 is only 16%, and you roll a 6, does that mean I failed to provide an accurate probability estimate?
The only way to judge the accuracy of a probabilistic model, is by quantifying the weighted error across numerous predictions. Not by looking at two yes/no outcomes. This is the exact point they are trying to get across, which you wrote off as annoying.
You do raise an interesting point about the betting markets - I too am very curious about whether they have a better track record than people like Nate Silver. Looking forward to someone analyzing their relative accuracy over a large sample of independent elections.
They also didn't seem very aware that mail votes (leaning Biden) would likely be counted after in person votes (leaning Trump), which 538 had been predicting would cause a temporary pro-Trump lean for ages.
The 538 model is explicitly based on the assumption that the state-polling errors are correlated to one another to some extent. Ie, if Trump performs better-than-expected in OH, he will likely perform better-than-expected in FL as well. This is why they rated Trump's chances as being 0.1, not 0.1^8. Given how polling works, this is exactly the right assumption to make.
Hence my earlier comment that if you want to evaluate the accuracy of 538 or any other model, you need to evaluate it across numerous different elections/events, over an extended period of time. Not a single day of elections in a single country.
The point is - 538’s input is a bunch of polls, their output is a prediction. Whether they do the polls or aggregate them is not relevant - they’re analysts whose job is to provide accurate estimates.
You can’t shift the blame on inaccurate underlying polls - Nate has time and again said, they look at many aspects in their estimates. Not just polls.
The polls did tighten at the end, but you can't just look at a snapshot (even near the end) to account for the trend. In all of these states the polls narrowed in some, but the final results were within MoE:
FL: https://www.realclearpolitics.com/epolls/2020/president/fl/f...
PA: https://www.realclearpolitics.com/epolls/2020/president/pa/p...
NV: https://www.realclearpolitics.com/epolls/2020/president/nv/n...
AZ: https://www.realclearpolitics.com/epolls/2020/president/az/a...
GA: https://www.realclearpolitics.com/epolls/2020/president/ga/g...
NC: https://www.realclearpolitics.com/epolls/2020/president/nc/n...
In Iowa the Des Moines Register poll had Trump at +7, which is within MoE, the rest were garbage: https://www.realclearpolitics.com/epolls/2020/president/ia/i...
In Ohio, none of polls were within MoE, so all of the polls there were garbage: https://www.realclearpolitics.com/epolls/2020/president/oh/o...
For the Senate I think North Carolina was wrong everywhere I looked, but 75/25 for Democratic control at 538 was the most wrong I think: https://fivethirtyeight.com/features/final-2020-senate-forec... https://www.realclearpolitics.com/epolls/2020/senate/2020_el... https://www.270towin.com/2020-senate-election/consensus-2020...
You can go state by state[1] to see which polls were more/less accurate within MoE compared to the final result and the trends in each state. 538 had a better chance of Biden winning by 400+! EVs [2], than the 306 he's likely to win with. It's this distribution that could have been better at accounting for severe polling errors in some states.
538 did do CYA posts[3], but here while bringing forward the error from 2016 in Ohio seems right, it still projects a win of 335+ EVs. Optimistic for Biden is the kind way of saying what the final 538 projection were, severely more wrong than individual polls is more accurate. If their distribution in [2] was better, I would be more willing to give them a pass.
[1] click on the state: https://www.realclearpolitics.com/epolls/2020/president/2020...
[2] "Every outcome in our simulations" https://projects.fivethirtyeight.com/2020-election-forecast/
[3] "What if polls are as wrong as 2016? Biden still wins" https://fivethirtyeight.com/features/trump-can-still-win-but... https://fivethirtyeight.com/features/im-here-to-remind-you-t...
I think this needs to be a serious topic of discussion. After this administration, “governing by polling firm” seems disingenuous at best and outright detrimental to all involved.
One of the most plausible ones mentioned by Nate was the fact that people who wfh weré more likely to be reached by pollsters this year, and may skew D.
There's definitely a demographic who voted for Trump, but aren't open about it for social reasons.
PA: 50.2% vs 49.7%
FL: 49.1% vs 47.8%
TX: 47.4% vs 46.3%
OH: 46.8% vs 45.2%
IA: 46.3% vs 44.9%
NC: 48.9% vs 48.6%
WI: 52.1% vs 49.5%
MI: 51.2% vs 50.5%
GA: 48.5% vs 49.3%
AZ: 48.7% vs 48.9%
NV: 49.7% vs 49.9%
Most are within ~1% or so (some for Biden, some for Trump) with some outliers being Wisconsin (-2.6%), Ohio (-1.6%) and Iowa (-1.4%) in Trump's favor still not being so far off (relative to margins of error).
I'm not sure what that "missing" bit in the polls really means (undecided?) but it seems the issue was that that bit of the electorate ended up going entirely for Trump in many places.
(edit: had the PA 538 # wrong)
I know FiveThirtyEight allocates undecideds evenly - if 8% respond undecided they assume 4% will break Biden and 4% will break Trump. In one of the podcasts they discussed this assumption and it's what they have found to be the most accurate historically.
We might be seeing a shift towards that no longer being true - maybe 75% break red and 25% break blue now.
Edit: Also, It looks like Jorgensen (Libertarian) significantly underperformed her polling (~3% > ~1%) in the first few states I spot checked. If those voters broke for Trump, that makes up about 2% of the error. That'd take a lot more in depth checking though.
Florida for examples may be off by 8 points, but 538 gave Trump a 1/3 chance of winning, so it was well within the error. What people don't realize is that the error on polls are pretty damn big, and elections in modern times have been extremely close. Like most of the swing states come down to under 100k votes / 1%. It's insane how close these rates are, and no polls will ever have any chance at predicting things like that.
The bigger mistake here though is looking at the mean and ignoring the error on those values. 538 is fairly conservative in their calculations and their error bars are pretty big, so almost all of these results are well within their error bars.
There is no reason that polls would be significantly different from the final vote results in most circumstances, if they are competently run.
Also, polling in most European countries doesn't go this far wrong, on almost all polls. So it's not like "accurate polling is impossible", it's just the polls in the US that seem to have some problem (though there have been some spectacular failures in other countries as well, which itself may lead to some thoughts on the value of polls in general).
Brexit and UK elections would like to have a word with you...
Polling everywhere sucks. Pick a country and chances are it either has no polling to speak of or else it has the same problems.
But I agree with your general point.
There is lots of reason.
If you could wave your magic wand and get a representative sample of voters in your poll you would be right, but you can't. Instead you do something like call people on a telephone, get a sample of people who both answer the telephone and who give answers you think indicate they are likely to vote. That sample isn't going to be at all representative of the voting population, so you take your best guess at what the actual voting population looks like, and how representative each of the people in your sample is of the voting population, and weight it accordingly.
Your best guess at the actual voting population is probably wrong. Your best guess at how representative each person in your sample is of the actual voting population is wrong. Your doing things like assuming that every <race> person with <education level> votes similarly regardless of whether or not they answer the telephone and respond to pollsters because you really just don't have a better option.
Polling is hard and the result are likely non-representative. Great! Then don’t display them to the public and have pollsters in interviews acting as if their insight is in any way valuable to most people. They’re either incompetently wrong or, worse, intentionally manipulative.
Just because something exists doesn’t mean it has value in the wrong context.
The forecasts you see on sites like 538 and the economist are even better, because it doesn't just communicate experts best guess (democrats win by +7), but experts best guess at how accurate that guess is too (something like democrats win by +7 +- 6, except in a lot more detail and nuance).
On HN the catchphrase for this is basically "don't let perfect be the enemy of good".
Note: I made up the numbers in this post.
If weather forecasters were wrong nearly 100% of the time, would they provide value? If a hurricane was approaching and they couldn’t predict probable paths within any margin of error, would they provide any value.
for instance, in your example, a hurricane's path directly affects citizens of cities that lie on that path—do they need to prepare to leave their homes, etc.? but is anyone actually basing their decision to vote or who to vote for on the polls in such a way that it significantly affects the outcome?
I live in Canada but am working for an American country. I am considering if/when I want to move to the US, they informed me about the likely outcome of this election, which influenced my plans a non trivial amount (not a huge amount either, covid dominates in the planning process, but a non-zero amount).
For many people it can influence how they vote. If you're in a state with a potential to change the outcome of the election you are more likely to vote "strategically", i.e. vote for one of the two most popular candidates instead of vote for a third party that you prefer. On the flip side if you're voting in a non-competitive race you are much more likely to vote for a candidate who is unlikely to win in order to better indicate your preferences (i.e. you should be more likely to vote libertarian/green/...).
For many people who wins an election affects there careers going forward, in the US in particular a large number of people are fired/hired based on the election. Even if you're not in that position your industry might get more or less government funding, if you're a government contractor the projects you are working on might or might not be in danger of getting cancelled, and so on and so forth. Having better information earlier makes it easier to plan your life.
The US election is one of the most significant worldwide events that happens every 4 years, the idea that being able to better predict it is not valuable is... insane.
i guess i was thinking more along the lines of, most of the time i personally don't think you should be choosing a presidential candidate based on predictions of who may or may not win. even some of your examples are more centered on the ultimate outcome of the election and how to plan for it; i really think people whose lives could be affected that drastically should be planning ahead for that situation anyway.
nonetheless, i agree they do have some value, and perhaps i should have clarified my line of thinking more clearly.
Do you expect the weather report for the next week to nail the temperature of every day next week to within 1 degree?
The pollsters can tell you if a hurricane is coming, they can tell you when the weather might shift significantly (albeit not be absolutely certain about which direction it is shifting) and they can give you a very good idea of what to expect even if they do not get the exact temp for each day correct.
What's more surprising to me is that he can do it consistently and that no one has seemingly appealed to these people before. Given that these are almost election winning demographics, it's odd that no one has spent more time on understanding these people.
But it is definitely within the margin of error for many individual polls, which often have a MoE > 2.5%.
This is some kind of non-trivial process problem.
For one, Andrew Gelman (professor of statistics and political science at Columbia University). Who knows more about the topic than anyone here in this HN discussion.
https://statmodeling.stat.columbia.edu/2020/11/04/dont-kid-y...
They discuss a lot of the things brought up here and go into details about the actual mechanics of calling people, etc. It's an easy listen.
I wish pollsters would move towards very very large digital samples. You can't get a 15 minute survey but head-to-head you can get a huge sample for not a lot of money. And don't try to weight it to what you THINK the turnout will be. Break out likely voter or not. Do not add any inference or opinion or 'math'
I've seen some good very large digital surveys. We do some - geared towards measuring ad recall/lift - and I find them valuable, especially a giant sample on only two recall + head to head questions
You can always ask for gender/age and weight it but that's part of what I see as the problem. so many 'traditional' surveys i see make fairly large adjustments, lots of 'looking back' at past turnout and over fitting based on personal bias. Especially when weighting from very small sub-sample xtab e.g. hispanics. im not a pollster and that's just my still fairly-insider / polling adjacent insight
Another problem with digital surveys is most of the firms do opt-in panels, like having people register to take surveys and get paid on mechanical turk. and then they weight from there. I think this is a problem. The surveys we do run inside of mobile ads, similar to Google / FB Brand Lift surveys they go after those who saw the ads or a truly random sample, not a small biased group of survey panel
Put simply, it was sampling bias. Pollsters screwed up. They either sampled the wrong voters, the wrong areas, mispredicted who would turn out, or voters did not accurately report who they planned to vote for. Garbage in, garbage out.
You can try to correct for sampling bias, but if you are blind to it, as your comment pushes more people to be, then you will fail.
Now you can argue that there's little utility if the error bars are so big, and you may have a point. The bigger issue here is that modern elections have been extremely close, and unfortunately we cannot get polling accurate enough.
[1]https://projects.fivethirtyeight.com/2020-election-forecast/
[2]https://fivethirtyeight.com/features/what-pollsters-have-cha...
Polling is really hard. There are fundamental problems with it that are impossible to fully solve. It's frankly amazing that they get as close as they do.
[1] https://poll.qu.edu/florida/release-detail?ReleaseID=3683
Accurate polls are important to running campaigns and to legislative positions. Selzer got her polls right in Iowa and she posited that she knew the local electorate better than other pollsters so she weighted correctly. Maybe Trafalgar Group has a better method or maybe it just weights GOP voters more heavily (we don’t know, it won’t disclose its methods). Or maybe Trump just breaks polling and the polls will be fine next time. Whatever the reason, I’d like to know, but I’m not sure we ever will.
It boils down to: a fraction of the electorate will vote for $bigotcandidate but would never ever admit it to anyone outside the voting booth, because they know it is wrong. They will tell people they will vote for $notbigotcandidate, but vote for him nonetheless.
This fraction is not insignificant and you cannot easily represent it in polls without guessing the effect. In my home province it eas consistently around 10 to 15% over a decade.
The publicly available polls are done for free by polling companies or by news organizations. Unfortunately some of these companies apparently have an agenda to push. So they skew their sample away from what the likely voters are to get the result they want for propaganda purposes.
I think that the rise of Nate Silver has ironically made this situation worse. He advocates taking an average of polls to get a true view of the situation. So if you are skewing you poll to get a certain outcome and you know the poll will be averaged, you need to warp your results even further.
Ok, disclaimer time - I don't have any inside knowledge or real evidence of any of the above. It does seem to fit however. I can't explain why it seems all the distorting is to make the Dems lead seem bigger than it is - that is, why aren't there pollers on the other side of the aisle doing the same thing?
So you'll have to decide for yourself if there is any merit to the above theory.
I've heard several hypotheses, but I don't have anyway to test them. They all seem to have merit.
I think we have to seriously consider that Trump just breaks things that otherwise would be working within reason.
This wasn’t some rare weird outcome. This was probably a standard deviation or so closer to Trump in a forecast. Perfectly normal thing to happen.
Alternate phrasings of identical questions asked in polls creates differences in outcome vastly larger than the effects they claim to detect. Somehow Nate Silver managed to get rich by selling synthetic CDOs of arbitrary polls (to people with his identical political outlook.)
The only real poll is to ask the same questions, in the same way, as the ballot sheet does. Ideally these should be asked as people are entering or leaving the voting booth, and people should be paid for their responses, in order not to bias yourself towards bored loudmouths with nothing to do. This would be good for accuracy, not for usefulness in anything but generating a suspicion of election fraud.
Instead, they're trying to figure out what percentage of black people go to the polls in Missouri, and pretending that the self-selected black voter has more in common politically with the general black phone answerer than the Missouri voter. Relying on that arbitrary assumption is the entire basis of their field.
edit: I haven't read https://site.pennpress.org/aha-2021/9780812250046/race-and-t... yet, but I heard an interview with the author around the time that it came out, and it's another hidden history of a field that has completely unearned legitimacy. "Political Science" might largely be considered an outgrowth of scientific racism.
A little scrutiny here is worthwhile. In 2016, for example, the polls were off in part because pollsters weighted the black demographic in proportion to their voting behavior in 2008 and 2012 (when a black candidate was running, boosting turnout). The actual 2016 black turnout looked more similar to 2004 or 2000 - other elections where a black candidate wasn't running. Though I haven't dived into the 2020 polling data, it wouldn't surprise me if a similar effect were at play.
It's also become noticeable harder to do polling. Once upon a time, everybody had a landline, so random dialing worked really well. Now, some people have more phones than others, and people are increasingly less willing to talk to pollsters, increasing the error rate.
These are all important things to talk about - when you dismiss it as "oh yeah, polls are never perfect," you prematurely shut down the conversation. Don't let perfect be the enemy of the good. The polls could be better and it's important to talk about how.
Besides the usual adjustments for likely voter models, we also do things like infer how people will vote in House and Senate elections based on Presidential polling, because we do a lot more Presidential polling than down ballot, because money.
So one early thing that seems to have happened this year (and again, we are still early in the process to be making final verdicts in "what happened" retrospectives) is that people split their ballot more than anticipated, which means that the inferences we were making about down ballot races based on Presidential performance were off.
So not quite a polling error, but an example of uncertainty that may not have been correctly taken into account in forecasting.
I think the complaint is that here the polls were not just imperfect. They were at best, misleading raising questions over methodology ( lessons, apparently, not learned in 2016 ), wishful thinking, and their usefulness, or, at worst, attempts at influencing desired income.
I can absolutely agree that polling is not an exact science, but now it is twice in a row that polling science has grossly miscalculated the mood of the nation.
I think it’s interesting too how some of the qualitative predictions (like Sabatos crystal ball) were pretty close to the outcome in the Pres race.
https://projects.fivethirtyeight.com/2020-election-forecast/...
https://www.predictit.org/markets/detail/4365/Which-party-wi...
You can also try to work the probability distribution from this betting market and I'm sure they were more favorable to R than he was:
https://www.predictit.org/markets/detail/6669/How-many-House...
That's a pretty big discrepancy regardless, did he ever give an explanation as to why he disagreed with betting markets so much? Seems like betting markets had a ~ one standard deviation pro Trump/R polling error baked into their pricing.
Because betting markets are kind of dumb? Nate Silver has uttered those words a few times, anyway.
https://i.imgur.com/O7UegtQ.jpg
We'll simply never know without 1000 realizations of this election but next cycle I'd like Nate Silver to bet money whenever he finds mispriced events/odds and show how much money he makes to strengthen the case that his model is an improvement.
Even if the pollsters are the most neutral nonjudgmental people on earth there will still be a fraction of the electorate that will (irrationally) shy away from telling them the truth.
This effect is certainly real.
I know those people in person. They are people who will tell even their spouse they vote the left wing candidate, but after the 6th beer they will tell you how they really want the strong guy in charge (note: Hitler is from Austria, voting the rightwing guy can still be a taboo for certain people).
So these people do exist. If it is just them who make up the difference — no idea, but they exist.
There certainly were some skeptics of the polls, but these skeptics were generally marginalized or dismissed.
The point is, most media outlets who reported poll results did so with the assumption that the poll results would reflect the actual results -- NOT that the polls results needed to be adjusted.
instead of asking directly who someone would vote for they asked a number of other questions first to warm them up etc.
They predicted way more Trump votes - and were summarily ridiculed by mainstream pollsters I think.
Further, the media could refuse to echo consistently unreliable poll results. Ultimately the media loses when people can't trust what's being reported.
That's why people love watching your elections so much. Or maybe also because of the fallout in their countries.
Much more interesting are the direct results predictions, especially per state where they matter most. Those are polls that give a particular result with a +- confidence interval. If the polls were to any extent competently done, you expect the result to be well within that margin of error.
The reality though has been that the results have been well outside the margin of error, sometimes multiple sigmas from the actual result - that is simply unacceptable for a poll. The only conclusion should be that polling data can be completely ignored, as it doesn't seem to have any clear relationship with the final results.
When we regularise these results like 538 or the Economist did, we get slightly better results, but they were still not particularly accurate.
This is a real problem, and not one that we should be ignoring. I gave the polling companies a pass after 2016, but that this happened again is extremely concerning.
> This is why sites like 538 were able to model the various outcomes with some degree of certainty and ONCE AGAIN manage to get results within the range of expected probabilities.
But they were super close to the edges of the intervals. Like the most expected number of electoral votes for Biden was 380-400, which is not what happened.
> You really seem to have a hard time understanding both polling and probability and in particular the relationship between sample size and error. I suggest a bit more research.
Please don't be dismissive of other people. Many people (including people who do stuff like this every day) were very surprised by the results. We definitely shouldn't tar and feather the pollsters, but I think it would be great if we could get an understanding of why the polls appear to have been systematically biased.
I think the problem is trying to make sense of outliers and uncertainty, and that, in an election with 150M+ variables, you will end up with a result that isn't quite what anyone predicted. In other words: it's not numeracy, it's epistemology.
‘We just can’t predict’.
I understand that the voting percentage we saw nationally was within the margin of error that 538 predicted, and that the specific configuration of electoral votes we saw was also one of the outcomes considered by the models.
The polls are never 100% accurate. The pollsters have to estimate how many of each type of voter will actually show up to vote and what the response rate for each type of voter was. It is not an exact science.
Yeah, so... numeromancy.
I am betting Facebook or Google would be able to better calculate the outcome than any polling firm. Bring in all the data they collect on each internet user, I am sure they can model the likelihood of voting and where the user is leaning.
As far as I can tell the most surprising result was Trump's popularity among Cuban Floridians.
When there is a major difference in voting patterns between the two political parties, then polls may not actually be inaccurate. In this case, Democrats chose mail-in voting and Republicans, by and large, chose in-person after being directed by the president.
Also, some percentage of mail-in ballots may not have reached their destination because of possible election interference through reduced post office efficiency--in particular, fewer mail sorting machines, lower hours, deprioritized mails, etc.
In short, the polls don't take into account election interference in the post office or unsuccessful delivery. If most people from both parties were voting in-person, maybe we might have had more parity between the polls and election results.
“What are the chances of a guy like you and a girl like me... ending up together?”
“Not good”
“Not good like one in a hundred?”
“I'd say more like one in a million.”
“So you're you’re saying there's a chance!”
I know enough about probability distributions and modeling to know you can improve your models to lower that uncertainty.
And sorry, you are the Lloyd Christmas in this one.
They would give you a probability distribution if they were truly random samples. But in reality, response rates are very low, and pollsters make up for this using demographic weighting and modeling - so the results don't actually represent any real population, but rather are extrapolated from people who respond onto a hypothetical population that the pollster thinks is likely to vote. This is why they are spectacularly and systematically off, more and more over time. It's not just variance.
... and a lot of assumptions about the supposed underlying probability spaces.
(I do not know why the forecasts are wrong, but I don't see why Parent is at all justified in his statement).