It's that they are willing to bet more that they'll perform well on tests of objective scientific knowledge of related fields... and then don't.
It's that they are willing to bet more that they'll perform well on tests of objective scientific knowledge of related fields... and then don't.
The latter was at least removed from the results, but emphasizes the point nonetheless that they found it appropriate to include questions that had nothing to do with science. So you end up with results that affirm that people who disagree with consensus disagree with consensus.
This is an incorrect reading of the paper. Those are statements intended to determine whether a participant has anti-consensus beliefs, but do not occur on the objective measure we're talking about.
Then participants were asked to bet on their ability to pass a middle-school level science test. People with anti-consensus beliefs bet higher on their performance and did worse on the test compared to people without those anti-consensus beliefs.
These are the test questions they were asked:
1. True or false? The center of the earth is very hot: True
2. True or false? The continents have been moving their location for millions of years and will continue to move. True
3. True or false? The oxygen we breathe comes from plants: True
4. True or false? Antibiotics kills viruses as well as bacteria: False
5. True or false? All insects have eight legs: False
6. True or false? All radioactivity is man made: False
7. True or false? Men and women normally have the same number of chromosomes: True
8. True or false? Lasers work by focusing sound waves: False
9. True or false? Almost all food energy for living organisms comes originally from sunlight: True
10. True or false? Electrons are smaller than atoms: True
11. True or false? All plants and animals have DNA: True
12. True or false? Humans share a majority of their genes with chimpanzees: True
13. True or false? It is the father’s genes that decide whether the baby is a boy or a girl: True
14. True or false? Ordinary tomatoes do not have genes, whereas genetically modified tomatoes do: False
15. True or false? Sound moves faster than light. False
16. True or false? The North Pole is a sheet of ice that floats on the Arctic Ocean. True
17. True or false? The ozone layer absorbs most of the sun’s UVB radiation, but not UVA radiation. True
18. True or false? Nitrogen makes up most of the earth’s atmosphere. True.
19. True or false? Antibodies are proteins produced by the immune system. True
20. True or false? Pathology is the study of the human body. False
21. True or false? The skin is the largest organ of the human body. True
22. True or false? Ligaments connect muscles to bones. False
23. True or false? All mutations to a human’s or animal’s genes are unhealthy. False
24. True or false? Uranium is an element found in nature. True
25. True or false? Radioactive milk can be made safe by boiling it. False
26. True or false? The process of splitting uranium or plutonium atoms to create energy is called nuclear fission. True
27. True or false? Venus is the closest planet to the sun. False
28. True or false? It takes 24 hours for the earth to orbit the sun: False
29. True or false? A “Red Dwarf” is a kind of planet. False
30. True or false? The universe is expanding. True
31. True or false? Earth is the only place in the solar system where helium can be found. False
32. True or false? Gravity is the theory that serves as the foundation for modern biology. False.
33. True or false? The earliest humans lived at the same time as the dinosaurs. False
34. True or false? “Survival of the fittest” is a phrase used to describe how natural selection works. True
The questions for this ("objective knowledge" of COVID) begin on page 36 of the supplementary materials page. [1] I predictably offered the most extreme example, but there are various other questions that are also more about liminal consensus than "objective fact".
But, perhaps most importantly, the observed effect from this study was extremely small. And with only 34 questions, even a small number of questions that measure anti-consensus views, rather than scientific knowledge, is enough to create a false effect through what's basically an obfuscated tautology: people who hold anti-consensus views, hold anti-consensus views.
[1] - https://www.science.org/doi/suppl/10.1126/sciadv.abo0038/sup...
> But, perhaps most importantly, the observed effect from this study was extremely small.
The substudies that looked at those set of 34 questions above had an extremely significant result-- both statistical significance and magnitude of effect. I also don't see anything too controversial in that set of 34 questions. Do you?
The study is also filled with various issues. For instance they also removed everybody who completely agreed with the scientific consensus on studies 1-4 (see buried in page 44 on the materials and methods supplement) who undoubtedly also rated their objective knowledge highly. They drew participants from outlets like Amazon Mechanical Turk where people are paid peanuts to carry out menial tasks ($0.85 for this study). That's obviously not going to be a representative sample of society, let alone America. The paper claimed to have preregistered their study, yet didn't provide a link to such - which completely negates the point of preregistering, and so on.
Something I find quite troubling is how increasingly common studies of this sort are. This study is mostly junk, but it's enticing click bait junk that will also undoubtedly bump up this journal's impact fact as other's unquestioningly cite it more with the goal of generating interest even if in lieu of science. What exactly were the peer reviewers/editors doing? It seems that so long as one concludes the right thing in exciting enough fashion, what editorial process there is becomes quite deferent, a la Sokal.
Which is part of why they had subscores for related fields. Obviously we can't quiz about the actual controversy, because we disgaree on the facts.
> It doesn't even make any sense.
I think it's interesting that people who disagree with these consensus have, on average, lower overall literacy in the sciences-- both adjacent sciences and other science, but believe they have above average literacy and will outperform the average on tests.
Ironically, the numbers don't even support that if you do use the subscales. In all studies the overall effect from opposition was a fraction of a single point difference. In most studies opposition to the consensus in climate change was also predictive of a higher than average level of field-specific knowledge.
???? ad hominem? Where? I'm saying it's very difficult to quiz about the actual controversy itself, because then we get very close to the point of disagreement and risk confounding. The research explicitly mentions this aspect of design.
> They used a non-representative sample, removed the results of those who fully agreed with the consensus, and just generally did everything trying to massage their numbers into the conclusion they wanted to make.
I don't love use of Mechanical Turk for social sciences. It's still an interesting finding. Of course, more and higher quality research should be used to confirm the effect and gain additional nuance.
> In all studies the overall effect from opposition was a fraction of a single point difference.
The overall effect from opposition was a fraction of a single point difference per unit of opposition.
As for the replicability - again, this study arbitrarily removed people fully in line with the consensus. That makes it fairly safe to say that replication is a nonstarter. But I'd also add that another red flag for social science papers is when they collect a large number of variables that end up having no relevance to the published conclusion. Large numbers of variables is a key resource in p-hacking. And this paper was collecting all sorts of data that went completely unused and had nothing whatsoever to do with what they ultimately chose to publish.
> Estimate Std. Error df t value Pr(>|t|)
> (Intercept) 25.90592 0.43454 13.11332 59.617 <2e-16 **
> opposition -0.664790.06842 2130.90896 -9.717 <2e-16 **
Looks like each point of opposition was about 2/3rds of one more incorrect true/false question. So the people who were most opposed scored >3 questions worse on a 34 question test, on average.
> again, this study arbitrarily removed people fully in line with the consensus.
Median filtering and trimming saturated measures is common in research like this-- hopefully designed into the original protocol. I do agree it would be nice to see their preregistration.
But, weird things happen at the tails-- truncation effects, etc.
And while I'm aware trimming extreme outliers, regardless of the side they end up on, I'm unaware of any study entirely removing a segment from their sample, from one side only, which is critical to your entire hypothesis, and doing so without any explanation, let alone justification, whatsoever. One thing I'd observe is that the observed effect is small enough that if this culled group had a knowledge specific score below the mean, then it's likely that their entire conclusion would be invalid.
One other issue we have not discussed is the methodology not only resulted in a very non-representative sample, but also was indirectly testing something else. The surveys were done on the internet, and all of the general knowledge questions (besides the covid ones, which were just...) have answers which can be looked up in a matter of seconds.
I think they're both interesting. Broad overconfidence in performance in basic science combined with low performance is an interesting characteristic of a population. The fact that this is represented in those with contrarian views is interesting.
> So that translates to a fraction of a question difference in results.
The subscale finding, just shows that this same phenomenon appears to also extend to the subjects they are contrarian about. And it's about the same magnitude of effect, because there's very few questions on the subscale.
On the big test, a participant does about 2% worse per unit of disagreement. On the subscale, they do about 1.8% worse per unit of disagreement. The effect appears to have the same magnitude and is statistically significant in both cases.
This is an invitation to do studies on specific fields of disagreement with better samples and more questions on the subscale. But this early research casts a wide net across many types of anti-consensus view. Study 5 is a small, possibly flawed step in this direction.
I'd be particularly interested in a 2 variable analysis that attempts to model how one's overall objective performance and level of disagreement models objective performance in the specific field where they disagreed.
> outliers, ...
This is all moving goalposts when we were originally discussing the question set. I've already agreed that sociology and psychology research via Turk is problematic and faces many confounds and selection problems.
And perhaps a bigger issue is that there were literally zero questions that would do the same "trick" vice-versa for somebody who ideologically agrees with one of the topics, but otherwise lacked much knowledge. Testing this would have been quite easy by simply throwing in questions that sound like they fit the consensus but have an important aspect that makes them false. An example would be "Evolution is a process of improving a species which, over time, may lead one species to become an entirely new one." That is, of course, false - but somebody of low knowledge and high consensus agreement, would likely consider it to be true.
My personal hypothesis would be that extreme views tend to be associated with high confidence and relatively low knowledge. And this could be extremes of agreement or disagreement with a topic. As Betrand Russell put it, "The fundamental cause of the trouble is that in the modern world the stupid are cocksure while the intelligent are full of doubt."
If it's factually true, it's not testing ideology.
> but it seems much more likely to test ideology than scientific knowledge, in part because of how it is phrased
Is there anyone credible who disagrees with the statement here?
I think a bigger issue is that it may be testing not knowledge but careful, critical reading. I can score well on this test, but my flippant quick answers are not very accurate. I am not surprised those that are worse at critical reading might be more likely to have anti-consensus views.
> My personal hypothesis would be that extreme views tend to be associated with high confidence and relatively low knowledge.
That doesn't seem to be the case here; while we don't know the "1" full agreement results, the rest seems to be a pretty dang linear trend line.
Imagine that there was a secular Islamic expert. And he was quizzed on Islamic ideology in a similar fashion, such that he was obligated to imply belief in such, in his answers. Assuming he is being as accurate as he can in his responses, he would end up scoring as completely ignorant of Islam as he marked all beliefs as false.
I disagree there's a clear trend here, even without the 1s, because there were clear exceptions such as e.g. climate change where, even by their metrics, anti-consensus views were predictive of greater knowledge. So you end up having to, at a minimum, limit the stated claim to specific fields.
OK, well, here they expressed disagreement with stuff that is not just the scientific consensus, but the overwhelming agreement. If you want to phrase it as disagreeing rather than knowledge being lacking-- I think this is unhelpful. I do not think the nuance you're driving at here is worth complicating the questions or instructions.
> because there were clear exceptions such as e.g. climate change
No.
All 7 areas of controversy did worse on the entire test. It was not statistically significant for climate change and evolution. It was highly statistically significant for 4 of the remaining 5.
6 of the 7 did worse on their on subscales (significantly significant for 4 of these 6), with effect sizes of -.83, -.65, -.28, -.82, -.88, -.55; climate change had a non-statistically significant finding of an effect with slope .03. I do not consider this a counterexample.
You've been really advocating for study power and significance of results... when it favors your argument.
It is absolutely worth making the questions more clear, because this is the entire point of the study. Do people who do not agree with the consensus on something lack knowledge, or do simply not agree with said knowledge? Conflating the two in your own study does nothing but set yourself up to observe a tautology (those who disagree with a consensus, disagree with a consensus), which a cynic may interpret as not entirely accidental.
And no, I'm not referencing statistical power. If I haven't made myself clear, I think this study and their numbers, are both going to fall well into the endless black hole that is the replication crisis, which is especially pronounced in social psychology - which has an aggregate replication success rate in the 20s. What I am saying is that you can't make your broad claim based even what I suspect are deeply massaged figures.
Sorry, no. Overwhelming agreement along with readily self-observable characteristics is good enough. You can claim that the kidney is actually the largest organ of the human body, but everyone else agreeing is good enough-- especially when we can look at pictures from other people we trust, arrange to check a cadaver ourselves if we really care, etc.
I mean, I guess it's possible Big Skin (tm) has rigged all the measures and faked everything. /s
> What I am saying is that you can't make your broad claim based even what I suspect are deeply massaged figures.
One submeasure on one subpopulation having an outcome that does not support the claim but is also not inconsistent with that claim doesn't invalidate the claim.
You are talking about:
* A non-statistically significant finding
* Showing an opposite slope, but of tiny magnitude compared to the slopes in the other direction
* On one subpopulation
* On one submeasure.
As for your hypothesis, what would you propose would be a null hypothesis to test your hypothesis? It seems to me that it would be "Somebody who opposes a scientific consensus will score about the same in a test of general knowledge as somebody who supports it." And the climate change example confirms that null hypothesis, which would thus reject your hypothesis.
> And the climate change example confirms that null hypothesis, which would thus reject your hypothesis.
???! fails to reject the null hypothesis. This seems to be a rather large and fundamental error. You don't ever "confirm[ the] null hypothesis"-- and you certainly don't in a way that lets you exclude other explanations.
e.g. I flip a coin twice. It is heads once and tails once. I perform a statistical test and this outcome is consistent with random chance. In no way does this let me reject the idea that the coin is unfair.
(Indeed, doing a trial of 4 flips and getting 4 heads in a row is consistent with random chance; such a trial would not be likely to increase my belief that the coin is fair, even if it is not strong statistical evidence that the coin is unfair).
Back to our case: we have a whole lot of strong results in one direction, and one ambiguous near-zero result in one submeasure in one subpopulation. Perhaps that subpopulation really is different. It's an interesting thing to measure again. But if you measure enough things, you can expect to find some of these in your result set purely by chance.
My argument, in the most specific form, is that I do not believe this study measures, or was intended to measure, what the authors' claim it does. This issue is quite subtle, but I'm inclined to believe in malice over incompetence when these authors have spent years working these numbers [2]. The current study is a rehash of a study they did two years ago. The max assent group does indeed demonstrate reduced knowledge / increased self perceived knowledge. And there are demographic biases abounds. The desired trend varies by an order of magnitude in countries outside the US.
So this study took the most extreme sample from a subpopulation they were already familiar with the results of, stripped out the subpopulation of that subpopulation which challenged their hypothesis, added in a number of questions that were more likely to result in noise than anything else (the population of those taking surveys for pennies a piece on Amazon Turk, that can explain the difference between e.g. UVA and UVB radiation is certainly near zero so true/false just gives you noise and randomness), and then published. Another fun sample bias: Literally every single person who took fewer than 3 minutes to complete the survey (again based on survey 1) and bothered to fill in some (generally random) answers indicated max dissent. An interesting issue in itself, but a further problem in this study given relatively few people indicated max dissent, so they're measuring people literally racing through the survey for their $0.85 likely without reading it beyond looking for the "Please select #2 as the answer to this question." questions.
[1] - https://osf.io/x23c8/?view_only=29a92a9a707547149f210e5bf76a...
[2] - https://osf.io/t82j3/
I note you've completely moved on from the huge fallacy of a statistical claim that you made, to choose to circle back to hit one of the other topics we've already discussed earlier.
We all know about the Gish gallop. Pick one. You basically committed data interpretation murder in the grandparent comment.
If I fail to engage with every single point of yours, it doesn't make those points right. Especially when they're repeated after having already been addresses!
> means this is an extremely significant issue.
OK, so there's some questions on the whole test which might inadvertently measure religiosity-- if people don't realize they're supposed to respond with the scientific belief. (Of course, once they're betting on how they will score on that test, it's hard to understand what else they might believe would be the rubric for said test). Even so, this would hardly explains things showing up in the subscores with the same slopes, unless you think the bulk of the test is contaminated.
You're the one pulling the CSVs and running analyses-- is that question missed more by people who have anti-consensus views?
> An interesting issue in itself, but a further problem in this study given relatively few people indicated max dissent, so they're measuring people literally racing through the survey for their $0.85 likely without reading it beyond looking for the "Please select #2 as the answer to this question." questions.
Look, it's just easy to eyeball the graphs and see removing the '7's does nothing to the slope of the lines.
Looking briefly over the data, anti-consensus views do seem to correlate well with questions that were basically just proxies for measuring that dissent again. For instance one question was 'man was alive at the same time as dinosaurs' type thing. The correct answer on 1-6 is a 1. People who rejected the evolution consensus scored 4.6 on it, with a mode of 7. The mean of all respondents was 3. The score of people who indicated max opposition to e.g. nuclear power was 2.85. Those samples were my first and not cherry picked, beyond intentionally picking a group that would likely have very different biases (opposition to evolution vs opposition to nuclear) to compare against.
Imagine you simply framed the question slightly differently, but testing the exact same knowledge, "Most scientists do not believe humans were alive at the same time as the dinosaurs." I suspect you would have gotten radically different responses. And this issue (in various incarnations) is something I think that is absolutely destroying the 'soft sciences.' Quite small changes in experiment or survey can yield radically different results, and that can be exploited to "prove", more or less, whatever you want to.
p.s. The reason I moved on is because it was simply poor phrasing in a casual anon chat.
It wasn't poor phrasing. Being unable to reject the null hypothesis (on one subtest, even) doesn't mean there's no effect. If you test something 15 different ways, and you get a statistically significant effect with consistent slope on 12 of them, a non-statistically significant effect with a slope in the same direction on 2, and a non-statistically significant effect with a tiny slope in the opposite direction on 1... that looks pretty universal. It's an interesting question for followup research whether those 3 things are actually different in some way, or whether it was luck.
This is one of the many issues I have with social science - lack of falsifiability. Social science effectively lacks any degree of falsifiability simply because everything is based on vague probabilistic distributions which vary dramatically based on subtle nuances like asking an identical question two different ways, or testing on different populations in cases where the hypothesis is suggested to generalize - as here.
If there was any notion of falsification, it would all be false. In many ways it's quite analogous to astrology which, if you are unaware, was indeed considered a science, studied in universities for hundreds of years. It's end, somewhat ironically, came not from scientific objection to the fact it suffered the same issues as the social sciences today, but from the Church who deemed it heretical in claims of defacto divination.
This analogy doesn't fit our circumstance.
You take two photos each of 7 swans, a close-up and a wide-angle one.
For all of the 7 wide-angle photos, the swan is obviously white.
For the 7 close up-photos: you get 4 good close-up photos of white swan. Then there's 3 bad photos. 2 of them look probably white, and 1 looks like the swan is maybe dark grey.
All the evidence you have is consistent with all the swans being white. Looking to see if you can find those 3 swans and take a better close look at them might be a nice bit of followup.
This study is intentionally juking their numbers. They "learned" from their first study and removed every single sample that contradicted what they want to 'prove', magnified those that confirmed it, and even under this form of "science" - they still failed to show statistical significance for multiple categories.
It'd be like noticing there were far "dark" swans in Eastland and so going out of your way to make sure in your next safari you limit yourself to Westland. If you genuinely believed all swans were white, and wanted to test that hypothesis, you'd have done the exact opposite. There is only one conclusion to make.
I'm sorry; I disagree and I don't think you understand the statistics.
> Look at the last study.
This is another new goalpost. I am not willing to discuss further.
I was not shifting any goalposts, but simply referencing more evidence on the fundamental issue: that this study was intentionally biased and misleading with the goal of "proving", by any means necessary, a conclusion of which the authors had 0 interest in genuinely testing.
As I suspect we'll be wrapping up, I simply want to thank you for the interesting and lively discussion on this issue. It's rare two people can disagree on the internet for more than a few posts without things devolving into 'You're a poopy face.' 'No, you!'
I suspect we don't agree on what the "exact same quality of evidence" would be.
But the prior probability matters, too. A trial where a die looks slightly-biased-towards-6 (with weak significance) after 3 prior trials where it looked strongly-biased-towards-1 (with high significance) doesn't do much to change our mind about the previous findings.
Base rate fallacy, yada yada yada.