Science Isn’t Broken
fivethirtyeight.com
fivethirtyeight.com
On the other hand, science writes want a headline "Something baffles scientists." Well, the scientists are probably baffled because it contradicts their intuition, which is a strong indicator that something is not really working. This something may be a honest mistake by the authors, a poorly understood experimental effect, outright fraud or it may be something worth reporting. The entire process of science writing is geared towards the papers which are most likely to be wrong.
It's a negative feedback loop where quality is hard for people outside the field to judge. So, it's just race to see who can lower quality more while still being published.
I don't know how it is outside astronomy (hello parent parent), but inside astronomy it's relatively well known that there are four journals that are likely to be scrutinously peer reviewed and which _will_ reject you and which _will_ force you to do further analyses. It is also extremely well known that some journals advertise "In print in two months or less!" Inside the field (so to the people who will give you the next contract), it's obvious where you publish.
In some cases, it _is_ very valuable to put the extra months of effort in to lay down a rock solid piece. For example, in CAS, you can make an extra 10,000 RMB for publishing in a low impact Indian or Chinese journal. You can make upwards of 100,000 RMB (a full year's salary) for publishing in one of the "big four" journals. And up to 3 years salary for getting a Nature or Science publication. This is money being invested by the Chinese government to intentionally combat the plague of falsified Chinese research. I don't know of any other country or system that does this. It is also remarkably easy to get tenure in China, which allows a lot of long-term type focus. But this is a tiny bubble for China as it attempts to brain drain the rest of the world and is not likely to continue for long.
It's a big problem that the media doesn't understand which are which. For example when everyone was raving about the "faster than light" neutrinos a few years ago the first thing _every single person_ in my building said was "The Italians set their clock wrong didn't they?" Half a joke, but dead serious that no one should even consider this until it showed up in a publication.
Another big problem is that only basing your research on verifying others research is suicide for an academic career sing you're intentionally doing _low impact_ science. The only people willing to fund that are people with agendas, ie. BP, Exxon, Coca Cola etc.
This is a common sentiment, and fortunately for scientists I don't think it's a very accurate description of the real problems we face. I really don't think that there are many good reasons to spend 10 years on one thing when there is a substantial risk of failure. Scientists spend 10 years on problems all the time! But it's either part time, or has measurable metrics for progress along the way that can be published. I mean, you try to go to the moon before Pluto, right? Science (as an institution) is really a lot more like Tetris than people think, except some of the blocks (publications) get moved or destroyed later. But if the blocks are smaller, they're much easier to fit together and build upon, and fault tolerant and so forth, so if they blow up it's less of a big deal. Smaller, more well-defined and robust studies are generally better.
To change analogies to paths instead of blocks, which is a better description of the decision making a scientist goes through (rather than building the edifice of human knowledge), it's a lot more like finding your way through a new city than choosing two paths in a snowstorm that will not encounter anything for 10 years' walk. At every fork, you have to decide where to go and there will be some reason you have made that choice. After some smaller period, you can if you choose report where you've gotten and why. Other people care! They will be happy to read about it, as long as your logic is pretty good, especially if you make nice observations along the way. You can publish this, and probably no one will steal it from you. Maybe someone will... but if you do it right, do it incrementally and let people know what you're up to, you'll at least have your batshit ideas published as unreviewed abstracts and folks will know who made the first progress.
Option A: Pick just one thing to look at say cardiovascular health.
Option B: Collect a lot of data, and then look for trends.
Option B has much stronger risks of false positives simply because your looking at more factors, but those false positives seem much more interesting which helps you get published. Worse you can't use any of that data for verification, so you now need to run another 10 year study. Upside, you got published, downside, your results are almost meaningless. Bigger upside, you then get to publish again in 10 years.
Well what about option C use approach B but publish sooner. Now your not only risking false positives, but also extrapolating from more limited data.
Hmm, guess what say nutritionists are going to pick... "Chocolate good/bad for you!"
A fairly standard definition of the p-value (from wikipedia) says: "In statistics, the p-value is a function of the observed sample results (a statistic) that is used for testing a statistical hypothesis. Before the test is performed, a threshold value is chosen, called the significance level of the test, traditionally 5% or 1% and denoted as α."
What this description is missing though is the crucial importance of the fact that the threshold value is chosen before doing the analysis. And moreover, that the entire analysis plan has been chosen before doing the analysis. Because what the p-value is really telling you is the probability that on repeating the experiment (and its accompanying analysis!) you would see a result as or more extreme than what you observed.
If your experiment comprises "try all of the combinations of variables to see what gives me the best answer", the p-value you compute would need to be some very fancy test that took that into account... and you would see your analysis as having much less statistical power.
For a simple example, look at a statistically rigorous method for dealing with multiple hypothesis testing when you plan it in advance: https://en.wikipedia.org/wiki/Bonferroni_correction.
Of course p-hacking is bad. The problem isn't frequentist statistics or p-values though, its scientists not understanding the statistics that they use. If you want to use a p-value to help make a decision about a hypothesis, you have to commit to your analysis plan in advance.
edit: Furthermore, p-values were designed to deal with experimental data. If you're doing an observational study, perhaps you should use statistical tools designed for that purpose.
To sum up: when you have people who have no idea what they're doing do statistics, they will do it badly.
Or rather, when you have people who have a good idea what they're doing, they will do a good job of getting the results they were looking for. And maybe that is why statistics are so popular in political debates.
The problem is not that people are not aware of the issue of testing multiple hypotheses. The problem is that (1) it's hard to say exactly what your hypothesis is before you've even looked at the data, and (2) it's hard to determine if people are choosing parameters for p-hacking or simply making choices based on their best judgement.
>Furthermore, p-values were designed to deal with experimental data. If you're doing an observational study, perhaps you should use statistical tools designed for that purpose.
This is simply wrong. p-values are equally relevant in both cases. E.g. I can use p-values to reject the hypothesis that consuming saturated fats is uncorrelated with weight gain, amongst the general population. It sounds like you are reaching beyond your actual expertise in statistics.
Here's what doesn't add up for me: this implies that for many configurations, p-values are invalid. But what was it about preselecting the experimental conditions that really makes them any different? What makes initially chosen hypotheses of higher quality than iteratively discovered, "p-hacked" hypotheses?
The need for careful selection of a limited set of variable combinations in advance seems symptomatic that the test being employed is not robust. I'm not convinced that even limited application of p-value significance testing is actually valid.
Imagine you are a scientist and want to find out whether the hypothesis that men are on average taller than women is true. If you could just exactly measure and take the average of the entire male and female populations you wouldn't need a hypothesis test.
Since you can't, you can do an experiment where you take a random sample of men and women. Now, you can do the average in the same way, but you need something to help figure out whether to trust the results. That's where the p-value comes in. The reason you need to be careful to select hypotheses in advance is because in order for statistics to help you account for error you need the noise in the data to be uncorrelated with the result you are trying to assess---which it won't be if you chose the result because it was the one that looked best after you account for the noise.
Bob has a class of 100 undergraduate students. He tells them to go and run one experiment each as part of their final year class. Assuming all the data in all the experiments is completely random, what is the probability that at least one student will nonetheless find a "statistically significant effect" (p < 0.05) that can then be written up and get published.
rough answer: 99.4%
p < 0.05 seems to be chosen as that's a barrier that is not too harsh on the individual researcher (at very high significance levels it would be hard for anyone to run a powerful enough study in order to get published and keep their research job), but p-values are essentially defeated as a filter for genuine effects by the sheer number of experiments being run around the world.
In some senses, the answer here is simply to stop brandishing science as a social cudgel ("X is science; how dare you not believe it" or more often "how dare you employ/fund someone who does not believe it, they must be sacked") which puts it on a pedestal of being canonical truth all the time.
It is only the fact that science has recently been used more and more as a political stick to beat opponents with (and no, not just in climate and evolution, but right down to things like the best way to teach reading, or whether schools should be regulated or independent) that has meant that people think "science is broken" whenever it turns out to have got a wrong result. No, it's just expected to get it wrong fairly often. And in practice the most common "self-correction" process in science isn't a repeated experiment and published retraction, but just academics reading the paper, thinking "this one's garbage", and not basing their research from it.
Publication is not strictly a test of truth -- just a test of methodology and analysis. The test of truth always occurs in the mind of the reader.
That's incorrect. If you perform the exact same experiment twice, the chances are exactly 50% that the first result is more extreme than the second one (neglecting equal outcomes).
This illustrates nicely how hard it is to properly explain the p-value. Things would be more intuitive if people would report confidence intervals. I prefer "x is with a likelihood of 95% between a and b" a thousand times over "we measured x=c and reject the null hypothesis with 95% probability".
The chance is 50% if you condition only on the information you have before doing either experiment. But once you've done the first experiment, the chance of a more extreme result given what happened the first time may be much more or much less.
(Extreme example: Your experiment consists of rolling ten ordinary 6-sided dice. The null hypothesis is that they're fair dice, fairly rolled, in which case you expect a total not too different from 35 pips. All the dice come up 6. It is not now true that if you run the experiment again, you're as likely to get a more extreme result as you are to get a less extreme one!)
> confidence intervals [...] "x is with a likelihood of 95% between a and b"
But that isn't what a confidence interval means! A 95% confidence interval [a,b] means "If we ran the experiment lots of times, using the same method of computing the interval [a,b] each time, then in 95% of runs (in the long run) the true value would be in the interval [a,b] obtained on that run".
(What you described is what Bayesians call a "credible interval". Of course that interval depends on your prior.)
So 95% of the runs produce an interval that includes the true value. Doesn't that imply that when doing one run, there is a 95% chance of the resulting interval including the true value? I.e. the probability is 95% that the true value is in the interval?
The situation is similar to the one we discussed earlier where you do the same experiment twice. Before you do the experiment, your estimate of Pr(true value is in confidence interval) is 95%. But after you do the experiment, you have extra information: you know how the experiment turned out. And this may lead you to a different estimate of Pr(true value is in the confidence interval).
The equal probability you are describing is for the p-value itself being more extreme, not the result.
Why batman? Because he had to fight cops just as often as baddies. Not that the current system is necessarily corrupt, but that our current system is almost as blind when it comes to recognition of good science. Our current system is a bit, well, dumb.
[1] It's really hard to end a fight in 4 pages.
So I don't think this really fits dluan's description very well.
> some people have begun to ask: “Is science broken?” > I’ve spent many months asking dozens of scientists this question, and the answer I’ve found is a resounding no.
News at 11: cognitive dissonance is a thing.
Things can certainly improve, and they usually do, perhaps a bit gradually. Take a long view. Science is funded by the public and not the church these days for example. But 'broken'? Not really.
Edit: and non-scientists are very quick to think they know the problems and offer solutions. Working scientists are actively discussing the problems and trying things out all the time. These are not the dumbest people you'll meet. The current model is a bit like democracy - it's the worst system apart from all the other things we've ever tried.
To be clear: I'm not an apologist and I'm not sweeping troubles under the rug. It's just that it is complicated, we're working on it, and it is currently much too productive to be called broken. We have to be careful not to break it.
Are scientific papers meant to be textbook quality results, with absolutely no errors? Or are they meant to speed the pace of discovery of true facts? Because I would argue that they are meant to be the latter, that they are effective at that, and that trying to make them more like the former would slow the pace of discovery.
For instance, studies that compute the gyromagnetic ratio of the electron (to 10 decimal places no less) and then compare that against the experimentally obtained value would be classified as "high rigor". Studies that assess whether watermelon is linked to heart disease would be "low rigor".
The question now is how to come up with a rigorous classification system...
I'd love this to extend to pursuits outside of science too, ranging from "mathematical result verified with multiple proof assistants" to "this thing I made up and argued persuasively". Unfortunately, most of my arguments for why Haskell is the best language would fall dangerously close to that end of the spectrum :P.
You mean like, say, measuring the elementary charge of an electron to high precision? ;) Richard Feynman has something to say about the rigor of those experiments...
> We have learned a lot from experience about how to handle some of the ways we fool ourselves. One example: Millikan measured the charge on an electron by an experiment with falling oil drops, and got an answer which we now know not to be quite right. It's a little bit off because he had the incorrect value for the viscosity of air. It's interesting to look at the history of measurements of the charge of an electron, after Millikan. If you plot them as a function of time, you find that one is a little bit bigger than Millikan's, and the next one's a little bit bigger than that, and the next one's a little bit bigger than that, until finally they settle down to a number which is higher.
> Why didn't they discover the new number was higher right away? It's a thing that scientists are ashamed of—this history—because it's apparent that people did things like this: When they got a number that was too high above Millikan's, they thought something must be wrong—and they would look for and find a reason why something might be wrong. When they got a number close to Millikan's value they didn't look so hard. And so they eliminated the numbers that were too far off, and did other things like that...
[from "Surely You're Joking, Mr. Feynman!" via Wikipedia–I know this isn't exactly what you were talking about in your comment, but I thought it makes for an interesting discussion regardless]
Interesting paper on this here: http://arxiv.org/pdf/physics/0508199v1.pdf and also another example here with the Hubble constant here: http://www.pnas.org/content/101/1/8/F2.expansion.html
Even in Mullikan's case, the paper I linked seems to suggest he might have intentionally omitted data that didn't fit his narrative (I only know what I've read in that paper, so I wouldn't like to make any stronger claims than that). Whether he was 'right' to do that is an interesting question. Clearly the oil drop experiment was ground-breaking, and the results hugely important, but omitting inconvenient results is clearly a huge ethical violation. However, his intuition was correct, and he was observing a real effect, and the measurements he kept were good. So who can say? Science seems to be about asking the right questions as much as finding the right answers. Sometimes good science is just to get smart people talking about the "right" things, and the correct answers come later.
Perhaps peripheral, but another favourite example of mine is the Lennard-Jones potential for modelling e.g. noble gases. It has an attractive term and a repulsive term. The attractive term is distance to the power 6 and based on good physical reasoning; the repulsive term is distance to the power 12 and entirely arbitrary, just because it was easier to calculate since you'd already calculated the power 6 term. So, in a sense, it was completely made up. But it was still a useful tool because it helped people start asking the right questions, and exploring these systems, even if it was, in a sense, wrong.
-"One of the hottest topics in science has two main conclusions:
Most published research is false
There is a reproducibility crisis in science
The first claim is often stated in a slightly different way: that most results of scientific experiments do not replicate."- [Ioannidis'] "Paper: Why most published research findings are false.
Main idea: People use hypothesis testing to determine if specific scientific discoveries are significant. This significance calculation is used as a screening mechanism in the scientific literature. Under assumptions about the way people perform these tests and report them it is possible to construct a universe where most published findings are false positive results.
Important drawback: The paper contains no real data, it is purely based on conjecture and simulation."
- then it summarises in the same way 7 other papers, including the Many Labs results.-"I do think that the reviewed papers are important contributions because they draw attention to real concerns about the modern scientific process. Namely
We need more statistical literacy
We need more computational literacy
We need to require code be published
We need mechanisms of peer review that deal with code
We need a culture that doesn't use reproducibility as a weapon
We need increased transparency in review and evaluation of papers"
- the final paragraph:"The Many Labs results suggest that the hype about the failures of science are, at the very least, premature. I think an equally important idea is that science has pretty much always worked with some number of false positive and irreplicable studies. This was beautifully described by Jared Horvath in this blog post from the Economist. I think the take home message is that regardless of the rate of false discoveries, the scientific process has led to amazing and life-altering discoveries."
[1] http://simplystatistics.org/2013/12/16/a-summary-of-the-evid...
Most people in business view any sort of rigorous research as a "cost" center. And this thinking has diffused into government too.
On the other hand marketing and advertising produces the most value for dollar. Why ?
In a market those activities generate new demand - something that a capitastic system needs in order to survive.
Why spend 20 years doing research while we could spend 1 year aggressively expanding our market. Even a 1% growth would he worth it given a 0% growth if the money was spent doing research.
I like that rather than villify a single person Nate silver points the finger at the system. This type of thinking is where we should all be heading towards.
It's a false dichotomy to say that you can only choose between 20 year research and 1 year marketing. Plenty of private money goes into research with long term payoff. There are private companies engaged in fusion research, because the payoff is huge if it is achieved. Even the public money that goes into publicly funded research comes from private earnings in the first place. So science and free markets are absolutely suited to each other. That's before we even start discussing corruption of objectives in totalitarian states, based on leaders whims (ie, Nazi science research, Lysenko etc)
Sorry to tell you but capitalist society as it exists is planned and managed economy. Large corporations are command economies.
You should see what science has discovered about the brain:
Objectively deconstructing the practice of p-hacking is science at work and a scientific work. So say, if we devised a way to measure p-value based research for it's p-hack-ability, we would have a way to validate what that research is saying (or not saying), as well as test the integrity of the research and its researchers.
Basically, they took a TON of demographic research in the health sciences, explored all possible hyperparameter tunings, and found that they could get p-values of 0.05 in either direction for most of the papers depending on choice of data sources and many other types of hyperparameters.