More than half of high-impact cancer lab studies could not be replicated
science.org
science.org
If you've ever been within a mile of a research lab this shouldn't be remotely shocking. Typical research is done rapidly, sloppily, and in a context where getting the "right" answer is the only incentive.
The replication crisis that hit psych is only being held back by (IMO) that bio is harder to replicate due to tooling, protocols, reagents, cell lines, etc.
In complex systems you have the degrees of freedom to be wrong at a scale that folks still do not appreciate.
Let me preface that by saying that 50% sounds like an awful lot, even bearing in mind the following.
On top of that, these studies focus on highly complicated phenomena, where almost anything could play a role. It is common that one factor that is not necessarily accounted for in a study turns out to be important, even in well-run studies. Flukes happen as well.
Failure to replicate per se is not a problem, if the original study was done well and honestly. It is still some data points that could be re interpreted later in the light of other studies.
Now, the fact that there is so little incentive to publish confirmation studies is a big problem, because it means that these non-reproducible works are not as challenged as they should be.
Remember that one study is meaningless. Even with tiny rates of false positive or negative, erroneous conclusions are bound to happen at some point. A result is not significant unless it has been observed independently by different research groups.
I don't quite get this point: this is cancer research, co in theory all related institutions should be interested in replicating breakthrough studies, right? And failure to replicate it, time after time, should be a cause of concern, shouldn't it?
The lifecycle for a junior PI is about 5 years, which really means they have about 3 years to produce a result sufficient to advance. This is enough for a bit of validation but unlikely anything beyond their lab, so the result is likely to be untested. But who in that lab cares? They just need to get the next job, so a 3-4 year timescale is fine. The PI cares long-term because a lack of new productivity means a lack of ongoing funding, which means they either improve the quality of their processes (which means herding the junior researchers through that system) or simply continuing to hack their way through.
There is a reason why people moan about scientific theories advancing one death of an old professor at a time.
For external industry funders, they have a strong and vested interest in identifying good output before it is public. The strategy here typically is a mixture of identifying good researchers and making them kings (consistent funding, direct lines of communication, access to nonpublic materials, etc), and of placing lots of small investments to get in the room for discussion. An example of this mixture is the BP deal at UC Berkeley [2].
As for shifting standard of care, this is a commonly-expressed goal but often the research is far enough away that the individuals doing the work lack close communication with the actual practitioners. Some institutes are trying to resolve this by close physical and social proximity (e.g. TU Dresden's medical campus) but this is far from the norm.
[1] https://www.helmholtz-muenchen.de/fileadmin/Jahresbericht/Ja... [2] https://www.berkeley.edu/news/media/releases/2007/02/01_ebi....
Well, no. Grants don't pay for replication. Replication studies are less likely to get published or get citations. In other words, if you want to do replication, you wont be able to find money for it. If you find money on it, there will be punishment in terms of lesser career progress.
The system does not value replication.
Well, most of us (scientists in general) do care, but we have opposite incentives.
But this is the opposite of how it should be! If we learned anything from scientific research is that if a breakthrough study is published, it means almost nothing until it gets verified. Actually, these days I tend to believe meta-analyses almost exclusively. But you need time for these to appear.
So when I read about some exciting new study, my first reaction is, "Well, that's interesting, we'll see what other scientists say." So actually we should look forward to the first replication as the first indicator of whether the original study was correct or not. The authors of the original study and the journal board should be grateful to the replication team for helping them do their job! Rejecting replication is a disservice to science.
I share your concern. The operative word here is “should”. I. Practice replication studies have an opportunity cost because they won’t end up in a prestigious journal, they are harder to use in grant applications, and the time and resources could be spent on studies that would be more useful in these ways.
> And failure to replicate it, time after time, should be a cause of concern, shouldn't it?
It would be, and it would also indicate that a technique is dodgy, or that a research group or scientist is unreliable.
Not necessarily: it could also mean that we don't properly understand the process, and therefore essential steps/prerequisites in the experiment are left undocumented. That type of error is more fundamental than merely the technique or the researcher.
No but in this case a common reason for failure to replicate was that the papers didn't actually include enough details to replicate them to begin with. This should be grounds for failing "original study done well" because the reason they're paid to do research in the first place, is to propagate accurate and useful knowledge. Missing critical details means the study is a failure.
This looks like such a fundamental problem to me. Zooming out, studies are explorative into the unknown, outside of our knowledge. We should expect most of them to find nothing.
Every ML practitioner should spend some time working with genetic data just to realize how weird things can get when you have millions of sparse features for a few thousand cases and all sorts of batch effects. Like to do it well you need to control for the lightbulbs in the microarray machine you were using or the sequencing center used, minor version of the sequencing technology and reagents order the subjects were sequenced in, who ran the machine, geographic origin of subjects etc.
And then when you go out and try to tie your data into the literature you need to be aware of all sorts of self reinforcing bias in which genes get published on.
Amazing and artful applications of mathematics have been developed to address all of these and the care bioinformatic papers address things like cross validation with makes ML look like a joke. Beyond that its even common to go out and get more data or validate with different technology before publishing but it is still easy to get things wrong.
If we truly want reproducible research we need to address the batch effects with less noisy sequencing machines and an assembly line like approach to generating orders of magnitude more data as cheaply as possible including new data on new subjects to verify studies after the fact.
The journalism side of Science magazine is far too credulous. Most science is wrong. Even when it is published in Science (perhaps, especially when it is published in Science).
I'm not shocked either though. Not because biology is hard, but because academia is corrupt.
Whether this is malice or not is kind of irrelevant - they have failed to meet the most basic requirements expected of researchers. They took the money but did not deliver the expected work artifacts.
There was a recent blog post / HN thread [1] about groups never admitting failure, it's very relevant here. The funding agency has to have a process for vetting applications, and what's more reasonable than the ability to manage their own work, i.e. propose a plan and successfully execute on it? Combine it with slow cadence of funding and you have a list of promised results for next 3+ years which you must achieve somehow or you'll never see a grant from that agency again. And they give out a lot of grants, so reviewing reports in detail to be able to distinguish "this result we promised was not achieved because we are incompetent" from "this result is negative, and we need to change direction based on new research" is just infeasible.
In commercial research the situation is much better in my experience, probably because your "funding agency" is more narrowly focused and is typically interested in getting useful results.
Lazy? Absolutely. Prone to exaggeration? Definitely. Self-delusion? Always. Corrupt? Maybe sometimes.
Science is a human process, and is subject to all of the problems created by any group of humans. But the magic of it is that, in the long term, everyone's personal incentives produce an emergent result of overall correctness. You just have to understand that most papers are wrong, and never take anything at face value.
People love to blame bad statistics, but it's much more obvious that the issue is the "publish or perish" culture and a perpetual race to manufacture credibility in academia. If you're curious you can trace some of this back to the 70s when peer review as we know it today was essentially born, precisely as a means of proving credibility during a period of academic funding crunches.
As much as it's easy to point a finger at academic I increasingly find the same thing in industry. Within my own career the number of companies that are doing nothing more than finding elaborate ways to burn investor funding (without even realizing they are participating in this) has grown to consume almost the entire tech industry.
At this point the most shocking news isn't fraud, it would be to find someone out their down real, meaningful work.
https://www.amazon.com/Idea-Factory-Great-American-Innovatio...
We owe most of our "modern world" to the hard science done at Bell Labs in the 1900s, before the corporation was singularly focused on quarterly profits.
There's an analogy in Software Engineering. Programmers are akin to Artisans or Craftsmen, but the industry wants and needs them to be something more like little factories. I don't think that approach makes good software, even if it's a good product. Just as I don't think the system we have created in academia doesn't produce good research, even though it makes good careers.
Still, if scientific progress is in any way real and not cyclical (as in Kuhn), then science should be getting harder. It stands to reason that if there are fewer and fewer secrets to find, finding them (or randomly stumbling upon them) will be a lot rarer.
We tend to assume that there is an infinite amount of knowledge to be gained of the natural world, but it's not necessarily so. It's especially not evident that our tools are evolving as fast as they need to, or even that they can evolve much further.
I'm not saying "everything has been found", more like "all progress makes further progress require more effort, by definition".
My day job involves designing measurement equipment, some of which is used in life science. I absolutely care about knowing whether a measurement is any good or not, to the point where it's not just a paycheck, but a true passion. And I work with people in actual research who care just as much, using my stuff to do front-line work.
There are certainly problems, and I don't wish to defend the "publish or perish" system, but I don't think a cynical view of the causes is a complete or even accurate picture of what's going on.
There is a stupendous amount of new science being carried out. Science that we'll look back on 80 years later (just as we look back on the science of the 1920s) and recognize how impactful that science has been. In fact, I would imagine that the science that is being worked on today will prove far more impactful than the research of the 1920s.
Think about the seminal technologies that are still in their infancy! Deep learning in it's current form (AKA the form that actually works) is ~10 years old. That's it. And in those ~10 years we've gone from cute theory to being able to convincingly emulate video, translate languages, solve dictation, and generate near-human quality text. Then there's CRISPR, which could potentially cure 1000s of diseases and dramatically improve diagnostic tools, and is similarly ~20 years old. There are many thousands of other groundbreaking technologies that are being worked on as we speak.
Contemporary scientists are conducting this research.
In general, all the interesting work that enabled today's deep learning boom happened towards the end of the 20th century and recent advances are primarily owed to increases in computational power and availability of data sets.
Says not me, but Geoff Hinton:
Geoffrey Hinton: I think it’s mainly because of the amount of computation and the amount of data now around but it’s also partly because there have been some technical improvements in the algorithms. Particularly in the algorithms for doing unsupervised learning where you’re not told what the right answer is but the main thing is the computation and the amount of data.
The algorithms we had in the old days would have worked perfectly well if computers had been a million times faster and datasets had been a million times bigger but if we’d said that thirty years ago people would have just laughed.
http://techjaw.com/2015/06/07/geoffrey-hinton-deep-learning-...
So.. I guess I disagree strongly with your premise? Look, it takes some time for these ideas to propagate to non-experts. In 20 years a hacker-news poster will likely post the same comment, referencing the 2017 Transformer paper to enforce the idea that all the _good_ science in DL was done years ago.
>> In 20 years a hacker-news poster will likely post the same comment, referencing the 2017 Transformer paper to enforce the idea that all the _good_ science in DL was done years ago.
On the contrary, my expectation is that in 20 years from now people will point to Transformers and laugh and joke about how misguided people were "back then" to think that just building larger neural nets would lead to AGI, as they do now for earlier AI approaches.
What I'm not convinced is that you are any different than the "non-expert" in your comment, to whom "it takes some time" for ideas to propagate. You are holding up Transformers as some kind of scientific breakthrough but that is only the most obvious conclusion to draw from the overhyping of the approach on social media and the lay press.
Transformers basically took anything that was "sequence learning" (large fractions of computational biology are) and made it work 100X better over night.
Regarding Transformers, the results on protein structure prediction are certainly impressive, but they are not enough to justify your comment that "transformers have literally transformed science". At most we can tell that they have considerably improved protein structure prediction. Perhaps your comment was an exaggeration?
Edit: I can't find the paper on the ICML 2015 website, although I can see that it has a few hundred citations, as a preprint, which is not uncommon these days. But, could you point to some published work you contributed to?
* Cosmology. Most of the cosmological theories (e.g., Big Bang) didn't develop until after 1950, and the Hubble Space Telescope has advanced the field considerably.
* Supersymmetry/superstring theory. I hesitate to put this here because I think this is a wrong turn for theoretical physics, but it is a very well-known model that developed in the 1970s, with the superstring theory revolution occurring in the 1970s.
* Central dogma of molecular biology. Proposed in the late 1950s, this states that basically DNA -> RNA -> protein, and nothing goes back the other way. It has also taken a massive beating in recent decades, as we've discovered there's quite a few ways that information gets passed around in cells that don't follow this path (e.g., epigenetics, and it's an epic acerbic word match if you want to restrict epigenetic only to things like methylation of DNA or include any other heritable non-genetic traits as epigenetic).
* Plate tectonics. This became accepted geological theory only around the late 1960s.
This is just a list of pure science discoveries I can think of off the list of my head. Stretch to applied science advancements, and there's several that have caused massive revisions to fields--organic synthesis techniques, NMR spectroscopy, radiocarbon dating, genetic sequencing are all things that came to mind while compiling the above list. If you switch over to math/CS side of things, you can also consider that the development of RSA quite literally opened up a new world of what can be done with computers, or the Cook-Levin Theorem that really kicks off computability theory (and leads to the asking of the question P=NP). Pretty much everything about CS only develops in the late 20th and early 21st century.
Hell, that last comment prompted me to remember something else that rather fundamentally changes how we view the universe. It turns out that it's far easier to teach a computer how to play chess than speak with the intelligence of a three-year-old, never mind an adult. What can be more mind-blowing than discovering that we can't make a more effective learner than a toddler?
I also have a hunch that replication problems are not equal everywhere. I suspect that biomedicine and the social sciences are worst off (although I think CS ML might also have some issues).
If you can't reproduce it, it's not science.
I believe in science, but I also know scientists are humans and thus fallible.
We live in a post-truth world, I have no certainty about anything anymore except what I can check by myself, which is not much.
Sadly I couldn’t remotely afford the equipment to do things on my own (nor do I think I could handle the initial learning workload these days without some structure).
There are also community initiatives like BosLab [2] which are more similar to maker spaces in structure.
Let me throw a tiny wrench into your logical reasoning: I can't reproduce most results, does it mean most results are not science?
Absence of evidence is not evidence of absence. You should not expect scientific experiments to be replicated every time.
Of course the result could be wrong, and, assuming scientist B knows what they are doing, it's an indication that the result is indeed wrong. But it could also be true.
This is not a comment on the state of science and whether incentives are set correctly at the moment (they aren't).
Correct, but until scientist A’s result can be reproduced by someone, anyone, relying on scientist A’s result is really more akin to a religious leap of faith, then science. When any belief based on faith alone is presented as an unquestionable truth, it deserves skepticism.
But if you can't reproduce with the exact same conditions than the original experiment then it's not science.
I agree, that's trivial. What's not so trivial: Where does the "experiment" start and where do the "conditions" end?
In theory the scientific method is easy, in practice (i.e. state of the art research) you are studying highly complicated phenomena, that are, per definition, not well understood. Often you don't know the mechanism of action, let alone the factors that influence it. I don't blame you - if you haven't done scientific research, it is very difficult to appreciate.
So again, from the failure to replicate a result it does not follow that the result is wrong, or that "it's not science". That's a misunderstanding of the replication crisis.
What bothers me is that some people are very dogmatic about things that are published and most of the time those people aren't scientists either.
Just a point that not all science can have empirical and reproducible study.
replication is boring, yet critical, and we should be compensating highly skilled scientists for doing it.
in seriousness though, there's enough money and trained scientists looking for work... moreover i think there's a whole personality type that is prevalent in science that would be very happy to replicate and reproduce results...
I have very little insight into this area of research, so I’d just have to guess everyone’s in a hurry and reaching for results or published papers.
i don't know what you'd find in a financial audit that a scientific audit wouldn't tell you, but a passing financial audit won't tell you if a scientific finding actually sticks. therefore, seems like a better deal to audit the results, no?
Financial audits would uncover financial fraud.
> "seems like a better deal to audit the results, no?"
Better to audit both.
this discussion is about ensuring scientific velocity by increasing the rigor required to declare success and add new knowledge to the commons.
unless there's large scale fraud going on, i honestly don't think bringing in additional auditors to look at books that are already professionally kept is going to benefit anyone other than said auditors.
If it was only their money on the line, that would probably suffice. But when researchers (and by extension the institutions they work at) are receiving government grants, I think there need to be independent audits.
these are all nonprofit institutions, some are government run, i'm pretty sure they're required to engage independent professional auditing firms just as a condition of their tax exempt status, let alone acceptance of federal grant dollars.
- breakthrough study is published
- because it’s breakthrough, it attracts attention and other labs look at it (we often tried to repeat breakthrough work in our lab)
- if it can’t be replicated, interest just peters out OR it drives someone else to “fix” it. If it’s so important people need to know it gets published.
- if it can be replicated further research continues
Given the various difficulties of culturing cell lines, this is not surprising.
What is infuriating to me is unresponsive author queries about data and protocols. That should result in some sort of ding to the paper on the journal's website.
The Repeatability Experiment of SIGMOD 2008 https://pages.saclay.inria.fr/ioana.manolescu/PAPERS/SIGRecR...
Also, the lab animal telomere issue has corrupted almost all animal studies in the last 50 years, but there doesn't seem to be any effort to correct course, or even to acknowledge the problem.
I think this is one symptom of a much greater issue: the inbred mice strains we use are disastrously weird (arising from the initial population artificially bred to make research easier). For the sake of controlling for genetics, we've chosen to make lab research translation drastically less effective.
2009 Nobel prize given for this work.
Lab animals, due to the way they are bred, are essentially susceptible to cancer and health issues not seen in normal wildtype populations. Shortened telomeres are the primary danger, increasing mutation rate and cell senescence.
Essentially, all data involving the expectation of normal cellular behavior is subject to question - signaling from lab animals is corrupted.
Apparently leukocyte telomere length is heritable though, is that what you are talking about?
Most cancer studies in mice are xenografts of human cancer cells so any relevance of mice telomeres is completely lost anyway.
Also this obsession with telomeres being the fundamental key for all our biological problems seems so conspiratorial. Like you think our cells which do unimaginable things can't keep their DNA ends frothy? Like that's what is gonna spell their doom? Or you consider telomeres are just a symptom of a cells longevity decided by other factors ?
lab rats and mice aren't useful model organisms, but it's not for this reason.
Unfortunately the argument of "too expensive" or "too difficult" fails to fly or appeal to me when astro/particle/nuclear physics always split their people/experiments/resources into multiple teams, eventually to compare and contrast results.
"Big-pharma", often supporting this research has enough money/resources to waste it seems that they will pursue any research to test it rather than verify it outright. And this culture has fed back into the mindset of most biologists these days it seems...
This has to be one of the biggest drains/in-efficiencies in the field and is driven by naive corporate policy/greed... such a waste.
Defining replication gets tricky when your N is 1, and your population is one system, perhaps at a particular moment in time.
Isn't cancer basically the result of entropy? With enough imperfect cell-divisions, sooner or later one of them is going to have a mutation that causes unbound cell-division (i.e. growth).
Tl;dr misaligned incentives in science, the academic job market, and the funding system, the publish or perish dogma, and an overproduction of PhDs.