I'm a statistician, and when I wrote my own book about bad statistics in science (see https://www.statisticsdonewrong.com/), I made sure to reference studies which quantify how often errors occur in real published research. The rate is stunningly high. The average biomedical experiment is conducted with (a) a sample size which is far too small to detect an effect of the expected size, (b) a vague analysis plan which leads to exploratory analyses with high false positive rates, (c) frequent copy-and-paste errors and math mistakes in presenting important results, and (d) an overreliance on statistical methods to make up for poor experimental design.
This means the average published paper is likely a false positive, likely an overestimate of the true effect if not, and is barely reproducible.
Is this the fault of individual scientists? Partly, yes -- these problems have been pointed out for years in leading journals, but nobody takes action to do better research. It's also the fault of the grant funding systems which incentivize salami-slicing of results instead of doing one big, rigorous, well-designed study, and of journals which prefer dramatic but unreliable results over mundane but well-executed results. (Of course, the journal editors and reviewers are usually active scientists themselves.) I think the average researcher would like to "get it right", but has to focus on getting a career instead.
Just a few papers on the problem of poor sample sizes in biomedicine: http://journals.plos.org/plosbiology/article?id=10.1371/jour... http://rsos.royalsocietypublishing.org/content/4/2/160254 http://www.nature.com/nrn/journal/v14/n5/full/nrn3475.html
Psychology is going through an epistemological crisis for because reproducibility only recently became a hot topic there, and psychology experiments are super-cheap and easy to replicate compared to typical benchwork. But there have been equally dismal results when there's been sufficient incentive to replicate in other fields. http://www.nature.com/nature/journal/v483/n7391/full/483531a...
On the other hand, molecular biology and forward genetics have great traditions of reproducibility, and results in those fields tend to be pretty solid.
Where is the evidence of this? There is plenty saying otherwise:
http://www.sciencemag.org/news/2015/06/study-claims-28-billi...
http://www.ncbi.nlm.nih.gov/pmc/articles/PMC4270077/
http://www.slate.com/articles/health_and_science/future_tens...
http://www.nature.com/nrd/journal/v10/n9/full/nrd3439-c1.htm...
http://www.reuters.com/article/us-science-cancer-idUSBRE82R1...
http://www.nature.com/nature/journal/v483/n7391/full/483531a...
http://www.sciencemag.org/content/348/6242/1411
http://journals.plos.org/plosbiology/article?id=10.1371/jour...
http://www.nature.com/news/cancer-reproducibility-project-sc...
http://www.sciencemag.org/news/2017/01/rigorous-replication-...
On the other hand, I find your skepticism and independent thinking refreshing.
I see now I had an earlier discussion with tstactplsignore, back then they were seemingly incapable of understanding what I was saying (they kept thinking I denied the endonuclease activity).
Interestingly, from table S12 and my (selection for pre-existing mutants) model we can also explain a mysterious result they observed:
"Silent co-mutations in the repair oligonucleotides were introduced into >99% of D54H mutant alleles and ~3% of the F482S allele (Fig. 2c). Carryover of these silent SNPs indicates that the alleles are the product of HR and not de novo mutagenesis. The lower rate of coappearance of silent SNPs in F482S is presumably due to the larger distance between the two SNPs in the oligonucleotide and is rather common to see with single-stranded oligonucleotide donors19."
Rather than that ad hoc explanation, it is simply that the D54H cells had more silent only mutations to begin with (0.1% vs 0.0%). Regarding that 0.0%, an annoying thing is that they only report these percentages to one decimal place.
I'm not exactly clear on the number of cells present before the CRISPR-Cas9 treatment, but it sounds like 10^8, and then they let them grow for 72 hr + 7-12 days (total of 10-15 days) after the treatment. They also don't tell us how many cells were left at the end... but anyway if we assume these cells divide once a day, and 0.1% are preexisting mutants we could calculate the possible number of mutants thus:
Nt = 10^8
p = 0.001
d = 0:15
Nt*p*2^d
After 12 days we can get ~400 million cells from those initial pre-existing mutants, and after 15 days over as 3 billion. Of course other factors would probably come into play that limit this growth, I'm just saying it would be no problem for that small subset of the population to become dominant during the experiment. That is even if the 99,900,000 "WT" cells were just growth arrested rather than died.So I find those results to favor the "selection for pre-existing mutants" explanation over the "gene modification" one.
Because in cells that contain the target sequence (the complement to the guide RNA), Cas9 is damaging the DNA, leading to cell death and growth arrest. The small percentage of cells that already contained indels (thus reducing affinity for the guide RNA) are "immune", so they preferentially survive and divide to take over the population.
Also, it is 100% possible to create a testable model of something at the level of toxicity without knowing the details of the toxicity. I mean, here would be a simple one (written in R) where the wild type cells divide at 1/10th the rate of the mutants for some reason, so the mutants take over the population:
# Basic parameters of the cell culture
Ntotal = 10^8 # Initial total number of cells
Pmut0 = 0.001 # Propotion pre-existing mutants
Rdiv_wt = 0.1 # Divisions/day
Rdiv_mut = 1.0 # Divisions/day
# Calculate initial numbers of WT and mutant cells
Nwt0 = Ntotal*(1-Pmut0)
Nmut0 = Ntotal*Pmut0
# Convert between days and divisions
t = 0:15
Dwt = Rdiv_wt*t
Dmut = Rdiv_mut*t
# Calculate number of cells at each timepoint
Nt_wt = Nwt0*2^Dwt
Nt_mut = Nmut0*2^Dmut
# Calculate proportion of mutant cells in population at each timepoint
Pt_mut = Nt_mut/(Nt_wt + Nt_mut)
# Plot proportion of mutants vs time
plot(t, Pt_mut, type = "b", panel.first = grid(),
xlab = "Days Since Treatment",
ylab = "Proportion of Mutant Cells")
If the parameters are known accurately enough (initial number of cells, initial proportion of mutants, division rates, etc) this is a perfectly testable quantitative model.This was explained earlier I believe, so I am not sure where the confusion lies.
1) You start with a mixed population of cells. From the literature it looks like about 99-99.9% will lack indels at the target site, the rest have them.
2) The Cas9 will cut the DNA of cells lacking indels at that locus (ie the wt cells containing a sequence complementary to the guide RNA), thus killing and/or growth arresting those cells.
3) Meanwhile the cells with indels will continue living and proliferating since they lack the target sequence. These are "immune" to the CRISPR-induced damage.
Thus the proportion of WT cells will decrease, while the "mutants" will increase. It may help to play with the code of the simple model I shared earlier.
So you're stipulating the consensus understanding of Cas9's initial action, but arguing that this results in cell death rather than nonhomologous end joining repair?
Pr(T[0]|O) = Pr(O|T[0])Pr(T[0])/sum(Pr(O|T[i:n])Pr(T[i:n]))
This tells us the evidence for a given theory given the observations depends on the ratio of two things:1) How well the observations fit a theory (you can substitute hypothesis/model/etc) and how plausible a theory would be without the observations in question.
2) The above for all other theories
In other words, the default for the scientist is to be skeptical of any explanation until the others have been rendered implausible (ie ruled out). There is a bit more to it (eg the Pr(O|T[i]) terms depend on the precision of the predictions, which remain vague in the case of NHEJ despite generous funding), but that is pretty much it.
I don't see how it is even possible for you to gather that from what I have said? All you could possibly have is a rough estimate of the ratio between the priors for NHEJ vs selection. Selection is a far more common and well studied process...
There also exist simple quantitative models of that process, which allow precise predictions, something lacking in the case of NHEJ models afaik, which must remain vague. So the likelihoods will also be narrower in the selection case.
Anyway, it was productive to discuss the specific paper and model, but is now getting philosophical and pointless. I only mentioned the scientific thought process because you asked why I would be skeptical.
Quantitative reasoning can give you a lot of leverage in systematizing and inferring from a body of knowledge, but if your reasoning doesn't start from that knowledge it will lead you nowhere. In this case, the necessary knowledge is the structure and mechanism of DNA repair.
When Dirichlet computed the probability that the Sun wouldn't rise tomorrow, given a flat two-event Dirichlet prior and the observation that it had risen every day for the last 6,000 years, he added that of course, for people who understood the workings of the solar system, the probability is far, far, lower. Your arguments here are like that. They just don't take into account the relevant facts of molecular biology.
2) But please, let us skip the philosophy and talk directly about this topic, because that argument is not even necessary. It sounds to me like you do not believe that double strand breaks can lead to cell death and growth arrest? You find this implausible? I am really surprised that this is an objection.
We're talking about the entirety of basic research. The "evidence" for this is the incredible advances in our understanding of how the cell and the molecules of life work in the last 50 years. Every single experiment done in the modern life sciences would not be even possible to contemplate doing if all of the molecular biology it is based on was not extremely reproducible. Entire fields of modern biology like genomics, structural biology, genome editing, and more could not even be fathomed to exist if this were not true.
Nobody tallies up how often biology "works", only how often a few experiments attempting to prove specific hypotheses don't work. Basic science in molecular biology and microbiology are extremely reproducible most of the time and nobody writing these articles questions that.
No offense, but this is some creationist-tier science denial you have going on if you don't believe this, and it isn't the responsibility of the scientific community to educate people who've read one time many articles about reproducibility that science works.
Yes, everything needs to be redone since at least the 1980s, and probably the 1940s (whenever NHST became the primary method of assessing what is correct or not). We have no idea what is actually going on. It is scary, I know. It took me years to accept.
>"this is some creationist-tier science denial you have going on if you don't believe this"
Ok, link to a specific paper and I will discuss it with you. In my experience, if it is biomed, no one will have ever published a replication of any results it contains. There will be no model capable of quantitative prediction of anything, and conclusions drawn will also almost certainly be fatally flawed due to the vagueness of their explanation.
Secondly, many biological findings are not quantitative or statistical in nature- they simply are observations that have been repeated tens of thousands of times.
Honestly I don't really care enough to respond over something that nobody else in the world really believes. Your claim that you were "trained in biomedicine" is cute but difficult to take seriously. It isn't worth my or anyone else's time to fight your bizarre and excessive ignorance on this topic. I just think it is important to point out to the community that you have zero credibility on this issue and are spouting nonsense that real statisticians, scientists, and computational biologists would laugh at.
Anyway, I agree it is pointless to continue a discussion at this level, which is why I tried to guide us to talking about specific findings (of your choice).
I believe this may account for the vast majority of the "editing" they are measuring, and claims of increased efficiency over earlier methods (eg TALEN) are due to this mechanism. They have not done proper due diligence of ruling out other explanations and prematurely starting running with their favorite, most hype-able, hypothesis.
Can you find one paper where they report the efficiency of inserting a cassette via HDR, or do they all only look at NHEJ when talking about that? I have only seen the latter.
Also, this alternative explanation has real-life consequences. If correct, the very mechanism via which CRISPR-Cas9 works is toxic, meaning all the hundreds of millions of dollars (billions?) currently being spent trying to make it less toxic will be wasted.
Cutting DNA is certainly a great self-defense mechanism used by the bacteria that causes strep throat to attack other bacteria. But in the ways that Cas9 will eventually be used, likely its entire endonuclease activity will be stripped, leaving only its homing ability. And for that you can swap in any other useful protein functions that now happen at a particular sequence location. And that is what makes Cas9 so exciting. That NHEJ works at all is just a bonus to get us to some early results quickly.
fCas9 or dCas9, or a nickase (different endonuclease, and endonuclease-dead respectively) as well as a host of other variants actually hold great promise precisely and solely in their ability to locate a sequence.
I know large sample sizes are often impractical, and you have to make do. But given that, results must be presented with all sorts of disclaimers, since results from underpowered studies are frequently wrong or exaggerated (https://www.ncbi.nlm.nih.gov/pubmed/18633328). That's not what happens -- scientists who should be aware that their studies have very little value as evidence instead present them as groundbreaking definitive results.
I would much rather see slow, difficult, tentative studies instead of the high-speed spray of nearly meaningless results we see every day. But that will take a dramatic change in career and funding incentives.
A major hurdle is the experiments do keep getting slower and slower. As a recently graduated biomedical grad student I spent literally 6.5 years of my life trying to get an experiment to work enough times to glean useful data out of it. I knew exactly how much power my data had (I had literally years to think and rethink about it). The number of stars that had to align to get equipment, protocols, controls, materials, animals, sleep-schedules, etc. all aligned to get data out of an incredibly complicated system meant that there is no fast data. Useful biological data is hard-won. And it's getting harder and harder (read, more expensive and labor-intensive).
You could say, well, then just don't expect every scientist to do their own (independent) research (see the author lists on CERN publications...). That indeed would change incentives a lot in the biomedical fields - and I actually do think for the better.
The last point is I think there's a big difference between how a practicing scientist sees a 'published paper' in their field, and how everyone else sees it. A published paper really is just a mark of what you did, and what happened, and maybe a basic interpretation. It really does not claim to be 'true' in a strong sense. And there are entirely reasonable situations where different papers come to opposite (justifiable) conclusions. This is itself weighted data to be used in the next motions of the field precisely because sample size and (more often than not) experimental complexity makes clarity hard to come by for a single paper.
Consider these questions:
1) Can you define a p-value and articulate why you would perform a significance test?
2) How many replication studies have you published?
3) How often do you perform your experiments blinded to intervention?
4) Is there any quantitative prediction that has been made by anyone in your sub-field of research ever? I'm not even asking for an accurate one, just the publication of a point or small-interval prediction.
Some of the best science I've seen were experiments where the result was guaranteed to be one of a few different possibilities - and therefore any concrete result was interesting. That's mostly attributable to good experimental design. Cell A does X under condition 1. Cell A does Y under condition 2. What is the protein/pathway/atom/molecule responsible for triggering condition Y under condition 2 (or conversely triggering X under condition 1). Pick your level of inquiry, discover the answer, remove the answer to test for necessity, and add the answer back to a null background to test for sufficiency. If all are true, you have no found something useful and true. Statistics was only a tool to help you accomplish the above and actually had little to no bearing on the knowledge being created. Good experimental design could create situations where the culminating experiment either presented an indicator or did not, confirming the questioned hypothesis true or false.
Those kinds of factual discoveries were the biomedical science I was taught to probe for, where whatever the answer was it would be significant, interesting, and useful. And there is no way to successfully conduct the above experiments in such a way that statistics even can be of much relevance to the scientific question itself. Either you discover a mechanism or you do not, and logically it can only be one of a very few options.