Millions of animals missing from scientific studies
sciencemag.org
sciencemag.org
One of the things that is latent in this article is that in the US you are supposed to have done a power analysis in order to justify the number of animals you are using in your study. Almost no one does this, and it is not surprising -- if you are doing cutting edge research you are in unknown unknown territory and any power analysis is likely to be no better than a guess. In a sense it is farcical that exploratory research needs to pretend that it is always successful.
Countless animal experiments are uninterpretable simply because they were poorly executed by first year graduate students. There is no other way to learn. Write them off as animals used for training if we want full accounting, but researchers' time is scarce enough as it is so asking them to publish uninterpretable results is a non-starter.
On the other hand I think it is important to publish as many results as possible to avoid file draw effects etc. Data publishing might be one way around this, but at the moment few labs anywhere in any field have the know-how to publish raw data for negative results, even if it is just sticking files in a git repo and getting a DOI from zenodo.
> I'm sorry but this simply isn't true. During my masters i did in vivo surgery on 30 rodents and tested how many the procedure was succesful in (10). I had to record how many the procedure was sucessful in and write this up in my methods. I'm not sure how you see keeping a record of this as something that takes a significant length of time that you wouldn't bother to record it.
In fact i'm pretty sure (at least for the UK) that keeping records like this is essential under the home office rules for animal safety in scientific procedures. These rules are there to help with the 3 R's (reduce, replace, refinement). This guidance aims to cut down on the use of animals in research or at least ensure quality experiments are being ran. If you don't measure - how can you improve?
Despite the possibility of differing opinions in this thread, I think the article has not done a good job of explaining the realities of the process, so anyone here who is not familiar with how animal work gets done is coming in with a purist misunderstanding. Everything is reported (in the US), but not in publications.
Looking at publications alone for animal accounting would be like if I looked at the checking accounts for everyone in a country and wondered where all the money went. Of course it's in savings, investments, cash under the bed... but I only looked in one place. I cannot conclude money is unaccounted for when my search was incomplete by design.
This is different from the UK where the 3Rs are much more strongly enforced, to the point where I remember asking a question back in 2015 to the then head of UK animal research about reproducibility, and getting the answer that he wouldn't approve the use of animals just to replicate an already completed study. In the US in some fields animals will be used just to replicate a result because another lab needs to know for sure that it is real before expending even more animals for a potentially useless follow up study. The way the numbers play out in practice, we would be much better off doubling if not 10xing the number of animals used in initial publications to avoid the 10x replication studies that will be done inside other labs to make sure that the result is real. Of course if we did this then the publication rate in many fields would be cut in half, or decreased by an order of magnitude.
I disagree with the parents that the remaining animals are even primarily used for training purposes. There are countless ways that experiments may fail with uninteresting results, that do not count as null results.
Did you interpret the article as saying that exploratory research must be successful? I read the opposite, that "unsuccessful" (defined here by me as negative or inconclusive) research should be published more. What am I missing?
Suppose you want to see how different types of neurons are distributed in the brain. You hypothesize that two specific subtypes of neurons are always found in close proximity in one condition (brain area, developmental stage, disease vs health, etc), but not another. There are a lot of ways to do this, so you pick one and start.
If things go well, your antibodies selectively label each neuron type. You count the pairs of neurons that are neighbors (or not) in condition A, those that are neighbors (or not) in condition B, and do some stats. If you get this far, I agree it ought to be possible to publish something, regardless of whether the proportions are wildly different, exactly the same, or somewhere in between.
However, things often go wrong. These protocols have a lot of free parameters and it's often not feasible to calculate the best ones from first principles. As a result, you try something and notice that the result is wildly implausible: maybe everything is labelled as one of your cell types, even stuff that isn't neurons. You tweak the protocol, and now nothing is labelled. This is also implausible--the tissue is from a normal animal--so you make some more adjustments and try again. Perhaps you even change techniques altogether and use FISH or a viral vector instead of immunohistochemistry.
The final protocol (if successful) is always included in a paper, but these intermediate failures are usually not and I'm not sure it makes sense to. Suppose the solution was to use a better antibody from a different company. The pilot experiments where we varied the incubation time, sample prep, etc using a dud antibody are fantastically uninteresting. Furthermore, people often change multiple parameters at the same time; going back and convincingly demonstrating which one "matters" would require a lot more work for a fairly limited payoff.
Finally, people also adapt their research question based on the data they can obtain. Maybe you can reliably label one type of neuron, but not the other, so you decide to focus on how those cells' locations vary during development. If so, it'd be weird to report a bunch of failures of an unrelated technique in the resulting paper.
yes, it's important for work to see the light of day, but we do not want to disincentivise risky or exploratory work.
Are journals even going to publish most no-result studies? (I.e. a result of no significant correlation found.) Journals pick and choose the best articles to peer-review and publish out of all submissions.
It feels like there ought to be some other system for tracking no-result and unpublished studies, something easy and low-friction -- not the trouble of an entire paper, but maybe something more like the length of an abstract?
Then if you think about running a study, you can run a search to see if others have already done something similar and found nothing... and reach out to them personally if you need more info.
Is there nothing like this?
That's not a trivial amount of work in the least. If you're a professor trying to establish your lab, would you use a few months of your post-docs time writing up experiments that don't work, or working on new ideas that might work?
(counterpoint: string theory... I kid, I kid!)
Yes they are. Not all journals of course, but some are very explicit that they will publish based on quality and not based on outcome, e.g. Plos One. And you always have preprint servers as a backup publishing option.
There are many reasons for publication bias. That you can't find a place to publish your findings is not one of them.
Should the same apply bacterial or fungal cultures? What about outside of biology? Should someone studying psychology of personality who gets a non-result from a survey of psych department undergrads publish it? Should a materials scientist who builds a bad battery describe it in detail? Should I create a PR for every branch which I don't want to merge?
At some point, any new endeavor can only start by choosing to not spend 5 years reviewing all the one-off failures.
Repeat.
I hard disagree that pilot projects and results from data with technical issues should be published. There is a real cost to write-ups (time). Pilot projects usually are done just to prove feasibility of a protocol, and data retrieved from experiments with technical issues is likely noise. This could include inaccurate measurements, cross-contamination, etc. Taking the time to write and publish all the junk is, frankly prohibitive. I do think null-results should be publishable, however.
The title is a bit of click bait also: The animals are not missing. Make a request of the relevant Institutional Review Board.
Are any of you voting my comment down actually scientists?
But this explains the value of scientific meetings and poster sessions where I can just tell this random tidbit to interested people without the burden of peer review.
An interesting recent development I've observed once was a tweet describing the fact that a dye (Texas red) labels brain vasculature even when injected subcutaneously. This random fact is not in the literature but is quite helpful from a procedural perspective. I think that science twitter has a potentially super valuable role to play in reporting unexpected or otherwise difficult to publish findings.
On the other hand, data sharing protocols mean the data is going to be made available, so I think that probably addresses that issue.
The problem with pilot data is that lots of things can change in the course of running the pilot; tweaks of the experimental protocol, bug fixes to code, etc.
Instead, they demonstrate that doing something this specific way is too much of a hassle, too unreliable, or too expensive to answer a particular question in a particular context: skills, budget, timeline, resources, alternative ideas, etc.
I don't see how you could write pilot results up ("Trust me, I'm usually pretty good at these sorts of things but this didn't work well"?). Meanwhile, disentangling these factors in a rigorous, generalizable way would turn it into an entirely different project.
Jumping to the conclusion that negative results aren't being reported (which is indeed its own big problem) from a difference in animal number suggests lack of understanding of the process, but perhaps formal accounting of research animals differs by country.
I do not come from academia so I may be mistaken, but why would you publish a non statistically significant result? There seems to me there's a very distinct difference between "this is most likely true" and "this might be true, we can't say yet". The first is useful, the second adds to the chatter needlessly. Some may say that it promotes further research, but what if a false weakly proven result misleads instead?
Now, if you had 19 studies which were inconclusive and 1 study that claimed to have conclusions, you would look at it very differently than one study that claimed to have conclusions.
> why would you publish a non statistically significant result
There's actually two separate things going on here.
The first is that researches (at least in the biomedical sciences) often decide not to waste their time writing up results that are either weak, inconclusive, or negative. This is only mildly controversial - everyone involved recognizes the inherent time availability constraint, but knowing what didn't work can help to inform the field at large in it's own way.
The second is that the term "statistically significant" itself has become _highly_ controversial in recent years. The issue is that when you perform statistical tests, they generally (oversimplifying) spit out a number telling you what the odds of your results being correct are.
In concrete terms, say you have some graphs that show a slight improvement in some condition when a drug is used. Is the drug actually effective, or is the "improvement" actually just random measurement errors? So you run a statistical test on your data and it gives you a score of 80%. It's telling you that there's an 80% chance that your results were due to an actual improvement and a 20% chance that they were due to random chance.
If you had unlimited time and money you could just keep collecting results endlessly. Eventually, that score would either go towards 100% (it definitely works) or 0% (it definitely doesn't work). But this is the real world, where we don't have unlimited time and money (particularly academic researchers).
So when is your result worth publishing? How do you decide when to throw in the towel? And if you're reviewing papers for a journal, how low a score is acceptable before you vote to reject the paper on the grounds that the conclusions are unreliable and not worth looking at?
Enter the term "statistical significance". At some point, people started classifying scores that were above some arbitrary threshold as "significant" and those below it as "not significant". But there's an obvious problem here - some measure of probability flipping from (say) 89.999% to 90.000% doesn't do anything magical! Worse, what a future reader intends to do with the results will determine how important any given score is in that particular case. Clearly, results and their associated score need to be interpreted in context instead of blindly. Using a term such as "statistical significance" flies in the face of that by actively encouraging lazy thinking.
So the controversy being referred to in that specific sentence you quoted isn't the decision not to publish but rather the usage of the term itself. (Which is confusing, because the article at large is addressing the controversy surrounding not publishing.)
What is really needed is publication of the data. The big problem with that is that there are no good systems for data publication that don't require significant extra work. Ideally the data would be made available directly from some electronic labbook system, with the relevant metadata attached to it. At the moment all systems are miles away from that (even if we don't consider how bad record keeping is in many research labs).
If funding agencies would really want to do something about open data they would significantly invest into a good data system and make data entry mandatory and in particular not just at the end, but during the study (private, but then made available at the selection of the PI or after some time)
My understanding was that the gold standard for research is to come up with a hypothesis, design an experiment to test that hypothesis, collecting data in such a way that you have statistical independence between the variables you wish to measure, and finally analyze the data and conclude check if you have confirmed your hypothesis or if it remains unconfirmed.
But data collected correctly to test one hypothesis is not necessarily good for testing other hypotheses - e.g. perhaps it was irrelevant for your research if all the animals were female and might have had birth defects, but that doesn't mean your data is useful for someone studying the whole population.
This would mean that not only do you have to publish the raw data, but also describe your data collection methods, your assumptions, things you didn't check etc. How far are you from publishing your whole paper at that point?
I would argue that the more data is out there, the less easy it is to "p-hack", i.e. it's much more difficult to find some weird correlation if you have a huge population. Also I would not call researchers who use others data "unscrupulous". In fact that is been done in meta studies already and is exactly the point of the exercise. Increase the amount of data published so if some "unscrupulous" or "erroneous" researcher publishes a study with some some spurious correlation, we can look at a lot more data to see if that also exists there.
>My understanding was that the gold standard for research is to come up with a hypothesis, design an experiment to test that hypothesis, collecting data in such a way that you have statistical independence between the variables you wish to measure, and finally analyze the data and conclude check if you have confirmed your hypothesis or if it remains unconfirmed.
>But data collected correctly to test one hypothesis is not necessarily good for testing other hypotheses - e.g. perhaps it was irrelevant for your research if all the animals were female and might have had birth defects, but that doesn't mean your data is useful for someone studying the whole population.
It is true that sometimes data collected for one purpose is not necessarily good for another purpose, however very often it is and many discoveries were made by looking for something new in data collected for a completely different purpose.
>This would mean that not only do you have to publish the raw data, but also describe your data collection methods, your assumptions, things you didn't check etc. How far are you from publishing your whole paper at that point?
Good reproducible science should do that anyway. You should always keep a labbook that clearly documents what you are doing so that someone else (or your future self), can reproduce the results. So in the requirement I'm asking for is forcing people to do good science. I can tell from experience this documentation is still very far away from a whole paper.
Many scientific studies obscure how many input subjects are rejected from final assessment. This would represent millions of experimental subjects, over time.
Suddenly a wildly extrapolated clickbaitykachu appears
Non-sequitur, I choose you!
Why to say millions when we could say thousands?