Evidence of Fabricated Data in a Vitamin C trial by Paul E Marik et al in CHEST
kylesheldrick.blogspot.com
kylesheldrick.blogspot.com
As for this particular paper about the Vitamin C trial (which hasn't been retracted yet): numerous other groups tried replicating it and none could. Millions and millions and millions of dollars, countless hours wasted on conducting controlled trials all because of this guy's fraud.
My two takeaways:
1) Kudos to Sheldrick, we absolutely need more people like him. We need statisticians to pick apart data in papers and see if it smells funny. This should have been caught earlier.
2) There should be a higher price to pay for this kind of fraud. For the confusion and mistrust that ensued, for the money and human capital that was wasted, he needs to be behind bars for a long, long time.
Also, if we really are fond of consequences, frauds carry over the price of having to confront, for future publishing, with a scientific community that has grown a bias against the frauder; and the latter will have to regain trust if he/she wants to publish again.
The criminal charges should be for those who use unproven scientific results to endorse their own foul agenda.
If a system allowed me to publish fraudulent research or publish bad research recklessly, with the only consequence being "public shaming" (which is sometimes not really consequential if one has the right connections to the right people), what's to prevent me (and everyone else) from doing it over and over again?
TBH from a legal perspective knowingly publishing false information for personal gain is already fraud. You're arguing that authorities should stay away because academia should somehow have immunity? That's a very disturbing line of thought.
At the moment, University incentives are to protect the prestige and not for honesty.
There should also be a University list of shame, based on author affil of papers accused of fraud.
The trouble and defense will be that any rando can accuse a paper of fraud, and instead one should use finalized decisions (such as retractions). That would cause delays and not lead to any resolution in many cases, and I doubt that the number of papers false accused of fraud based on any substantial statistical analysis is large.
Karma would help. One could then keep track of people who falsely accused papers of fraud in the past, and correctly pointed out papers which were later retracted, and then weigh the University shame list based on that.
> There should also be a University list of shame, based on
> author affil of papers accused of fraud.
The ad-hominen attack is generally considered a form of fallacy. And many significant scientific discoveries were discovered by then-considered-quacksThat said, this particular author is egregious and I'm sure his infamy will precede his name in future papers.
That's how science work. No matter what we do to discourage fraud, we'll always invest tons of capital and time to find and filter it out.
Sometimes it is difficult to distinguish between both but of there is a track record of fraud no reputable journal should accept more papers from that source.
Edit: As I re-read my comment I worry because research should not really depend on "prestige", it should be based on cold facts. On the other hand I don't see any other way to deal with malicious actors.
It's absurd to think there will be a way to know which studies need reproduction or not. The whole purpose of reproducing is to find out which ones are true and which are bogus
But it was known not to be true, and so people were sent under false pretenses to go and reproduce and verify it. That is far from well invested.
We will always need to invest in finding error and fraud, no matter how we perceive scientists honesty.
The moment we stop doing these investment, we're doomed.
Unpossible!
We need to reproduce every single study, including the fraudulent ones. Gosh, ESPECIALLY the fraudulent ones.
But we can't possibly know in advance which are which. That's the whole purpose of trying to reproduce.
Without reproduction efforts, scientific papers are of no value, as we can't trust them.
In this sense, a paper without reproduction effort is not a "window" in the analogy. Only after reproduction efforts it becomes a window.
So saying it's a waste to try reproducing an article is just absurd. It's like saying it's a waste to produce the "window" in the first place because a kid might break it in the future (in case of the paper, that we might figure out the paper was fraudulent).
It's still valuable to find out a paper was fraudulent. It builds trust in the whole scientific endeavor.
The overall process assumes trustworthiness. The better that assumption holds, the better science performs.
But it's reasonable to think that some reproduction is better than no reproduction.
The point I think folks were contesting was the tone some may've inferred from [this comment](https://news.ycombinator.com/item?id=30787447 ), which may sound a bit like downplaying the damage caused by fraud under the premise that reproduction ought to catch such things anyway.
For a computing analogy, TCP uses mechanisms that enable it to tolerate packet-loss. But packet-loss still slows things down, limiting bandwidth and increasing latency, until a connection would be impractical. For a connection with many hops, even a small packet-loss-rate along each hop would be a huge problem, even though TCP actively checks for packet-loss and tries to fix it.
Likewise, science-with-replication presumably ought to be able to fix a lot of fraud and errors -- but those problems take a heavy, debilitating toll on scientific progress; it's not a small problem that replication-studies alone could satisfactorily rectify.
The real damage would have been everyone believing the paper conclusions were solid and start to act according to this false belief.
Reproduction efforts saved us from that damage.
These efforts were not the damage.
Whoever is seeing differently doesn't understand one of the basic foundations of how the scientific endeavor works.
And again: anyone expecting a world where there are no dishonest scientists or where the dishonest ones call themselves out in advance to save from the "damage" of people spending time reproducing their papers, is just playing wishful thinking, which is not a recipe for success...
I'm pretty sure everyone firmly agrees with this.
The general presumption would tend to be that other scientists are trying to do good, honest work. Scientific progress does better when that presumption turns out to be more true.
To be clear, I was responding to their apparent disregard for the harm done by the original alleged-fraud.
You can't flag "this scientist is honest, never reproduce anything from him anymore". He can turn dishonest the next day you do this.
If there are more dishonest scientists, we'll just find more useless papers. But the time to reproduce will be the same.
So I agree with OP the time wasn't wasted, but it could have been spent on better opportunities.
Happily from the sidelines - it probably is fraud and this isn't really an article targeted at the general audience. But the statistical evidence proves beyond a doubt that the data wasn't collected by the methods described in the paper. That doesn't prove it is fraudulent because formally speaking he can't rule out that they just did a very bad job of explaining their methods.
Just from life experience, any time anyone goes in guns blazing based on statistical evidence - without actually talking to the people involved - it is often a mistake. Just because someone can't think of an alternative reason says more about their imagination than the reality of the situation. Particularly coming in hot off a twitter argument.
That being said, these are unusual and unexpected patterns that show irregularities in the data, and I would certainly share his concerns.
As the author describes, the p-values listed are effectively impossible to get from real data. Because of the magic of mathematics, we know exactly how p-values are meant to behave when you calculate them repeatedly - uniform random between 0 and 1. There are a battery of tests you can use to interrogate this very well understood distribution. Those tests, as the author states, prove without any doubt that the p-values in table 1 are effectively impossible. You can't make it to those p-values by mistake.
That does not mean fraud (although, I stress, in this case I can see why it is probably fraud). For example, maybe the study is done in some sort of hospital with some sort of patient scheduling system and the front desk is playing a funny joke on the doctors in this study. It is quite common for experimenters to not be as much in control of their environment as a paper suggests - that would be an honest reason for why the data is weird.
In a more plausible example - I don't know how "matching" works, but maybe the patients were matched, the report writer didn't know that and the sentence slipped through proofreading. Unlikely, but certainly possible. That mistake would happen in the wild from time to time.
You don't need much imagination to find examples other than fraud. Although it does take a little imagination, which is concerning. But, frankly, there should almost always be an attempt at dialog with someone before accusing them publicly of being a fraud. Unless they're going to shoot you or something extreme.
Sloppy research methods leading to bogus results are can be construed as fraud, particularly if you don’t know the motivation for the lack of care.
A diligent researcher might take these results which were “incredible” and attempt to replicate them before publishing.
He's doing that based on an interpretation of some study, which I suspect is approximately a 20 page document.
Now on the one hand, he is probably right. On the other hand, this is a very brave approach - the damage and the care taken aren't in good proportion here. This sort of thing is easy to get wrong. He should have talked to the paper's authors first.
The author published an email he also - before publication - sent to the people involved.
Science happens in public for a reason. The reason is that others should be able to check the results being published. Authors should therefore be very aware of the fact, that their results might come under scrutiny.
If one were to publish claims that look so unlikely based on valid scientific methods they should at least address the unlikeliness of the results and provide a rationale why this is still valid, even if it looks unlikely.
That didn't happen - and so they get called out. The extreme unlikeliness is what validates the claim of this being fraudulent data. Now it is on the original study`s authors to react to that public answer to their publication.
Sheldrick did exactly that. It is the first line of the article.
"Below is an email I have sent to Sentara Norfolk General Hospital, the editor of CHEST Journal, and prof Paul E Marik"
That's a dangerous statement. I'm pretty sure I got p-values of over 1 or under 0 by mistake. People will make every mistake you think is impossible. (Not that I think it's a mistake here)
In the best case of "poorly described methods", we're talking about authors who blatantly lied about how they sampled their patients and deliberately omitted what would have been an extremely laborious process, involving many thousands of extremely non-consecutive patients, to produce an extremely well-matched treatment and control group.
You've got your choice as to the exact nature of the enormous lie or fabrication, but there's no possibility of innocent omission.
As noted in other comments here, one possibility is that the authors failed to correctly describe the procedure they used. Sheldrick does consider this, pointing out that even if they actually did try to match subjects, unlike what they said in the paper, it would be very unlikely that they could match them well enough to obtain the reported p-values. But another possibility is that they were totally incompetent at using their statistical software - that the reported p-values are not actually the real p-values for the tests that they said they did. Maybe they accidently used the wrong dataset, for example. Or perhaps they copied numbers from the wrong column of the table the software produced.
In a wider sense, though, this would still essentially be fraud - it's dishonest to publish a paper that purports to be a competent study when really you are totally clueless about how to do research, or care so little about correctness that you put no effort into checking that you did things correctly. (And surely they would have to suspect that they don't really know what they're doing...)
Something like a column duplication would have resulted in perfectly identical data between treatment and control groups, which this data does not show.
The best case scenario is an errant copy-paste or Excel formula that somehow presented both groups of data in a way that actually only looked at one of the two groups, but also somehow introduced small differences along the way.
How exactly this could happen is beyond my imagination, but perhaps it introduces a sliver of doubt as to whether the paper is fraud by intention or merely by negligence.
“Remember that Vitamin C cures sepsis paper that could never be replicated in 9 RCTs?
Turns out there is a good reason why: it’s very likely fraudulent.
More brilliant statistical sleuthing by @K_Sheldrick.”
https://twitter.com/nickmmark/status/1506402015936581635?s=2...
The outcome is that if you are honest it's hard to get funded because your results look bad.
Then, if we talk about honest results only, crappy statistics also leads to crazy optimistic effect estimates. Regularizing priors and pooling should be the norm. Andrew Gelman posts a lot about this: https://statmodeling.stat.columbia.edu/
This is flat out fraud, and if the research used any public funds the guilty party/parties deserves criminal prosecution.
That sort of prompts the question though: Should it be a crime for any scientist to deliberately publish falsified data?
Obviously there would need to be a very high bar for proving that it was deliberate and not an honest mistake (or even complete incompetence), and we wouldn't want juries to have the power to stop new scientific theories emerging, but other professions are subject to criminal penalties if they fail at their job.
I guess academic fraud is fraud in the legal sense. The problem is that the 'fraud' statute requires 'damages' and 'intent' to have standing, and they're difficult to prove in academic fraud.
That role is reserved for the tribunals. Judges are supposed to determine what is protected speech, and what is not.
I'd go with maybe based on the context. If you take money, such as grants, etc, for studies, then lie about your results that might be fraud (IANAL). If you use those incorrect results to apply for more grants for further study that is definitely fraud. Padding your CV with publications where you lied about the results may not be fraud but would seem the same as any other lies on a CV and could justifiably result in you getting fired.
I'd not be worried about "honest mistakes" so long as a juries are involved, with respect to specific studies with direct, actionable consequences.
So there would be a difference between some nut jobs proposing a flat earth theory (unqualified), which they simply can't reproduce, to running a medical trial where potentially the medical technicians, MDs, surgeons, etc. might cite and apply a study.
Its fine to have some crazy theory so long as you are not going to present it as some actionable truth - if you publish a study and show negligence and cannot reproduce, then it is known as such. You are welcome to incomplete, possibly outright incorrect theories, so long as they are academic and not cited as mainstream "technology". Its when your bullshit becomes measurably impactful, is when you decide to impose consequences.
Not sure how to write that up in legal terms, but.
That's not really different from most crimes... sometimes being hard to prove doesn't mean we just make it legal.
(As my username suggests, I have published scientific papers where I'm the corresponding author. I've also been added to truly shitty but honest papers against my will or knowledge, in cases where I did some of the work but left before the paper was written. I would never have approved the papers that were published had I been involved. This is an interesting wrinkle if criminal liability is on the table!)
You'll notice engineers tend to be scarce in arenas such as politics, it's safe to say people accustomed to and willing to face repercussions are a poor fit.
That's exactly why I'm arguing for raising standards in academia. When you publish a paper with a traditional publisher, you have to sign a bunch of stuff. You are aware that you are literally signing off on something and claiming responsibility for it.
In my opinion that responsibility should extend to legal liability for fraud with criminal penalties if mens rea can be established.
And, yes, I'd freely apply that standard to every paper I've ever published. Because they might or might not be wrong, but they're not fraudulent!
Plenty of ~intelligent people on HN and elsewhere, who would never sign up for a Gwyneth Paltrow detox regime, talked themselves into believing that there was some forbidden knowledge being exposed by these quacks. It was always so transparently silly.
I'm afraid that this is thanks to political tribalism. Especially in the system where you have only two sides.
"Strong convictions loosely held" is an admirable way forward on testing everything we have to save lives. Unfortunately when it became so politicized and the unethical doctors and scientists realized they could print money by claiming they had a secret cure being suppressed, the "loosely held" bit was left behind.
And imnsho a bigger problem: that they were in turn trying to dupe others. This amplification effect is at least as harmful as the initial fraud.
Also people really want to "own the libs" at any cost.
I’ve had this issue, arguing with bad faithed opportunists, whose approach was to just bury competing ideas in FUD.
Seems someone summarized it in this succinct statement:
[“The amount of energy needed to refute bullshit is an order of magnitude larger than is needed to produce it.”][1]
Good luck to this man!
https://en.wikipedia.org/wiki/Replication_crisis
"A 2016 poll of 1,500 scientists reported that 70% of them had failed to reproduce at least one other scientist's experiment (50% had failed to reproduce one of their own experiments)."
or at least - this is what it used to say! This has been memory holed.
and
https://www.bbc.com/news/science-environment-39054778
"According to a survey published in the journal Nature last summer, more than 70% of researchers have tried and failed to reproduce another scientist's experiments."
If you look at this page, you feel as if you have stumbled into some psychological/sociological discussion.
This, to me, is an example of how wiki editors are trying to dampen down the implications of the replication crisis.
Take a look at the current version against previous versions, eg: https://en.wikipedia.org/w/index.php?title=Replication_crisi...
The strong "Overall" statement early in the article is gone, and instead we have lots of dull, makeweight context and background.
Most likely, one report-writer didn't understand that these were MATCHED consecutive patients. The data is so clearly matched, it's hard to believe it's fraudulent.
This is sorta done for meta-analyses that aggregate the results of several studies systematically. I think there are people trying to automate more of it.
I'm 100% convinced hydrocortisone is bioactive is sepsis patients, like it is in anyone else . Also 100% convinced vitamin C boosts antioxidant levels in sepsis patients who're vitamin-C-depleted.
This isn't up for debate nor is it absurd to call for genotype specific studies of efficacy. We, collectively, have allowed real miraculous discoveries to flounder as a result of this failure to appreciate personalized medicine.
We will look back on this Era in awe with the ineptitude of our methods.
Funnily enough, if you should just go online and check the typical duration of a visual migraine, the results may shock you!
(where the placebo effect begins)
Is the study fraudulent? Could be, and his fight to use Ivermectin on Covid patients always struck me as quackery as well. That said, it seems to me that there is something about the Vitamin C protocol that should not be dismissed.
And this is exactly why scientists rightfully ignore "single data points" such as this. People get sick and recover all the time. And sometimes people get very, very, very sick, and they still recover on their own.
So there would be a non-trivial percentage of people who would have gotten very sick, and just based on timing it wouldn't be unusual that they would be given some treatment and then credit that treatment when they recovered.
Not only is this study blatant fraud, nine other studies tried to replicate the results and showed no benefit to sepsis outcomes.
The same applies for quite a wide range of other nutrients as well.
And as a statistician I can tell you with absolute certainty this paper is fraudulent.
Apparently this test is how likely it is for the same number of people to have the primary diagnosis in each group because all the 1.0 have equal number of patients with that diagnosis or off by 1. Like if they said, we have 8 people with COPD in the first group and then waited for COPD patients to show up at the ER for the second group until they had 8 then you would get a perfect 1.0 match since they made that variable the same? And then the other variables would be randomish, but maybe correlated so more than 0.5?
Maybe you can explain how this layman take is wrong, but if all that's going on is they selected patients with the same diagnoses and didn't make that clear in the paper then I don't understand what the big deal is.
Did you not read this? I asked whether if they did this they would come up with the perfect 1 scores on those items. Whether they actually did or not is a separate matter.
What troubles me more than whether this study was faked or not is the certainty HN readers have that it was without being able to answer simple questions about why they believe that. Is this test just measuring the likelihood of each group having the same primary diagnosis? If they added people to the second group so they had the same number of a primary diagnosis, would that result in a 1.0 p value on this test?
These should not be difficult questions to answer for somebody certain that fraud occurred.
2. Patient subgroup counts matched perfectly not only on COPD, but on about a dozen other conditions. Difficulty of matching subgroup counts grows rapidly with number of dimensions. To match subgroup counts for Group A and Group B near-perfectly on a dozen different dimensions, there's no substantially easier method than just one-to-one perfect matching, which requires a very, very large number of patients. You have to take Patient A1, who has subset X1 of 12 different conditions, and find Patient B1 with that exact same subset. Then repeat, 47 times in this case.
It is already quite hard to find the person B1 who has the exact subset X1 of conditions that person A1 had. For example, if there are 12 conditions, each condition is present in half the people that come into your clinic, and the conditions are independent, you'll need to go through 2^12 = 4096 people, on average, to find another exact match. The conditions may be a bit correlated but this can only help you so much when you're talking about 12 different conditions.
To repeat that feat 47 times is very hard. You'd have to churn through 10s or 100s of thousands of sepsis patients to get your matching subset. This would require access to an enormous pool of sepsis patients and constant reporting of all the conditions they have, that you want to match on. For this to be done without any mention in the paper is utterly beyond belief.
And the fact that the counts match near-perfectly in an off-by-1 fashion does not help the situation at all.
> Difficulty of matching subgroup counts grows rapidly with number of dimensions. ... and the conditions are independent, you'll need to go through 2^12 = 4096 people [for 12 conditions]
Except many of these are "primary" diagnoses, which presumably you have one of hence the name, so it's not possible for a person to have both a primary "Pneumonia" diagnosis and a primary "Other" diagnosis. So for each person maybe you're matching 2^3 or 2^4, not 2^12 - in any case, far less than your conclusions are based on.
> You'd have to churn through 10s or 100s of thousands of sepsis patients to get your matching subset.
So if the matching was 1/1000th as hard as you thought it was, that would be 10s or 100s of patients. A million cases a year, several doctors in major metro areas, there's probably 10s to 100s of patients at any given time in their hospital systems.
> Consecutive means you don't skip anyone who has the condition you're trying to treat.
Sepsis is very serious and common, so they'd have to treat them all basically simultaneously with the experimental treatment. I'm no expert, but I'd expect they'd want to be able to abort the trial if people started dying from it. It also doesn't even make sense as a study, because you want to test like for like as much as possible; maybe the treatment works fantastically on patients with cirrhosis and nobody else.
I think a plausible scenario is each day before normal rounds they did a med search for a patient matching the next one from first group, generally found a good match, worked with their doctor to change the treatment and added them to the study. Maybe two in a day, or skipping a day, and a month later they have really good matching data. Lots of cases to choose from, many exclusive variables. What do the statistics say for a charitable interpretation? Pretty good odds, right?
So maybe they didn't report their methods accurately, maybe it was assumed from domain knowledge or was a mistake when they cut and pasted from a template. Could be fraud too, but that seems like a huge leap to be certain of.
What you describe is selecting the group in advance so they match (though you use only one variable). There are cases where doing this selection is reasonable and is useful. However, the study describes a different selection process. If the selection process described was not the one actually used, what else in the study does not match reality?
If you trust anecdotes, your conclusions will be wrong most often than right. You might be lucky here and there, but will lose in the longrun.
Science is about being right most of the time and winning in the long-run.
However people get more than a bit overboard and woo-woo about Vitamin C.
Linus Pauling was a chemist and was obsessed with it but at least he was scientific about his quest.
https://en.wikipedia.org/wiki/Linus_Pauling#Medical_research...
and it gave us the Linus Pauling Institute which is a great thing for factual nutrition, I use it often as a reference:
I have never heard of this practice, and can find no evidence of it online. Can you provide a citation?
> If you consume vitamin C via citrus fruits like oranges, then you are making a more hostile environment for bacterial infections.
This is utter nonsense. The pH of a living organism is tightly regulated, and a failure of this regulation can be rapidly fatal. Consuming small (sub-gram) quantities of vitamin C does not have any meaningful effect on blood or extracellular pH.
Translated for laypeople:
This is BS. The pH of your body is constant and your body has mechanisms to make sure it stays constant. If your pH changed as claimed, you’d die.