Almost entirely irrelevant, provided it's properly powered--and based on the effect size and p-values, it is. Furthermore, for an 18 month trial, 40 people is larger than most.
> probably not preregistered
I mean, true, and we should definitely promote pre-reg, but, 1. so few are preregistered that I'm not sure it's prudent to disregard a study based on this and 2. it's an academic study, neither a clinical trial nor a public-health decree--not a typical study type that necessitates pre-registration.
> Likely Publication Bias or p-hacking.
What? You're arrived at this conclusion based on... the fact that it's "small" and you couldn't find preregistration about it?
> I'm starting to get interested when it's been independently replicated.
Yeah, if only there were another trial that showed a positive cognitive effect of bioavailable curcumin... https://www.ncbi.nlm.nih.gov/pubmed/25277322
I mean, the title of this post is by far the most wrong thing about this article, as...
1. it's not turmeric, it's curcumin, which is about 2% of turmeric by mass (so, you'd need to eat ~4.5g turmeric daily to match study doses).
2. curcumin doesn't even have any reasonable bioavilability by itself, so they're using a bioavailable analogue.
Check the last few paragraphs--"The authors thank Shayna Greenberg, Dev Darshan Khalsa, and Anya Rosensteel for help in recruitment, data management, and study coordination; John Williams, University of California, Los Angeles, Nuclear Medicine Clinic, for performing the PET scans; and Vladimir Kepe, Cleveland Clinic, for developing the FDDNP-PET analysis procedures for proteinopathies in patients with neurodegenerative diseases." Add all that to the effort involved in the actual chemical extraction of the bioavailable curcumin analogue and it just doesn't make sense to add any more people than you need.
Edit: Heck, you can even lease top grade medical space in all major metro areas and offer it as part of the package in a WeWork style arrangement. I'm sure they are tons of nuances to rolling something like this out, but it sounds like it can cut costs substantially.
The University of California, Los Angeles, owns a U.S. patent (6,274,119) entitled “Methods for Labeling ß-Amyloid Plaques and Neurofibrillary Tangles”, which has been licensed to TauMark, LLC. Drs. Small, Satyamurthy, Huang, and Barrio are among the inventors and have financial interest in TauMark, LLC. Dr. Small also reports having served as an advisor to and/or having received lecture fees from Allergan, Argentum, Axovant, Cogniciti, Forum Pharmaceuticals, Herbalife, Janssen, Lundbeck, Lilly, Novartis, Otsuka, and Pfizer. Dr. Heber reports receiving consulting fees from Herbalife, and the McCormick Science Institute. The manufacturer of Theracurmin, Theravalues Corporation, provided the Theracurmin and placebo for the trial, funds for laboratory testing of blood curcumin levels, and funds for Dr. Small's travel to the 2017 Alzheimer's Association International Conference for presentation of the findings.
I don't' work in the field so I find it hard to interpret these studies. What's your opinion about the size of improvements in practice?
For example, How meaningful is the 20.3 point difference in Buschke SRT test?
Cohen's d (mean difference divided by SD) above 0.60 with the memory or attention seems good (73 % of the people who take the dose get above the mean test score). But difference that is detectable in a test is not telling me how big the effect would be in the daily performance unless I'm familiar with the test.
Typical Indian cooking uses turmeric in pretty much everything. I wouldn't be surprised if daily consumption of turmeric is very close to the figures you mentioned.
My kitchen medicine for running nose, soar throat is hot milk with turmeric. Works for me most of the times.
I've been mixing it through my oatmeal, together with chilli-pepper and black pepper, a bit of cardamom, and a teaspoon of real cinnamon (the non-cumerin kind). While it feels like it works, I know that the placebo effect is probably much stronger than any real effect these foods have. Still, a placebo effect is still a real effect on my mood, so that's still a kind of health benefit I guess. If nothing else the spices help me wake up!
Tangent regarding health food fads: these days I mainly use nutritionfacts.org[0] to determine which of those are actually supported by the latest nutrition research. It is the only health food/diet website that I know of that directly cites nutrition research and continuously scours the latest papers for new findings - sources are always linked, and quotes are directly lifted from the papers with no modification (so less likely to suffer from "stronger or opposite of what the paper actually says"-shenanigans often seen elsewhere).
Some of the other things I tried out based on that website have too big of an effect to just be placebo: adding blue-berries to a meal really reduces sugar rushes/crashes in the hours after it[1]. Taking a table spoon of freshly broken flax seeds[2] does a lot to counter the rise in blood pressure due to my ADD medication (I measured it, plus I feel a lot less discomfort in my chest area and less jittery).
Last time I mentioned that website here, it was in a discussion of which diets are healthy. Ironically, it was the only comment in the discussion that attracted multiples downvotes without explanation, while everyone else was sharing their unsourced opinions.
[0] https://nutritionfacts.org/
[1] https://nutritionfacts.org/video/green-smoothies-what-does-t...
[2] https://nutritionfacts.org/video/flax-seeds-for-hypertension...
http://www.ijamhrjournal.org/article.asp?issn=2349-4220;year...
> Worldwide, esophageal cancer is the eighth most common cancer [...]. In India, it is the fourth most common cause of cancer-related deaths.
Granted it's probably due to alcohol and tobacco but curry probably won't save you from cancer.
>Almost entirely irrelevant, provided it's properly powered--and based on the effect size and p-values, it is.
Sorry to state this so strongly but I feel it's important to point out that this is misinformation of the most egregious kind.
There is no examination of statistical power in this article whatsoever so there is no basis for your claim that this study is properly powered.
Large effect sizes are not evidence of adequate statistical power. Quite the opposite is true. E.g. See this article: https://www.nature.com/articles/nrn3475/figures/5
'Effect inflation is worst for small, low-powered studies, which can only detect effects that happen to be large.'
It is widely misunderstood though, so often people try to use an observed result to justify low power.
http://andrewgelman.com/2017/08/17/just-google-despite-limit...
"The power of a binary hypothesis test is the probability that the test correctly rejects the null hypothesis (H0) when a specific alternative hypothesis (H1) is true. The statistical power ranges from 0 to 1, and as statistical power increases, the probability of making a type 2 error decreases. For a type 2 error probability of β, the corresponding statistical power is 1-β. "
So according to that, if I always reject H0 whether rightly or wrongly, that means β=0 and power=1.
Sorry to state this so strongly but I feel it's
important to point out that this is misinformation of the
most egregious kind.
There is no examination of statistical power in this article whatsoever so there is no basis for your
claim that this study is properly powered.
The power is only relevant when you fail to detect an effect. This is a positive study, so the power is simply not relevant to criticizing that particular study.The nature article you reference talks about literature in aggregate, and the problem created by small studies in general, but it does not allow you to say anything about a single positive study. In fact, the place where those very authors draw their primary conclusion, they state:
Inflated effect estimates make it difficult to determine an adequate sample size for replication
studies, increasing the probability of type II errors
Which means: when you try to replicate it, you might fail because you thought the real effect size was bigger than it is, based on your estimated effect size with small n.That's well and good, but is a problem for the replicating authors, not for these.
An even more direct point related to the current article: it was preregistered, so you know the authors have not been farming well-designed but small curcurmin studies until they did one than that was positive by chance, and then published it. The act of pre-registration goes a very long way toward addressing the issues raised by Button et al.
This is simply not true. It is sometimes called the fallacy of "what does not kill my statistical significance makes it stronger”.
See http://andrewgelman.com/2017/08/17/just-google-despite-limit...
The power is only relevant when you fail to detect an effect.
This is absolutely false. Small, low-powered studies are more likely to produce false results than larger studies. This was well-documented by John Ioannidis over ten years ago, in work that should be required reading for all publishing scientists:http://journals.plos.org/plosmedicine/article?id=10.1371/jou...
At baseline, they performed an extensive neuropsychologist test battery, including a bunch of subtests:
-Trail Making Test A
-WAIS-III Digit Symbol Substitution
-WAIS-III Block Design Test
-Rey-Osterrieth Complex Figure Test (copy)
-Trail Making Test B
-Stroop Interference
-F.A.S.
-Buschke-Fuld Selective Reminding Test [SRT]
-Wechsler Memory Scale-3rd Edition [WMS-III]
-Verbal Paired Associations I
-Benton Visual Retention Test
-Buschke-Fuld SRT
-Rey-Osterrieth Complex Figure Test [recall]
-WMS-III Verbal Paired Associations II
-Boston Naming Test
-Animal Naming Test
The outcome measures they report are: -TMT-A
-SRT
-BVMT-R
I would be _very_ surprised if those were the only tests they performed at followup after a 18 month (expensive) clinical trial. If these measures had been specified as the only outcome measures of interest, noone would question the results. But, when it was not registered, and when no other measures are reported, not even in the supplementary, the reader is left to wonder why that is.At least a practice to include that in the abstract
[0] https://www.biorxiv.org/content/early/2016/07/29/066803
[1] https://cran.r-project.org/web/packages/scifigure/index.html
Chocolate has the Mars Company and other Big Sugar companies behind it pushing for justifying chocolate as a "health food" in the public mind. It is very obvious that they are wiling to distort the truth for better sales; people who can ease their conscious about eating unhealthy sweets buy more chocolate[0]. There might be some real effect to eating raw cocoa, but that does not really translate to your average chocolate bar with insane amounts of fat and sugar.
Tea/Coffee is kind of similar, as would be the "drink x glasses of water per day"-advice, which is a distortion of science by bottled water companies (fruits, vegetables and other food contains a ton of water already).
For this reason I also don't buy that salt is safe: there are very strong vested interests in the food industry that want salt to be (considered) safe to sell more food, whereas I cannot see a comparable bias on the "eat less salt"-side.
For this research the situation seems to be have some of the mentioned problems, but not to the same degree:
On the one hand, that tumeric has an in-vitro effect is pretty well established. The issue is that the normal form we put on our food is probably not absorbed in a dosage that has a significant effect. However, this was a trial with a bio-available form, in a high enough dosage, so that obvious flaw is out of the picture.
On the other hand, that bio-available form is a product, so the company behind it wants it to be effective. Still, Big Pharma is under more scrutiny than Big Sugar, so I have cautious hopes this might actually work out in the end (and it won't be sold in the form of calorie-rich sweets at least).
Admittedly, I'm biased here: my grandmother on my mother's side, and my grandfather on my father's side both had Alzheimers (or something very similar to it at least). My parents are now their late sixties, I would really like to see them stay healthy and happy for a few more decades.
[0] https://www.vox.com/science-and-health/2017/10/18/15995478/c...
Another argument in favour of eating potassium and less sodium, which is pretty well established:
https://nutritionfacts.org/video/lowering-our-sodium-to-pota...
https://clinicaltrials.gov/ct2/show/NCT01383161?cond=NCT0138...
It's great to be skeptical, but it's really lame to be skeptical and lazy. It's a terrible combination that causes patients not to trust their doctors.
Statistically, there is no such thing as a 'small sample size', in isolation. It is always related to the power and effect that the researchers want to work with.
Moreover, even if a study is statistically insignificant, it DOES NOT mean that the opposite conclusion is true. It might often merit some further examination of cultural and historical trends. In this case, for instance, turmeric has been used for medicinal properties in Asian cultures for centuries. Theories of evolution suggest that this wouldn't be the case (and it would've died a natural death), if there wasn't some merit.
Lots of things that are completely useless have been used for thousands of years for their medicinal purposes. The evolution argument only holds if there were some negative effects here. Anything that doesn't do harm to the body can easily integrate into the huge body of knowledge called superstition. Eating it is tasty and doesn't hurt, there's nothing more needed to explain that Asian cultures have it as a medicinal herb.
If it for example was good for you but caused somewhat unpleasant side-effects, your evolutionary argument might have more merit. Anything that works as well as homeopathy (it doesn't), but doesn't harm either (like homeopathy), won't be sorted out.
These assumptions are more relevant in small sample sizes. Making them inherently unreliable. What's missing is not a p-value, but an estimate for the uncertainty around that p-value.
That is utter nonsense. To give just one counter-example, bloodletting was used for millennia for conditions where it’s completely ineffective (and, in fact, actively harmful: bloodletting has killed scores of people). In fact, you simply can’t apply an argument from evolution here. Evolution doesn’t optimise for “goodness”. It optimises something (and only over the long term). That “something” can often be positively detrimental to other outcomes, and even to general fitness (e.g. the laryngeal nerve).
Here is a great demonstration of how correlation does not imply causation: http://www.tylervigen.com/spurious-correlations.
I believe the research I was thinking of was specifically this part of the overview above.:
Various studies and research[9,10] results indicate a lower incidence and prevalence of AD in India. The prevalence of AD among adults aged 70-79 years in India is 4.4 times less than that of adults aged 70-79 years in the United States.[9] Researchers investigated the association between the curry consumption and cognitive level in 1010 Asians between 60 and 93 years of age. The study found that those who occasionally ate curry (less than once a month) and often (more than once a month) performed better on a standard test (MMSE) of cognitive function than those who ate curry never or rarely.[10]
Sources:
9. Pandav R, Belle SH, DeKosky ST. Apolipoprotein E polymorphism and Alzheimer's disease: The Indo-US cross-national dementia study. Arch Neurol. 2000;57:824–30. [PubMed]
10. Ng TP, Chiam PC, Lee T, Chua HC, Lim L, Kua EH. Curry consumption and cognitive function in the elderly. Am J Epidemiol. 2006;164:898–906. [PubMed]
"Curcumin seems to be somewhat more effective than placebo in reducing symptoms of depression. It may take 2-3 months to see any outcomes. Skepticism is warranted though, as the studies comparing curcumin to placebo were not well designed and produced effect sizes not too far apart, even though the differences were statistically significant."
If you meet statistical significance, then diminishing n means increasing clinical significance.
That’s only true if it holds across multiple studies. Otherwise it might simply be random variability. By your argument, a study with n=4 that showed an effect would imply a huge effect size. No. It implies that the measured signal is highly variable, and that you got lucky on 1:3 odds. Check out regression towards the mean [1].
(Note that if you’ve got a good underlying model of the variability, n=4 can work. In fact, it’s routinely done in specific applications, such as transcriptomics. But even there it’s far from ideal, and it’s supported by a stringent error model.)
[1] https://en.wikipedia.org/wiki/Regression_toward_the_mean
Significance is the probability of the result being coincidence. So "highly significant" only means "most probably not coincidence".
That means you can have a "highly significant" result where the actual result (= effect size) is tiny.
In study design, both are related: The higher the number of test persons, the easier it is to archive significance. Therefore, proper design would be to first guesstimate (based on existing literature etc.) the effect size, then define the minimum number of test persons necessary.