There's no aspect of this that would be statistically distinguishable from noise.
Then again, the original study suffered from this as well.
n=11 may be succinct, but it is not shallow. There is a world of statistical DoE, analytics, etc. behind this.
My concern is the rush to headlines about studies without any real resolving power. The bad part of this is that we don't have time to do proper studies (this takes months and many people to make work). The n=11 comment actually infers this issue.
Sort of like -1 = exp(sqrt(-1)*pi), you don't need many letters to convey a tremendous amount of meaning.
Something not shallow would be like "only eleven people participated in this study, this is too little to get a clear result that's not due to luck."
Here's a comparable comment that at least begins to engage with the specifics: https://news.ycombinator.com/item?id=22798031. I don't know if it's a good point, but it's at least actually about this particular study. We wouldn't moderate such a comment.
11 is very small, and useful only for evaluating the presence of very very prominent effect sizes, regardless of domain. A comment saying n=11 is not a novel insight, but I believe it is useful, similar to people who add "in mice" to articles whose headlines might imply the test was on humans. I appreciate those comments.
In this case I think it is fair to conclude that the headline "Severe Covid-19 Cases Don't Respond to..." is an accurate takeaway purely because the sample size is so small. With this sample size you can, at best, conclude that these combinations of drugs don't reliably cure the disease in severe cases, which is a very different idea.
While I have you: could you please stop creating accounts for every few comments you post? We ban accounts that do that, and in fact this account was banned. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.
You needn't use your real name, of course, but for HN to be a community, users need some identity for others to relate to. Otherwise we may as well have no usernames and no community, and that would be a different kind of forum. https://hn.algolia.com/?sort=byDate&dateRange=all&type=comme...
Obviously there are statistical concerns with small samples. Equally obviously, not every such study is worthless and the authors are not all idiots who knows less than any internet commenter. So to get substantive discussion, we need to engage with the specifics of a particular study, not post a generic dismissal. "n = 11?" looks specific but it isn't. It's a template instantiation, as stock as "Correlation is not causation".
The answer to "the rush to headlines about studies" is not the opposite reflex, it's to slow down and reflect, and then maybe post if one has something thoughtful to say. https://hn.algolia.com/?dateRange=all&page=0&prefix=true&que...
I don't mean to pick on the GP commenter. Maybe they had something else in mind. The issue is how such a comment plugs into a known shallow-discussion dynamic. The problem is at the group level, not the individual, and we all suffer from this reflexiveness, so none of us gets to feel smug.
Generic discussion is shallow discussion.
Imagine you have a purchase form that converts at 0.1%. You test a new change, and with the first 22 users (11 in each group), if the control group randomly gets a conversion, it will have AT LEAST a ~9% CVR now. Okay, maybe you ignore that data, since you can just compare to historical data. Whatever.
Even if the new form is dramatically better and has a 30% CVR -- There's a 3% chance it would show a ~9% CVR also, and about the same that it would show a 0% CVR. Similarly, there's about a 6% chance it would show a 91%+ CVR -- skewing your results just as badly in the opposite direction. There's only like a 30% chance you get close to the "true" CVR.
Unless these hydroxycloroquine is close to 100% effective (and that doesn't seem to be what any of these studies say) I just don't understand how there's any confidence with small sample sizes.
Does anyone have a link to how this works? I've asked a few friends in medicine, and they haven't given a good answer. Haven't found anything on Wikipedia. Probably don't know what to search for.
If an A/B test took a form from almost 0% conversion to almost 100% conversion, you might only need 11 users to show that. But even in that case, the devil is in the details. Were your users really a completely random sample from your population, etc?
The other side of the coin is, can we say these are non-interesting treatments with only 11 subjects? Or are the studies too underpowered to say much of all?
Depending on what we mean by "back to life", a trick is in fact much more likely than something thought to be biologically impossible.
If we eliminated the possibility of a trick -- that is, proved beyond a doubt that 11 people were dead, and then alive -- I would absolutely not be ready to attribute it to some new treatment. Natural law is being violated here! I would instead start checking for other possibilities, like:
- Are other people spontaneously reanimating around the world, or just in this study?
- Did these 11 people have anything else in common? Extraterrestrial origin, perhaps? Unusual religious beliefs?
- Hang on, let's double (triple, quadruple, etc.) check that they were really dead. And then go back and check again.
- And by all means, start a proper clinical trial of the new treatment while we're trying to understand how resurrection is even possible.
A bit of self-promotion: https://news.ycombinator.com/item?id=22787162
In fact, I am impressed, this is not a thing yet.
I mean, that plus Zoom and you're like halfway to a world-wide sample size, more-or-less
On a serious note: I'm extremely grateful for everyone who's working to solve this big and complicated problem, and (like the other poster said) it's entirely reasonable that it's really difficult to get larger studies going.