It does seem that the particular study was centered around a single electrically stimulated exercise session. It would be difficult to try to guess what happens if combined with different (more sustained? more repeated?) exercise.
If you're asking if this is generalizable to any other group? No idea. Probably okay to generalize to all healthy men. But that's about it.
As for why this is so small? Looks like they just piggy backed on another study. Pretty reasonable use of resources.
In short: typical alpha is 0.05, so a p-value < 0.05 is considered "statistically significant" and the results are not attributable to random chance in a t-test. An unpaired t-test compares two populations, and a p-value below the alpha level indicates that differences are not due to random chance.
Take everything with a grain of salt, though. Low sample sizes dramatically change how test statistics produce results. P-value hacking is the practice of eliminating data from an experiment to produce statistically significant results or otherwise altering the data by adding buffers to ensure that a result confirms a belief a researcher has.
Given a statement and a sample with a mean and standard deviation, what is the probability that you manage to choose a sample with that mean and SD, assuming the statement is true?
For example, my statement is the average weight of chicken an american eats every month is 20 pounds and my sample gives me a mean of 16 and SD of 4. The p value is the probability of choosing a group of people who eat on average 16 pounds of chicken every month if the REAL mean is actually 20.
Edit to answer question: The p value doesn't "care" about the sample size - it is adjusted and remains accurate. And so if the p value if high for the sample in the study, that means it is a likely occurrence and the result is not very noteworthy.
"Unlike traditional scientific publishing, in which manuscripts are peer reviewed only after studies have been completed, registered reports are reviewed before scientists collect data. If the scientific question and methods are deemed sound, the authors are then offered "in-principle acceptance" of their article, which virtually guarantees publication regardless of how the results turn out."
https://www.theguardian.com/science/blog/2013/jun/05/trust-i...
Let's say the drug studied was a guaranteed 100% limb regenerative drug for amputees. Would you really require a 200 person study to prove that Examplinol successfully regenerates limbs? I hope the answer to that is "obviously not; it will be very clear whether or not it works with a low sample size." Of course, how would we change the study if it wasn't supposed to work in 100% of people? What if Sample Pharmaceuticals indicated that it only worked for 50% of people? Or 10%? Or 1%? Would it suffice to use 200 people each trial?
Imagine you're investigating a drug for hair loss. You gather a group of 31 bald white men. 15 of them take a placebo and 16 take the experimental drug for 24 weeks. If, after 24 weeks, you find 1 of 15 controls with a full head of hair, and 14 of 16 test subjects with a full head of hair, is 31 too small a sample size?
Let me rephrase what you phrased:
Imagine you're investigating a drug for hair loss, like 20000 other researchers hellbent on achieving their goals. You gather a group of 31 bald white men. 15 of them take a placebo and 16 take the experimental drug for 24 weeks, but because your random allocation put more extra bald people in your non-control group you fixed it by balancing the two groups. Your counterpart that didn't have the same problem, uh, didn't make any adjustments. If, after 24 weeks, you find 1 of 15 controls with a full head of hair by some function you define, and 14 of 16 test subjects with a full head of hair again by some function that you define oh and we didn't adjust for anything but of course in the case where the drugs didn't do anything we did because the non-control group were all above average in both age and some other random things like testosterone levels. Is 31 too small a sample size?
31 is always too small of a sample size. These scientific studies are almost always bullshit and we all know it and we'd do something about it if we really cared but we don't because nobody is incentivized to fix it.
I recruited 31 white men who self-report having no hair on at least 500cm^2 of their head. They all had hair as adolescents, and lost it gradually throughout their 20s or early 30s. They're all at least 37 now.
I then sort them into the two groups by assigning sequential IDs in random order. Those with an odd ID are in Group A, those with an even ID are in Group B.
I then give each participant in both groups a placebo for 24 weeks, having them check in every 2 weeks to measure hair growth, if any. At 24 weeks, I randomly assign two new courses of treatment to each group. I do this blindly, giving Group A pill X and Group B pill Y. Another 24 weeks (with check-ins every 2 weeks) and I find that Group A (taking pill X) all have full heads of hair. Well, except for Paul and Maury.
Meanwhile, Tom, in Group B (who took pill Y), is the only member of his group to have a full head of hair.
Statistically-speaking, my n=31 was a great sample size, and I'm also pretty sure Tom discovered he had sugar pills and stole the real stuff from Paul or Maury. :)
To address the possibility of seeing these results due to a confounding variable (like above average testosterone), this is about as likely as shuffling a deck of 31 cards, and all the red ones ending up in the front and all the black ones in the back. About 0.0000003%
Clearly this is a sham and you needed at least 31^47 participants in this study you charlatan!
First, that's obviously not a rephrasing of what they said. They, very clearly, illustrated a valid case for a well-powered, low sample size study. They presuppose a case where 31 virtually identically-haired white men undergo a treatment and, voila, it works decidedly well.
Second, you're begging the question. You state that the drug doesn't work, yet have concocted a scenario in which it clearly and explicitly works. Even in your "cheatsy" example, 15 of the 16 people go from "bald" to "full head of hair." If the "full head of hair" metric is wrong it has nothing to do with the sample size, rather, we'll find this by examining how the study partitioned groups, pre-trial measurements and its claims to have calculated effectiveness over placebo (ie: actually reading the study).
Third, you've neglected statistical power (which is actually your argument here) for an incorrect statement about low-sample size bias. You state that the mechanism of action for this bias is a pre-partition between [truly bald, nearly haired] and then, given maybe a bit of luck, let good ol' father time take over and show, wow, the (useless) drug works! Maybe everyone experienced a 1.5x follicular density growth, but only because you've been able to presort the group do we see the non-placebo benefit more than the other... except this has nothing whatsoever to do specifically with sample size! If you can sort a trial with 31 people into 15 "truly bald" and 16 "near-haired" you bet you can sort a trial of 200 participants into 96 "truly balled" and 104 "nearly haired" or, heck, just don't even randomize your group selection at all! Or run the study 20 times and pick the successful one! Or any number of other tricks, which are real and meaningful and require someone to actually read the study to uncover instead of just looking at a clearly, very misunderstood, number.