Important things to check when it comes to experiments
- preregistration (it helps to counter fishing-for-signal bias, ie. maybe there were 1000 kids, and they only picked this 50 young girls, because this is where there's a big positive effect; also statistical significance levels, how much "sigma" is needed to say that the null hypothesis can be rejected)
- control group (here there's a nice direct comparison group, but there's no no-treatment group)
- blinding (did the kids know which one they were using? of course, but let's say it's not really relevant here, and they even mention that "no concealment was possible")
- power analysis (if we want to demonstrate a small effect we need a large sample size, power analysis can tell us was the experiment even capable of producing significant results or not, because it was underpowered? this requires an estimation of the effect size from the hypothesis)
- basics causal inference sanity checks (are they making claims/conclusions that are actually supported by the data? is the experiment even capable of showing any kind of causal direction? due to the lack of no-treatment group we don't really know if the decrease in the numbers is due to weather or the treatments, so in this regard if we want to be very pedantic, it's a bit of a stretch to say that the treatments caused the decrease, but it's true that there was a significant decrease)
- stats check (they are using "Wilcoxon matched pairs signed ranks test" ... which is good for exactly this kind of "is there an effect at all" hypothesis testing (and arguably better than a t-test, because this is ranked data, not pulled from some nice normal distributions), basically it says that how likely it is that the day 0 and day 30 numbers are "really different", and we see that they got p-values smaller than 0.001, which indicates that it's very unlikely that they are from the same distribution; and then there's also a test for comparing the groups, but I haven't checked that)
Also, we can look at the full-text to see the numbers: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5109859/
it's interesting that there's a typo in deviation in Table 1, but not in the others, but there's also one in Whitney in Table 6 :)
all in all seems okay, though I think a nice chart would have been good with error bars and effect sizes (because then you can see that the treatment arms leave the null hypothesis region, and by how much, like with this one https://cdn.the-scientist.com/assets/articleNo/69229/iImg/43... )
of course, it's just one study ... https://slatestarcodex.com/2014/12/12/beware-the-man-of-one-...