The basics here are very strong for me: run larger tests (which btw, are still too small) and the numbers become at BEST equivocal. Run small tests, you can get some "stunning" evidence it works.
Stunning, because 6/10 vs 5/10 is an apparent huge leap in efficacy, when in fact, its a single flip in the margins. 6/10 is probably 5.01/10 when you do it to 100+ test subjects. 60/100 vs 40/100 would of course be more interesting, but the likelihood of 6/10 converting to 60/100 is actually low: it was most likely a single flip, not a statistical flip of a cohort. Thats what the bigger studies appear to be telling us.