Warning Signs in Experimental Design and Interpretation (2007)
norvig.com
norvig.com
My one comment about the article is that most of what gets submitted to HN as a breathless press release on a research "breakthrough" is often not even based on experimental research, but rather on correlational research, so the study goes wrong with problems that Peter Norvig's excellent article doesn't even discuss much. Many, many submissions to HN are based at bottom on press releases, and press releases are well known for spinning preliminary research findings beyond all recognition. This has been commented on in the PhD comic "The Science News Cycle,"[1] which only exaggerates the process a very little. More serious commentary in the edited group blog post "Related by coincidence only? University and medical journal press releases versus journal articles"[2] points to the same danger of taking press releases (and news aggregator website articles based solely on press releases) too seriously. Press releases are usually misleading.
But, yes, definitely read the submission here, as it will help you check each submission to Hacker News you read for how many of the important issues in interpreting research are NOT discussed in the submission.
[1] http://www.phdcomics.com/comics.php?f=1174
[2] http://www.sciencebasedmedicine.org/index.php/related-by-coi...
Can... can the whole Web be like this?
As part of internalizing the bit about P(H|E) vs P(E|H) [Warning sign I4], I wrote up a quick gdocs spreadsheet to let me play with the numbers: https://docs.google.com/spreadsheets/d/10JrG42iKY-LhcnaKU7O7...
Dear frequentists and Bayesians: could you people kiss and make up already? The rest of us are pretty tired of having to run bad numbers to get papers published and of having computationally intractable statistics to run. Please come up with a compromise.
At one end of the spectrum, we rely heavily on uncontrolled single- or multiple-ascending dose studies to prove to ourselves that a treatment is likely to be safe in next-step studies, and to guess at optimal dose. At the other, we learn a great deal from post marketing surveillance about unanticipated toxicities - because our priors based on big phase 3 studies may still be insufficient to accurately estimate risk. Neither of these designs are randomized or placebo-controlled - and in neither case is it an indication that 'something is wrong', even though an RCT would be better in both contexts. Better, if cost were no object and patient safety were not a concern.
I realize my fellow Brunonian does include some offhand caveats - but it worries me to read comments about how this negates most social science research.
As I said, I think this is a nice introduction but can lead one to discount, well, anything other than large randomized placebo-controlled trials.
I don't ever seem to see scenarios like that addressed.
Still good, nonetheless.
There won't be much discussion if no one can read the article.