Also, doesn't their suggested approach amount to multiple testing? In other words, a kind of p-hacking:
https://en.wikipedia.org/wiki/Multiple_comparisons_problem
Edit - and this: http://www.stat.columbia.edu/~gelman/research/unpublished/p_...
Edit - and this: http://www.stat.columbia.edu/~gelman/research/unpublished/p_...