The statistical significance threshold usually used is p<0.05, meaning that something is (generally, this is beginning to change since the replication crisis) considered to be a real discovery if it has less than a 1/20 chance of being a false positive under the chosen model.
As soon as you start trying multiple hypotheses, then that 1/20 chance of being a false positive begins to become meaningless. If you can just keep rolling d20s until one of them comes up with a critical hit, then you can easily generate false positives that still look very robust.
This is exactly the sort of bad science - p-hacking, fishing expeditions, and the garden of forking paths - that led to the replication crisis. (And that makes sense, as this paper is from 2013, and predates the widespread discovery of the crisis)
It is also why we see repeated, spurious insistence that anti-depressants don't do anything.
Experiment design is a subtle skill.
As other comments have pointed out, once you start testing multiple hypothesis on the same dataset, you cannot apply the same significance threshold that you would if you had just begun with a single hypothesis before observing the data. Instead, you need to apply some sort of correction that takes into account the number of hypothesis being tested:
https://en.wikipedia.org/wiki/Family-wise_error_rate#Control...