In my mind, 50% replication is not that horrible. Even decent seeming sample sizes might have an 80% chance of replicating at a certain effect size, and the effect size originally reported could be (by chance or other reasons) over-stated, leading to insufficient power. More certainty ends up meaning vast increases in sample size). Another issue that should be tested (have not read paper but doubt it) is not just whether the result is statistically significant, but whether it statistically differs from the original published result. If the effect is weaker, but not statistically different from the published result, I would hesitate to say the effect is not real. Then there is that some of the psychology findings investigated were rather dubious to begin with, and the ones that even non-specialists can identify as being likely to be true replicated.