First, the reported p-value might be wrong. E.g. basing it on assumptions of normality when the data is non-normal. However modern non-parametric approaches like the bootstrap can avoid this issue.
Second, testing multiple hypotheses. If you test 10 hypotheses then you cannot reject the null (that all 10 null hypotheses hold) simply because one single hypothesis is rejected in isolation. But this is well known, and failing to account for it is an issue with the researcher, not with frequentist statistics. I actually think that the main practical difference between Bayesian and Frequentist statistics is whether accounting for the issue of multiple hypotheses is done formally or informally.