I don't think machine learning suffers from the same kind of p-value-driven replication crisis as other fields. It is true that people don't generally perform proper statistics to compare machine learning models [1], but ML research has two things going for it that other scientific fields do not. First, comparing machine learning models on the same test set corresponds to a within-subjects analysis, generally with tens of thousands of subjects, so the noise level is low. Second, because ML researchers don't generally perform hypothesis tests, they care solely about effect size and not about significance. If my model gets all the same examples right as the previous state-of-the-art, plus 10 more, then my model is statistically significantly better, but on a test set of 10,000 examples this corresponds to a 0.1% accuracy improvement, which is generally not big enough to publish. By not doing hypothesis tests, ML researchers actually tend to be more conservative than their p-value-driven counterparts in other fields.
In the Recht et al. study, the reason the new test accuracy is wildly outside of a binomial confidence interval around the original test set accuracy is that the distribution is different. The CI only applies to data drawn from the same distribution.
ML research still suffers from replication issues; such is the nature of the scientific incentive structure. However, these issues generally come in the form of poorly tuned baselines, buggy code, and claims with insufficient experimental/theoretical justification. Outside of some isolated cases, publication bias and cheating at hyperparameter tuning do not seem to be major factors.
----
[1] Statistically speaking, to compare two models on the same dataset, one does not care about the accuracy numbers but instead about the number of examples model A gets right that model B does not and vice versa; see McNemar's test.