The authors introduce a measure called the "fragility index" (rather, a modification of an existing fragility index) that's supposed to diagnose a lack of robustness in clinical trials satisfying the usual standards of p < 0.05.
First, the problem with clinical trial results failing to replicate or discover true effects goes far beyond the statistical methods used. For example, "virtually all major RCTs funded by an NIH institute (NHLBI) before 2000 were false positives. Once hypothesis preregistration is required in 2000, everything becomes a null." [0] Arguing about statistical methods is just rearranging deck chairs on the Titanic until all that institutional stuff is taken care of.
Continuing anyway: The introduced measure is extremely ad hoc and not supported by any theoretical framework (e.g. proofs of good asymptotic properties). I don't know how to interpret it, especially because (as the authors note) it can easily flag a study as "fragile" when it isn't. Say what you want about p-values, but they at least have mathematical guarantees when used properly.
Further, there's no discussion of how this compares to existing approaches. Why not just do careful Bayesian modeling? Or, why not try to learn something from the amazing success of machine learning researchers in predicting out of sample generalization (e.g. see [1])?
Maybe I'm missing something, but I really don't see what the contribution is here.
[0] https://twitter.com/paulnovosad/status/1427332860902584329 [1] http://jakewestfall.org/publications/Yarkoni_Westfall_choosi...