The issue is: you can't do stats. The reality is it is highly likely that all data here is consistent with any effect size you care to choose.
What this study does, nevertheless, is quantify the 'model risk' ie., the "p-value of the method itself" -- how often does research provide reliable effect sizes? Rarely. So the research methods themselves have a huge 'epistemic discounting' associated with them.
But we can go much further than this -- recurse: what is the risk of using this methodology to assess methodology? Actually, it's fairly high. This isn't a very reliable method for judging research methods. And so on.
What we find here is that in the vast majority of cases, where we'd care to research anything, recursing risks lead to an essential insurmountable error: measurement error, model error, methodology error, research-assumptions-error, etc. There isnt enough data to resolve between competing possibilities at the outer-most stage of error (often you'd need much more data than practically possible to collect).
What do we do with this? I'd say, mostly, throw away stats here. Do case analyses, risk-based recommendations, even broadly philosophical analysis.
The alterantive is just staring at tea-leaves.