In grad school I was told to generally assume that the tool with the second highest line in the key graph was the best tool. The top line was invariably the author's tool and its results rarely translated to other experiments. But the tool with the second line was usable by somebody other than its original author on a different problem space.
Empirical SE research is part of the social sciences. The subset of SE research that doesn't push hard for collaboration with social science departments is doing just fine.
I listed Software Engineering because it is one of the fields I know well but is way less esoteric than the other one I know well (PL). You'll find the same kinds of replication problems in POPL or OOPSLA papers too.
I suspect that other subfields don't fare any better. I certainly hear from my ML friends that a huge portion of ML results are untrustworthy.
Those papers have lots of issues, especially cherry-picking and narration of the results' significance.
But the core claim -- "I found N bugs in the following codebases" -- is generally reproducible.
It will not happen, though.