> we do have a lot of statistics now on the outcome of patient not treated with that drug. There's no need for establishing a benchmark. Just proper categorization of existing patient should be enough to compare.
This is very wrong; we have good aggregate statistics across lots of different hospitals, but for a study with a small number of sites where the patients are treated with the drug, there's enough inter-site variability that may obscure any effect, or create a false effect.
This is doubly true as health care varies as sites get overwhelmed, and healthcare workers get increasingly strained from long shifts. There will be far more variability as time goes on.
And for how small we expect the effect to be, based on the data in the paper, we really should include controls.
All that said, I think that the publication of this flawed study is very good, and why pop-science takes that excoriate science as "most research finding are false" really communicate the wrong way to think about science, and limit its applicability. Flawed datasets like this still give us clues, and publishing flawed data is still useful to others to accelerate the pace of science. Which is why we need to take the attitude that publications are point in time guesses at what's going on, and only rarely does one come out that can be considered on its own as definitive, and that it's good to have mostly "here's some data and analysis but we don't have a complete theory yet"-type-papers. Without those intermediate publications, science would grind to a halt.