So yeah, maybe you get a model with 5% false negatives where humans have 10% false negatives. That's good. However, when you go to apply that model in reality, what happens to those 5% where model and doctor disagrees? First, we should think about the kind of disagreement. In most of those cases I'd bet that the doctors see something wrong but can't say what, not that the model says "this is bad" and the doctor "it's not".
Second, we need to think about what happens when doctor and model disagree. As the model is not 100% accurate, the doctor doesn't know if the model is mistaken or if they're the ones mistaken, so they'll probably order more tests anyways. If it can be something serious, it's worth it to do an extra test to make sure. They'd probably ordered those tests anyways if they weren't sure of what was happening, model or not.
So what did the model for a single test change? Did it really change the diagnostic outcomes? What's the actual benefit of the model? How much it's worth to get from 10% to 5% false negatives with a model for a single test if just adding more tests (say, with the same 10% false negatives) to the mix can give you a 1%, 0.1% false negative rate?
That's my point. Unless accuracy is really high, ML models are not going to remove uncertainty in diagnostics, few diagnostics consist of just a single test. Benchmarking models against human performance in a single test is not a metric that can drive implementation in the real world.