ML is capable of doing that if you train it to do that.
I don't get this perfection requirement that gets put on so many systems here. No system is perfect, the question is what failure rate you accept.
So yeah, maybe you get a model with 5% false negatives where humans have 10% false negatives. That's good. However, when you go to apply that model in reality, what happens to those 5% where model and doctor disagrees? First, we should think about the kind of disagreement. In most of those cases I'd bet that the doctors see something wrong but can't say what, not that the model says "this is bad" and the doctor "it's not".
Second, we need to think about what happens when doctor and model disagree. As the model is not 100% accurate, the doctor doesn't know if the model is mistaken or if they're the ones mistaken, so they'll probably order more tests anyways. If it can be something serious, it's worth it to do an extra test to make sure. They'd probably ordered those tests anyways if they weren't sure of what was happening, model or not.
So what did the model for a single test change? Did it really change the diagnostic outcomes? What's the actual benefit of the model? How much it's worth to get from 10% to 5% false negatives with a model for a single test if just adding more tests (say, with the same 10% false negatives) to the mix can give you a 1%, 0.1% false negative rate?
That's my point. Unless accuracy is really high, ML models are not going to remove uncertainty in diagnostics, few diagnostics consist of just a single test. Benchmarking models against human performance in a single test is not a metric that can drive implementation in the real world.
But in general, these are matters of life and death, literally. You're still going to have someone looking at the images and verifying the output. That limits a lot the potential cost benefits of these applications.
My point isn't that you need 100% accuracy. It's that a diagnosis is a process, not a single test. If your model applies to a single step of the process, and if it doesn't remove enough uncertainty about that step, you're still going to continue the diagnostic process and the model is not going to change anything really.
E.g., “it will have no effect on outcomes until pigs fly.”
There will always be uncertainty, and uncertainty isn’t the only relevant parameter.
But it's a really important one and a lot of medical ML research doesn't seem to address in the proper sense. A very simple example: a ML model that classifies lung nodules as benign or malign, with 95% accuracy vs 70% accuracy of regular radiologists. Very good, right? But for actual, real world results, you need to see how the patient outcomes change. If the patients where the model and doctors disagree were going to have extra tests or followup regardless of what the model says, the model is not actually offering anything new despite the increase in accuracy.
So no, it's not absurd to say that models that work for a single type of test need to be very, very accurate to actually bring changes to the procedures that are worth the investment. What matters is whether they actually change patient outcomes, and sometimes it seems like the ML researchers barely consider it.