I never understood these arguments. Isn't it like claiming self-driving cars via neural networks would never be possible because you can't test that the neural network would take the correct decision in every situation?
I view the whole issue stochastically, with the immediate aim being to make (significantly) fewer errors than the current approach, which is having a human decide. I don't claim that I can design an experiment which could selve as an indication whether we are improving upon human judgments, but I think this should be the goal.
Reflecting upon my view, i think it comes from the experience of training ml-algorithms. You are always minimizing errors, but you goal is almost never to make 0 errors, because often your data is noisy and you are probably overfitting. I know the medical enviroments are more sensitive, but I can't really wrap my head around how we could design a learning algorithms that does not make any error and works on all abnormalies. I think it will always missclassify.
Rephrasing my argument: I think the approval should be given if an significant expected improvement over the distribution of real-life abnormalies can be detected and not over the uniform-dsitribution over all abnormalies.
EDIT: detecting out-of-distribution samples is hard and I don't think this is a solution and leads to a false sense of security.