At best, you have AI's that can easily recognize pathologies that an average rad could recognize, but are useless when it comes to trying to recognize a pathology that 99.999% of rads would miss. Why? Probably because the system is biased towards pathologies that rads recognize. Why? Probably because that's what's in the learning set. I understand all that, but that makes the AI almost useless in a production setting simply because of the way many healthcare delivery networks are structured with respect to radiology.
At worst, you have AI's that seem to recognize pathologies that an average rad could recognize, but then inexplicably have a horrible miss on an obvious study that any first year radiologist would have caught while half asleep. It's the sort of miss that makes the astute observer wonder if the other vendor's AI, the one that didn't have any misses on your test set, was simply not fed an example study that triggered its blindspot? You start to wonder where the blindspot on that one was? You start to wonder does it even have a blindspot? How can you work around blindspots like this in general? Etc.
But here's the thing, it's never good when you're thinking about mitigations while you're still testing the AI. My first RSNA was over 20 years ago. To this day, we're still hearing the same promises, and the production testing, (it never fails), is still uncovering the same issues.
Now I would have thought that recognizing a human would be easier than trying to ascertain, say, calcification in a DX, but apparently these same issues crop up in a number of different applications of ML based technologies.
I understand that a training set, by the very definition of the word "set", does not contain everything. Obviously, there will be bias towards whatever is in the training set. Essentially, most techniques today, simply train AI's to tell us whatever we tell the AI's to tell us. But for this reason it should be unsurprising that these sorts of AI's work best in settings with well defined domain spaces and are challenged when the domain space is less well defined. (And in either case, these AI's will have an inescapable bias towards whatever was in the training set.)
What does a tumor look like? Is probably too broad a question to give these AI's. What does a stop sign look like? Is probably a question that these AI's could answer relatively reliably. (I hope?) You would have thought that "what does a human look like?" would be closer to "what does a stop sign look like?" But I guess it's not surprising to hear that it seems to trend towards "what does a tumor look like?" in practice.