"Please use the original title, unless it is misleading or linkbait; don't editorialize." - https://news.ycombinator.com/newsguidelines.html
The 6% only scored better when compared to a single radiologist review. That's why we use double reviews.
We being Europeans. Double reads are not standard in the majority of the world, including the US.
I work an an AI company where we screen for Diabetic Retinopathy. The company is a decade old, and has validated in huge studies (over 100k patients). It's a hard problem that we have all worked very hard to solve it. At the same time, it's easy to build an AI tool that looks at an image and says healthy or unheathly (or poor quality image). So the bar is low for making something that is appears functional.
But hardly anyone makes good AI tools, so studies that look at different AIs systems see lots of the bad ones. It's always a bummer when the headlines are all dismissive of AI in general. Googling "diabetic retinopathy AI" the top result is "Artificial Intelligence Falls Short in Detecting Diabetic Eye Disease" (https://healthitanalytics.com/news/artificial-intelligence-f...), yet if you read the article it says one tool is better than humans, which to me, is the real takeaway.
This is like the comparisons of 200 top hedge funds to index performance and concluding that hedge funds can’t beat the market. Ignoring the other conclusion: most hedge funds suck
Edit: or most hedge funds are fraudulent?
I think you've misread what conclusion you're supposed to draw from this kind of assessment.
For both the hedge funds and this the conclusion isn't that the job is impossible - but that we lack the tools to predict who will do the job. You can't know which hedge fund will outperform (though some will) and we don't seem to be able to predict which AI models will perform well enough to use. Also that your odds of picking a winner "from the field" are quite bad.
It doesn't mean there's anything wrong with using a hedge fund - but there's a level of risk that one shouldn't ignore.
Like, my conclusion from this is that I would not accept a pure AI solution (because most of them are bad), but I would be interested in AI assistance. For many people, the promise is in replacing not supplementing doctors and this is the same as failure.
Oh no, that could be the conclusion. Ever study quantum physics? Lots of people have tried to come up with some kind of model with inner "hidden" variables that tries to make HEP physics into a completely deterministic, almost classical, theory. By your logic, if we just trained and AI with enough subatomic interactions, it would eventually be able to predict with 100% accuracy the results of a quantum process.
That's utter bunk.
Markets are the same way: people think there must be some rules that can be worked out to make them 100% predictable, when in reality there is an element randomness that makes cryptographers jealous. It will never be possible to predict the next number in a random process. If you flip a coin 99 times and get 99 heads, you still can't predict what the 100th toss will be (and before someone "well ackshually"s with something about an unfair coin: don't)
I mean, the job certainly could be impossible, but those studies don't prove it. As it relates to the discussion at hand we know that we can do better at evaluating mammography because humans do it.
Maybe I am not understanding what you are saying?
So 2 of the 36 were better than one radiologist, none were better than two radiologists collaborating.
I am sure the average radiologist, tired and bored at 8am in the morning, is much worse.
Imagine like "94% of all attempts at heavier-than-air flight crash on first attempt". Sure... kinda misses the point.
I asked about AI and they couldn't use the one they had on the cyst because it was not trained for ankles.
> Thirty four (94%) of 36 AI systems evaluated in these studies were less accurate than a single radiologist, and all were less accurate than consensus of two or more radiologists.
(Or just 94%, and then have people say 'but N= only 36!'.)
Percentages for small numbers, in particular below 100, are just annoying and often designed to mislead. (Though I'm not suggesting that here, it's standard in medicine if not other fields.)