Clearly there is room for improvement. Maybe this study will also spur the development of new types of systems which augment human/radiologist decision-making.
Clearly there is room for improvement. Maybe this study will also spur the development of new types of systems which augment human/radiologist decision-making.
indeed that is the problem in many areas. Almost always mgmt looks with "tech = fewer human = instant cost savings" expectations.
and anything that increases cost in the short term are shot down. sad . all blame on the cost-center vs profit-center mindset.
When I was CTO of DocHuddle (ML + Radiology), we reviewed every paper and preprint we could find and a huge number appeared to be overfit on tiny datasets.
Any meta-study which doesnt throw away obviously poor attempts will find bad-skewed meta-findings.
Do Hospitals keep radiological images? Are they researching this stuff? Presumably they'd have a large enough sample size.
I'd disagree this is due to privacy. IMHO it is due to conflicts of interest in major medical systems (certainly in the US) where incentives are skewed away from efficiency towards more billing.
Getting images is one thing. Getting labels is harder. Getting annotations on regions-of-interest is yet harder. Often the labels are stuck in unstructured data (notes.)
When I was doing my startup (https://www.dochuddle.com/) we trained our classifier and object detector on 1.2 million images. We worked with a large sovereign on getting the images, reports, and annotations.
Just dealing with 20+ TB of images was a job in its-self...but it was indeed unreasonably effective. The success was only technical and ML success.
Even harder -- commercializing it effectively after accounting for legal/contract costs. Yet harder -- getting past conflicts of interest inherent in the US medical system. We were not successful on this front. Possibly too early (we started in 2014)
We screen people with a cervix for cancer by scraping away a tiny sample of the cells periodically and having that examined at a laboratory for anomalies. If caught early, cervical cancer isn't fun but it's extremely survivable. You don't technically need a cervix (after all about 50% of humans don't have one) so in the worst case hysterectomy (removal of the womb and cervix) is an option.
Machines aren't very good at looking at a slide full of cells from the human cervix and giving it a score like 1-5 where 1 is "Fine" and 5 is "Cancer". This is a standardised task that her (human) team do every day, and she works with other bodies across the continent to ensure they're all doing a roughly similar job by looking at each others examples and checking they get the same number, so this way a doctor in Paris and one in Birmingham should get interchangeable results despite using different labs to process the samples.
However, it turns out the machines are stellar at a related task. "Is this sample infected with HPV?". Humans do not find this task easy and historically this test wasn't done anyway.
So maybe a human would do the first thing, and, if it was a bit borderline, then they'd check the second thing. Cervical cancer is almost always caused by HPV, so if you don't have HPV then you almost certainly don't have cervical cancer. But since the machines are great at that second problem you can reverse things. The machine processes every sample for HPV and then a human only looks at the positive ones, rating those.
Automation though can also make experts stupid depriving them of the practice to keep skills sharp.
The gains could be incredible though, not only would we get better diagnosis, it could be done faster, cheaper, and remotely.
"On average"?
What if the machine is better than average for common things but consistently, 100%, misses uncommon conditions with very high short-term mortality rates?
The evaluation of AI medical devices is determined by regulatory agencies like the FDA.