Use of artificial intelligence for image analysis in breast cancer screening
bmj.com
bmj.com
This is not a "weird relationship" thing, rather women got organized and: https://en.wikipedia.org/wiki/Breast_cancer_awareness
It’s more broad than that, it’s the relationship with women’s health as a whole.
If you're really interested, watch some older Dutch movies to see the normalness of naked, like Turkish Delight [0]. Or even have a closer look at the relation between the Professor and Raquel in the recent Casa the Papel. It's different from US series. More respectful, mature wrt women if you ask me. More emphasis on intelligence. Or on a beautiful woman in her 40s with normal wrinkles. The US, like on many other topics, seems to be polarized, caught between the hyper sexuality of Cardi B and the prude nature of American culture in general.
[0]: https://en.wikipedia.org/wiki/Turkish_Delight_(1973_film)
Fixed that for you.
You may want to consider finding a different medical professional!
So our best shot from a public health perspective is to say "Here's what we recommend for everybody" and pay for that.
All screening programmes have two difficulties, which must be balanced against the benefit, and this trade is somewhat personal, so when the balance is quite fine the arguments can be vociferous as a result.
1. The screening itself may seem unpleasant. One woman may find it a very mild annoyance, a drive ten minutes out of her way, the staff are very pleasant, the scan itself is far less traumatic than a bra fitting, and she receives easy to understand results after not very long and isn't anxious about them; but for another maybe it's an hour's bus journey to the city hospital, the staff there are short-tempered and say she has the wrong paperwork, then another hour in a queue, she feels like she's just meat, squashed around for the convenience of the machine for what seems like forever, and then after anxiously waiting for what seems to be too long the results are confusing to her and she has to have a friend interpret them.
2. Over-treatment is always a problem. Screening by definition detects something that isn't causing noticeable symptoms. If you have a noticeable lump, or mysterious bleeding, you don't need screening you need a doctor's appointment. So a positive screening result might be nothing important. However either you've now got the burden of a diagnosis you ignored or, you accept the medical advice and are treated, even though it's possible (not likely, but possible) that you would have been just fine without treatment.
So, screening programmes are set up based on guessing how to trade these factors plus a third, how much should we spend on this medical intervention? After all, in some sense every dollar doing breast cancer screening is a dollar you don't have to cure blindness in poor orphans (or of course, to bomb somewhere)
If your experience of a screening programme is that it's a minor inconvenience at most, and yet you know people who died of undetected disease, more screening seems like a no brainer. Particularly if you live somewhere where screening stops at age 50, and somebody you know died of undetected disease aged 54, you might reason that the screening should go to age 55 or 60 to detect such cases, no matter the public cost.
On the other hand if your experience is that it's an awful ordeal even when negative, and you know people who spent their last years horribly scarred by surgery as a result of suspected disease but then they died in their sleep from something else anyway, you may feel that there's already too much screening and it should be trimmed back, not to save money but the extra money for other programmes is welcome.
I certainly don't worry about the cancer I could develop as a result of an xray in a regular dental cleaning (even though that's a possibility).
However if you are a doctor trying to figure out how often to screen women, the danger if xrays is a strong reason to not do it too often. Daily screenings would catch breast cancer a lot sooner on average, but just isn't worth the risk even if the screening was otherwise free.
The same issues bedevil prostate cancer screening.
While Reagan did a lot to try to defund cancer research regardless, the first lady's mastectomy drew a lot of attention to a previously taboo topic.
[0] https://www.nytimes.com/1985/07/16/us/reagan-s-illness-medic... [1] https://www.latimes.com/archives/la-xpm-1987-10-18-mn-15261-...
Back when I was a life scientist the best advice I got was to work on diseases the family members of congresspeople had.
But since screening has been framed as caring for women, pointing out the flaws with screening is automatically seen as hating women.
https://en.wikipedia.org/wiki/Mammography_Quality_Standards_...
Our hospital bought an AI stroke detector (viz.ai). The "AI" part is laughable but it allows the hospital to collect some extra HHS/medicare fee for using "imaging algorithms" or something. I suspect that company is potentially using anonymized scans to continue training their data, because the initial product was trained on some tiny sample of brain CT scans, like under 500.
The one plus for actual physicians and nurses is that at least they wrote a non-sucky PACS imaging interface for the iPhone so we can just pull up the plain scans and view them easier.
Why would that even pass the FDA smell test?
Right now much of the mammogram reading is extremely straight forward. Radiologists can fly through those exams with high quality, especially when you consider the macros they have setup that allow them to write most of the report with with a two second voice command.
This is one of the biggest things radiologists complain about- from their perspective most AI is solving problems that aren't actually useful to radiologists. I think the companies that will do best are the ones that augment radiologists with a focus on making them more accurate and efficient, rather than trying to cut them out of the loop (I'm a bit bias though as my previous company, Rad AI, is taking the augmentation approach and it has worked out rather well).
Often, the ML provides an extra layer that slows things down, even if it's working at state of the art AUROC or whatever we're trying for.
But those are also harder for software based systems.
I think you're disagreeing with a point that the parent comment didn't make.
It's worthwhile in general to continue AI research. The properties of this application make it ideal for testing and progressing AI technology. But if AI doesn't end up being used in this particular area in the short or medium term, that's another matter entirely.
There's a difference between: not ideal to tackle and not ideal to build a business around.
It is absolutely ideal to tackle which will hopefully build confidence in a system that can then be expanded to solve harder problems. If you can't even solve this problem I don't see how you're getting public confidence in the harder ones.
The third world might need more effective training, so they can convert more cheap labor into specialists, but I don't see why per se they would prefer to use a computer sold at Western prices to a local specialist.
There's also the fact that radiologists get paid a pretty decent salary. If I could devise a system that can do a highly paid worker's job faster, cheaper, and at a similar level of effectiveness, why wouldn't I? It also opens up that skilled worker's schedule and allows them to tackle more difficult to interpret scans.
You are suggesting something that's against the groupthink for that profession. You can't graduate from an MD program while holding these thoughts.
I'm being 100% serious here.
By your logic I must not exist then.
Like who do you think is labeling your training images?
Look at the authors of pretty much all cited studies in that meta-analysis. Sure see a lot of MD and MD-PhD there.
The point I'm making is that focusing on mammograms is not the way to do that. It's already a very optimized specialty. If AI researchers worked more closely with practicing radiologists (not academic ones) I imagine there's all sorts of areas where existing technology today can improve things.
It's the reason professional baseball is experimenting with electronic strike detection, despite the fact that umpires are very accurate (~94%) calling pitches.
I get that there are useful applications of AI that could augment a radiologist. I'm simply pointing out that if one could completely remove a radiologist from the equation, that could be a super helpful tool. I suppose my view is that removing/limiting the human element in routine inspections is a very worthwhile goal.
There's also the side benefit of increasing trust for these technologies. I think a lot of people are rightfully scared of AI technologies; showing that they can be as effective as a radiologist for routine screenings would be very beneficial.
That's almost certainly true, and surely there must be people in the medical industry already working on this. In hospitals, there are internal organisations that vaguely attempt to link up technology workers with practicing doctors to improve medicine using technology. One of my most competent former developer colleagues retrained as a doctor and worked in medical digitisation, as well as being a practicing doctor.
However:
* Often coming at a problem from another angle and ignoring existing professionals can be useful too. Existing workers are very often stuck in their ways and unwilling to accept the way they look at a problem may not be the best. I have heard from many in medicine, that doctors are extremely distrustful and hesitant towards technology, as well as fiercely protective of their jobs.
* The medical industry is internally highly dysfunctional and trying to work with practicing doctors can be extremely politically difficult. The hospital bureaucracy often actively tries to prevent practicing doctors from working directly with technologists, preferring instead to silo them and direct all communication via senior management - who are often neither technology nor medical experts. Consequently it seems more likely that major changes will come from outside.
* AI has occasionally (rarely?) shown great success in solving problems without needing domain experts. See AlphaGo.
Games are always orders of magnitudes easier then analyzing reality. They are highly constrained to the narrow domain and space of possibilities. Reality doesn't. There is work on understanding physical laws by the AI but it's very long path and they are much less successful than their artificial reality colleagues.
In game domains we have the ability to generate infinite meaningful training data and have easily measured outcomes. Waymo is generating near infinite data for driving, and doing quite well with it. In medicine, we have piddly ass datasets because of privacy concerns. 'Reality is hard' had nothing to do with it.
This isn't going to be a commonly done thing.
This is one of the areas where working with radiologists actually makes things easier. Radiology groups are normally independent practices that contract with multiple hospitals- they link in with hospital systems but mostly run their own thing. They also tend to have really solid IT departments and enough funding to make things work. Working with hospitals is a nightmare though, as you say.
to have them read an AI's diagnosis before having to assess a mammogram themselves to confirm the AI's finding is adding an unnecessary step to their workflow.
That to me sounds like an ideal automation use case. You have an entire class of highly skilled professionals performing a repetitive task. Sure, they may be able to do it quickly, but I imagine they must have to do a bunch of them each day.
You could imagine a world where hospitals didn't have radiologists on staff, but rather they acted in a consulting capacity for difficult cases.
Walking on variable surfaces is also a highly repetitive task and automation is yet to master it.
That's empirically not true; the first such product approved for clinical use entered the market in the late 90s, and they've had small-to-medium success since. Definitely none of them changed things radically, but it's pretty common CAD functionality now and has both detractors and champions.
> Radiologists can fly through those exams with high quality,
This is also not entirely true, one of the biggest problems for screening (not diagnostic) programs is that the false negative rate for radiologists is high. They do fly through these scans, but that's to make it economically feasible. Also, the average radiologist performance on screening mammo is not very good - most have to do it regularly (probably 100s per month) to do it well.
You are right that augmentation is the more plausible path to clinical success, but that's been what most of the clinical systems aim to do anyways, provide a sanity check on the radiologist.
E.g. If a system could tell me which 967 results out of 1000 are normal, and send the other 33 for human review.
True positives require solving every problem in the solution space. True negatives only require solving the most frequent problems.
Plus--wouldn't the software company be liable? (I don't think the tech is ready for office use, but when it is--I still think it will be a battle.)
Kinda like all medical devices. Whenever there's a problem, the equipment is looked at first?
Personally, I don't think the AMA wants computers taking away their doctor's precious income, and will only acquiesce when the tech makes the doctor more money, or is so good politicians/insurance companies start demanding it.
But IMHO medical law in general has a big device / software blind spot. It's become decent at figuring out liability for humans, while permitting normal operations.
But all that machinery and case law doesn't really exist for black box software systems (no human in the loop). And if there's one thing that medical providers and insurance companies hate, it's unbounded liability.
But yes, augmenting human intelligence instead of substituting it often does work better.
This looks good from 70,000 ft, but in practice you need labels much more specific than BIRADS.
Clearly there is room for improvement. Maybe this study will also spur the development of new types of systems which augment human/radiologist decision-making.
indeed that is the problem in many areas. Almost always mgmt looks with "tech = fewer human = instant cost savings" expectations.
and anything that increases cost in the short term are shot down. sad . all blame on the cost-center vs profit-center mindset.
Automation though can also make experts stupid depriving them of the practice to keep skills sharp.
The gains could be incredible though, not only would we get better diagnosis, it could be done faster, cheaper, and remotely.
"On average"?
What if the machine is better than average for common things but consistently, 100%, misses uncommon conditions with very high short-term mortality rates?
When I was CTO of DocHuddle (ML + Radiology), we reviewed every paper and preprint we could find and a huge number appeared to be overfit on tiny datasets.
Any meta-study which doesnt throw away obviously poor attempts will find bad-skewed meta-findings.
Do Hospitals keep radiological images? Are they researching this stuff? Presumably they'd have a large enough sample size.
I'd disagree this is due to privacy. IMHO it is due to conflicts of interest in major medical systems (certainly in the US) where incentives are skewed away from efficiency towards more billing.
Getting images is one thing. Getting labels is harder. Getting annotations on regions-of-interest is yet harder. Often the labels are stuck in unstructured data (notes.)
When I was doing my startup (https://www.dochuddle.com/) we trained our classifier and object detector on 1.2 million images. We worked with a large sovereign on getting the images, reports, and annotations.
Just dealing with 20+ TB of images was a job in its-self...but it was indeed unreasonably effective. The success was only technical and ML success.
Even harder -- commercializing it effectively after accounting for legal/contract costs. Yet harder -- getting past conflicts of interest inherent in the US medical system. We were not successful on this front. Possibly too early (we started in 2014)
The evaluation of AI medical devices is determined by regulatory agencies like the FDA.
We screen people with a cervix for cancer by scraping away a tiny sample of the cells periodically and having that examined at a laboratory for anomalies. If caught early, cervical cancer isn't fun but it's extremely survivable. You don't technically need a cervix (after all about 50% of humans don't have one) so in the worst case hysterectomy (removal of the womb and cervix) is an option.
Machines aren't very good at looking at a slide full of cells from the human cervix and giving it a score like 1-5 where 1 is "Fine" and 5 is "Cancer". This is a standardised task that her (human) team do every day, and she works with other bodies across the continent to ensure they're all doing a roughly similar job by looking at each others examples and checking they get the same number, so this way a doctor in Paris and one in Birmingham should get interchangeable results despite using different labs to process the samples.
However, it turns out the machines are stellar at a related task. "Is this sample infected with HPV?". Humans do not find this task easy and historically this test wasn't done anyway.
So maybe a human would do the first thing, and, if it was a bit borderline, then they'd check the second thing. Cervical cancer is almost always caused by HPV, so if you don't have HPV then you almost certainly don't have cervical cancer. But since the machines are great at that second problem you can reverse things. The machine processes every sample for HPV and then a human only looks at the positive ones, rating those.
I would suggest these two things might be linked - the only way you get publishable results is by failing to do rigorous studies.
I hope I don't offend any other people in the field, but I think historically fields like histology/radiology/MRI screening via AI have been subject to less scrutiny by virtue of their multi-disciplinary nature... particularly in the past, it was hard to find reviewers that both (A) understood the biological/clinical validation (B) the technical validation.
I think things have drastically changed over even the last 5 years, which is why you're more likely to see headlines like these. We have a lot of discussions like these during our lab's journal clubs. The optimists among us (e.g. myself) argue that this opens up a lot of opportunities whereby more fair validations/benchmarks allows us to compete with "SotA" methods that are actually fragile and easy to surpass on even ground. More pessimistic members are quick to note that it's far harder to introduce new, fairer validations and that reporting lower metrics is far less buzz worthy (all true points).
This reminds me of a fundamental error that was recently found in methods that tried to predict protein-protein interactions. It was found that the train/test/validation method used by ostensibly every paper was leaking huge amounts of data [0]. When we plugged that leak, we saw much more modest metrics, but focusing on regularisation allowed us to beat the competition [1].
The kicker? The information leak were identified in 2012, yet you'll see papers written every year that have the same leak and report >90% accuracies, and >0.9 ROCs.
[0] https://www.nature.com/articles/nmeth.2259
[1] https://www.biorxiv.org/content/10.1101/2021.08.13.456309v1
"Please use the original title, unless it is misleading or linkbait; don't editorialize." - https://news.ycombinator.com/newsguidelines.html
Like any domain in applied AI, there will be a lot of approaches that miss the mark, or are simply stepping stones to better approaches. There are thousands and thousands of papers on language modeling, but we only needed one superior approach (GPT) to change the game entirely.
The search through any cutting edge problem space is messy and full of failure, and that's fine. You only need one breakthrough.
I guess a good questions could be: Will they be more reliable than radiologists in scenarios other than the ones studied?
It's fair to say that yes, AI systems aren't good enough yet. On the other hand, it's pretty clear some technological approach will outperform a radiologist at pattern recognition at some point in the future - whether that's "AI" or "if statements" or some third option.
It's just a matter of time.
How much time does it take to train a radiologist to that level of performance?
How much time does it take to clone that ML model?
After all, no AI beat two radiologists.
AI based analysis - whether it’s better than a human radiologist or not - is far more scalable and cost effective. Even if used as a screening mechanism to be escalated to a human radiologist, this approach will be very helpful to much of the world.
Here, the context appears to be a somewhat arbitrary selection of published algorithms. All they've really determined is that at least 94% are not ready to replace radiologists.
That's pretty much confirmation of the default assumption. If they were, they'd all be trying to get these into hospitals, and they're not.
Without any more details about the error rates, we can't be sure how likely this is due to chance. I would caution making any conclusion about AIs without better understanding the underlying statistics.
FTA:
> Thirty four (94%) of 36 AI systems evaluated in these studies were less accurate than a single radiologist, and all were less accurate than consensus of two or more radiologists.
So yeah, no AI system beat consensus of two radiologists. That's pretty damning.
Indeed. And we all know how quickly radiologists are improving at their job. At this rate the 6% of AI systems that beat one radiologist will be down to 0% in no time.
It's the same problem I have with self-driving cars. We are teaching cars to behave around (and like) humans, instead of just fencing off the highways and inventing an autonomous system of vehicles that doesn't have this much tougher constraint. I think AI can do a lot of good things today, if the people trying to apply it looked to completely redesign existing systems around AI, instead of trying to replace the humans operating the existing systems. The latter is much more difficult.
As a popular joke likes to say: The best place for AI cars is on special roads. You can then make those cars drive very close togehter, even touching. Oh and you can then move the special road behind buildings instead of the front. Give room for more pedestrian friendly pleasant streets. Then you may as well move the road underground.
You have invented the subway.
Another is that you could have your car use its AI in the special roads in the city and highways, and the rest of the time drive yourself, where AI is harder to implement. A bit like some hybrid cars where you can do most of the day to day commute on electricity, but for longer road trips you can use gas.
It has the advantage that it covers the very common use case where someone says they can't take public transportation because they need mobility at their destination.
And that still leaves aside tricky things like construction sites, double parked cars, etc.
Commercial airline autopilot has a lot more straightforward job: at normal operating altitudes there are no animals to worry about, no construction or other unexpected obstacles, and few other aircraft all of which can be essentially guaranteed to be squawking transponder codes of their own, and they still are mandated to have two pilots in them at all times. This is not because airlines like it- look how fast the flight engineer disappeared once flight management and navigation software became sufficient. It is instead because airlines have learned over the decades, from a lot of blood, that sometimes you need to have a pilot ready to intervene RIGHT THIS INSTANT, not 3 minutes from now when you've regained situational awareness after zoning out for a bit- and the AF447 disaster shows that they are right about that. And the only way to make sure you have such focus and alertness over many hours is to have two pilots, so they can trade off the responsibility of maintaining SA.
And that's a much easier problem which has been actively attacked for decades.
https://www.mic.com/articles/120898/7-simple-ways-to-stop-se...
With diagnoses, collating a consistent set of data is not a requirement you can impose because of how fractured the healthcare system is. If you had great AI systems in place maybe they could be convinced to invest in these systems, but you don't and so creating them is quite difficult.
Similarly, in the case of cars having their own highways and roads: how do we do this in a way that people find acceptable? If you prevent current highways from being driven on by normal cars, you cut off the majority of people from their daily commute which is a nonstarter. Making new highways in most places will also be a nonstarter for both cost and space reasons. At best you might hope that decades from now every single car has the means for self-driving in the limited environment and then one day the government turns it on, but that also feels iffy to me.
Building parallel systems incrementally creates harder technical problems but allows you to side step even harder coordination problems.
Well, if you did that, you probably wouldn't even need AI for the control system. If you know the parameters and have some control over unknowns, you want a deterministic control system. "AI" will hopefully handle the cases where you are in an uncontrolled environment. Current "AI" doesn't seem quite up to the task of self-driving. It will probably get there, but when is anyone's guess.
Wouldn't need it, no, but could use it to analyze traffic patterns and optimize flow for fuel efficiency, travel time, and safety. Those are much more constrained problems for which we have some useful tools.
...that we would have no way to evaluate the systems' accuracy because no other non-statistical system can understand the input data?
The Oracle at Delphi is always correct. Even when the Oracle is wrong, the Oracle is correct.
"...just fencing off the highways and inventing an autonomous system of vehicles that doesn't have this much tougher constraint...
Kind of expensive to build a second interstate highway system, though.
No, we can always evaluate the accuracy based on our current human-understandable images and scans and biopsies.
Of course this is all assuming there is something there. It might be there is nothing to find. I wouldn't be surprised either way.
Who would submit to a test that nobody knows what it means? What ethics department would let you run that test?
Now, if you want to use metrics already on file that radiologists don't generally consider, that might be interesting.
Of course, fencing off existing roads to exclude human-driven cars is an economic impossibility, like converting all our electrical systems to use compressed water instead.
Regardless, today's autocars are clearly NOT ready to fly solo in the absence of humans, as the Tesla that crashed into an overturned truck on a clear day a few months ago so clearly demonstrated.
And then they can platoon together to reduce air drag 4x, be electric and be built on a special surface to reduce friction 4x compared to 'normal' road?
I think that's called a train
If you shift the curve you could detect earlier and reduce false positives.
>> In two of the largest retrospective cohort studies of AI to replace radiologists in Europe (n=76 813 women), all AI systems were less accurate than consensus of two radiologists, and 34 of 36 AI systems were less accurate than a single reader
I'm still a little unclear how you measure the accuracy of a consensus of radiologists, without relying on the consensus of other radiologists. Maybe just a bigger consensus, or by using known end results but looking at imagery that could have been an early warning?
Therefore Hospitals should be investing in those 2 systems?
It could be that 2 of those systems are actually better. It could also be that the state of the art in deep learning just isn't there yet for radiology. But this is a strong suggestion that research is being dramatically overstated, so I would get real careful with my investments in this area overall.
"Please use the original title, unless it is misleading or linkbait; don't editorialize."
Same last year, when I clearly broke a rib: first x-rays revealed nothing, one week later a second set revealed a really nice oblique displaced fracture, which was still hard to find.
Today the tibia also got x-rayed and it was a pleasure to look at, but as soon as it has to do with the chest, where all the organs are messing with the x-ray, one can really hope that AI will finally solve this issue, since every city has a couple of doctors with x-ray machines, contrary to MRI/CTs.
I'm reconsidering my life choices regarding MTB, since I'm already in my 40s and made zero exercise for at least 20 years.
But this time I learned. I will never risk it in the summer, since those sunny days are way to valuable for spending the time on the bike instead of recovering. I told me this last year, and this time it definitely sunk in. "Be careful, you're not a teenager".
But it makes so much fun. It is so good to breathe in that fresh air and to exercise hard, I always end up with a smile. Too bad I didn't start with this 20 years ago. I had the chance.
Just now getting set up to get more time outside . Sounds like i should take a lesson from your pain. :) Hope you get feeling better soon and get back out there (if maybe a tick lower kinetic energy :)
(also i meant to put a smiley in my first reply, sorry for coming across as a dick)
was that taken in different angle / perspective ?
Stage 1: "So one managed to do it, but 96% of implementations still can't"
Stage 2: "Well, looks like chess is pretty much dominated by computers. But the Go Game is still way too complex for a mere machine to tackle: only the human mind can"
Stage 3: "It was one tournament that AlphaGo won and really..."
Stage 4: "Yeah, the AI can play better Go and Chess than humans. These are solved problems."
I guess we're at stage 1 for self driving and apparently for mammograms as well. Only difference is, chess players didn't have a legalized cartel working for them in the 90's.
Something that is "not as good as a radiologist but maybe better than me," and with results in minutes instead of hours, would be HUGE, if even to just prompt me to review the images myself before (instead of after) seeing the next patient.
The reason I think this (as a non-specialist) is this: when I upload a random photo to Google or Facebook, those systems recognize the faces of people in those pictures without any prompt. This seems like a much more constrained problem (given a pretty specific set of images of chest cavities, state the probability of a cancerous tumor.)
I am guessing that there's nothing inherent in this problem that is much more difficult than other image recognition problems. I suspect that this outcome is because we have fairly new technology competing against highly specialized humans doing a very specific task and doing it well.
I suspect that will happen is that the technology will relatively quickly catch up to humans and do equally well, after which point it's just a speed of response and economics question, where the technology wins over the person.
> I am guessing that there's nothing inherent in this problem that is much more difficult than other image recognition problems.
This is factually wrong. You're assumption I guess is that its a data issue: if we can just get more data we'd solve this issue.
The reality is that these diseases are complex, their presentation is complex, and compounded by different technology. There's also the issue of data complexity, x-rays contain less information compared to a picture of your dog.
We need better algos, not more data.
Also, different modalities will produce different quality images, how was that accounted for in training models? Did they use all the images for a single scan or a subset of the images of a single scan?
The problem is you're trying to train a model where you have many images for a single scan, like slices. Depending on the modality, you'll get different resolutions, different "visual" inclusions, etc. etc. So labeling the individual images and labeling the entire collection is really hard.
I played around with a number of image segmentation models a couple of years back and I can only imagine how much more subtle densitometry stuff can be, but overall I'm not surprised accuracy can be way off because _any model_ can be way off depending on what it's presented with.
Most commercial Machine Learning doesn't actually learn on the job, and is actually positioned for the triage/faster handling scenario. It will spot what it was trained for, but _most likely only in the conditions it was trained for_. Change the angle, change the patient's age, add any other visible medical conditions, etc., and ratios are likely to drop.
Then again, maybe the study could be a little more systematic (as other comments point out). I'd like to see larger numbers, if only because it would help statistic significance.
It seems really good news to me actually.
Not much consolation for the families of the dead patients that your software misdiagnosed.
Another interesting question, if you had unlimited resources for studies: how does AI compare to 1 or 2 poorly trained 3rd world radiologists?
Today 94% of 36 AI systems evaluated were less accurate than a single radiologist.
FTFY
In the beginning of the eDiscovery industry was here too. Lawyers read through insane amount of PDFs, TIFFs, emails and they were better than the AI/ML software presented.
Today most courts will reject human eD if AI/ML was available, and no proper legal team will use lawyers to search through terabytes of data 'manually'. The false negatives/positives are significantly better with AI/ML.
1. Presumably we are most interested in the 6% that are more accurate.
2. The more interesting part, although not entirely surprising, is that 94% of research is overstated, at best.
Why? At 94% failure, this 6% sounds like lucky guessing by a computer.
From the actual article's abstract, here is what was meant:
> Thirty four (94%) of 36 AI systems evaluated in these studies were less accurate than a single radiologist, and all were less accurate than consensus of two or more radiologists.
I wonder if the HN title was written by one of these AI systems.
That's the main conclusion to draw. Better studies are needed until we can know whether Geoff Hinton was right [1].
____________
[1] “Machine learning pioneer Geoffrey Hinton said, ‘If you work as a radiologist, you are like Wile E. Coyote in the cartoon; you’re already over the edge of the cliff, but you haven’t looked down.’
“Hinton went so far as to recommend that med schools stop training radiologists right now.”
https://blionline.org/are-cpas-like-wile-e-coyote-off-the-cl...
It appears that AI screening of breast cancer has reached the "not always better than a human expert" stage which is pretty good, that would be sci-fi like 10 years ago.
Makes me wonder how a human-AI "centaur" or the averaged output of multiple independent AIs would do, since the standard already seems to be multiple humans to catch everything.
a different question I would have is, do the AI models know when they are unsure. If so, you could use the AI, and keep the prediction when it is certain, and call a radiologist otherwise. Obviously, in this scenario, you have measured the accuracy of the certainty prediction.
For example, check the photo-recognition systems where minor changes to the picture change the recognition from a parrot to a Ferrari.
That is to say, pair human judgement with AI. It may not be as good as two humans making a judgement call, but it's probably better than just one.
Furthermore with AI you could do this on 100% of exams, as compared to the small amount that go through peer review.
When in reality 1 AI consistently beats all Human players every time.
Also check if ensembles of the lesser-performing systems (with other systems or with radiologists) can do better, and whether any of the systems can rapid-classify 'easy cases' so expensive radiologists need only check the tougher ones.
In my mind, many AI systems will still be faster than a single radiologist (though this assumption is a completely random guess from my part, I don't know how fast these analysis robots can work.)
While I agree wholeheartedly that there are a lot of false claims about medical AI swirling about, I don’t trust the authors of this paper to tell us where it’s happening based on the mistakes I’ve seen here
“ The remaining studies used enrichment leading to breast cancer prevalence (ranging from 7.4%26 to 73.8%37), which is atypical of screening populations. Five studies used reading under “laboratory” conditions at risk of introducing bias because radiologists read mammograms differently in a retrospective laboratory experiment than in clinical practice. Only one of the studies used a prespecified test threshold which was internal to the AI system to classify mammographic images.”
I’m very close to the authors of McKinney et. al and I happen to know that 2/3 of these claims are either false or specious:
1. The operating point was decided ahead of time before the reader study that was conducted in the paper. I was literally there and saw Scott choose it and I saw his methodology for doing so. So this claim that they did not use a prespecified threshold I have first hand evidence of it being false.
2. Enrichment:
Enriching for positives is absolutely a standard practice and it has mathematically 0 impact on computed metrics. It’s not possible to conduct a reader study without enrichment because cancer is so rare that you would never be able to recruit readers to your study. It just doesn’t make sense at this stage of research to not do enrichment because regardless of what you do the study will still be retrospective.
The paper also shows that results on the unenriched data is basically the same.
I think the understanding that we had at Google and that was mentioned in the paper is that these types of retrospective studies are not sufficient to deploy the AI. It’s probably a good thing this paper is stressing those points.
We need prospective studies, but practically speaking it makes sense to first do really solid retrospective studies. Otherwise a hospital system is not going to just let you run your AI on their scans.
Google is proceeding to do this sort of testing with the work that this paper is just falsely thrashing:
https://news.northwestern.edu/stories/2021/02/artificial-int...
I feel that this article has good intent (stressing the importance of prospective trials rather than just retrospective studies)
However, the execution is significantly flawed and the outcome seems to be in this particular forum people believing that AI will never work for this application.
These things take time and there is a pipeline of testing and trials that takes many years to go through.
1. Research and development
2. Retrospective testing: as a replacement system (so you don’t have to figure out the UX of how to make it help the human)
3. Retrospective testing: as a helping system (generally scientific papers have been less interested in this step and it would be nice to see more interest here)
4. Prospective testing as an assistive system.
I think the Google system is on step 4 now, and this paper is looking at evidence published from stage 2 and claiming the system does not work.
It’s probably a fair claim because it has not completed all the testing to be deployed. There is as of yet no intent to deploy it as a replacement system, but rather as some sort of assistive system so that prospective evidence can be collected.
However the way they’ve done this is by misrepresenting the paper, and I really think they’ve overstated their case here.
A couple background facts:
1. Mammographer subspecialist salaries are now the highest of any radiologist (median ~$440K) - https://www.auntminnie.com/index.aspx?sec=ser&sub=def&pag=di...
2. Mammography is one of the easier fellowships to undertake and is "lifestyle-friendly" overall (no call as an attending!)
3. Mammographers are highly in-demand, even in desirable locations.
So it's (relatively) easy, very well-paid, and you can live where you want. What's the catch? Well...
1. Mammographers are the most likely to be sued. Missing a mammographic diagnosis of breast cancer is the #2 reason for medical litigation, with some sources citing up to 10% of radiologists per year being named in a mammography litigation case. See: https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3138733/ and https://pubmed.ncbi.nlm.nih.gov/22596034/
2. Mammographers directly interact with patients to relay findings and biopsy lesions, and they must collaborate with referring providers and pathologists. Breast cancer is a highly emotional topic in the United States due to a combination of public personas ('Angelina Jolie effect') and political maneuvering (see the history of the federal Mammography Quality Standards Act and Program - MQSA). You can't be a "behind the scenes, in the dark" radiologist if you go into mammography.
From a computer science perspective, I one-hundred-percent agree that mammography is a "solvable" problem by AI. Standardization of mammography images (CC and MLO screening views) alongside BI-RADS makes for an amazingly well-labeled dataset. I fully expect AI to surpass the abilities of one-or-more mammographers in the near future.
And yet, amongst radiologists, mammographers are the least likely to be replaced. Why? For the same reason mammography is lucrative to begin with - Litigation, emotional patient-doctor interactions, and the social/political climate. Any one of these is a steep hill by itself. Most AI companies want none of the legal risk, and practically zero of them offer patient-facing results. And, I wonder if any of the AI startups have considered hiring a celebrity as their spokeperson. Perhaps politics/legislation might work in favor of AI if it becomes standard of care, but I wouldn't want to bank my startup on a federal regulatory change.
The other issue with Mammography is that even human accuracy isn’t great. which is why they keep getting sued.
So what specialty did you chose ? Call me if you are interested in body / oncological imaging.
This is like the comparisons of 200 top hedge funds to index performance and concluding that hedge funds can’t beat the market. Ignoring the other conclusion: most hedge funds suck
Edit: or most hedge funds are fraudulent?
I think you've misread what conclusion you're supposed to draw from this kind of assessment.
For both the hedge funds and this the conclusion isn't that the job is impossible - but that we lack the tools to predict who will do the job. You can't know which hedge fund will outperform (though some will) and we don't seem to be able to predict which AI models will perform well enough to use. Also that your odds of picking a winner "from the field" are quite bad.
It doesn't mean there's anything wrong with using a hedge fund - but there's a level of risk that one shouldn't ignore.
Like, my conclusion from this is that I would not accept a pure AI solution (because most of them are bad), but I would be interested in AI assistance. For many people, the promise is in replacing not supplementing doctors and this is the same as failure.
Oh no, that could be the conclusion. Ever study quantum physics? Lots of people have tried to come up with some kind of model with inner "hidden" variables that tries to make HEP physics into a completely deterministic, almost classical, theory. By your logic, if we just trained and AI with enough subatomic interactions, it would eventually be able to predict with 100% accuracy the results of a quantum process.
That's utter bunk.
Markets are the same way: people think there must be some rules that can be worked out to make them 100% predictable, when in reality there is an element randomness that makes cryptographers jealous. It will never be possible to predict the next number in a random process. If you flip a coin 99 times and get 99 heads, you still can't predict what the 100th toss will be (and before someone "well ackshually"s with something about an unfair coin: don't)
I mean, the job certainly could be impossible, but those studies don't prove it. As it relates to the discussion at hand we know that we can do better at evaluating mammography because humans do it.
Maybe I am not understanding what you are saying?
So 2 of the 36 were better than one radiologist, none were better than two radiologists collaborating.
I asked about AI and they couldn't use the one they had on the cyst because it was not trained for ankles.
Imagine like "94% of all attempts at heavier-than-air flight crash on first attempt". Sure... kinda misses the point.
I am sure the average radiologist, tired and bored at 8am in the morning, is much worse.
The 6% only scored better when compared to a single radiologist review. That's why we use double reviews.
We being Europeans. Double reads are not standard in the majority of the world, including the US.
> Thirty four (94%) of 36 AI systems evaluated in these studies were less accurate than a single radiologist, and all were less accurate than consensus of two or more radiologists.
(Or just 94%, and then have people say 'but N= only 36!'.)
Percentages for small numbers, in particular below 100, are just annoying and often designed to mislead. (Though I'm not suggesting that here, it's standard in medicine if not other fields.)
I work an an AI company where we screen for Diabetic Retinopathy. The company is a decade old, and has validated in huge studies (over 100k patients). It's a hard problem that we have all worked very hard to solve it. At the same time, it's easy to build an AI tool that looks at an image and says healthy or unheathly (or poor quality image). So the bar is low for making something that is appears functional.
But hardly anyone makes good AI tools, so studies that look at different AIs systems see lots of the bad ones. It's always a bummer when the headlines are all dismissive of AI in general. Googling "diabetic retinopathy AI" the top result is "Artificial Intelligence Falls Short in Detecting Diabetic Eye Disease" (https://healthitanalytics.com/news/artificial-intelligence-f...), yet if you read the article it says one tool is better than humans, which to me, is the real takeaway.
"Please use the original title, unless it is misleading or linkbait; don't editorialize." - https://news.ycombinator.com/newsguidelines.html
It strikes me, without any real knowledge but this is the internet, that it just goes down the typical rabbit hole of human specialization. Make the data portable, move the data handling overseas, add computer optimization over time. You do have to wonder if radiologists (any here?) really need general medical knowledge or if high specialized training with pictures would be enough.
When you see a bear riding a bicycle, you don't complain that the bear rides poorly. You're in awe that it can ride at all.
Wait until people who know what they're doing move in. Right now a lot of those people are in academia, trying to squeeze the last 0.01% of lift out of Imagenet dog breeds via architecture searches at $100K a pop.
It's not like we haven't done the AI Winter thing before. People should be used to this stuff by now.