MIT AI tool can predict breast cancer up to 5 years early
techcrunch.com
techcrunch.com
From [1]:
> Class imbalance can cause ROC curves to be poor visualiza- tions of classifier performance. For instance, if only 5 out of 100 individuals have the disease, then we would expect the five posi- tive cases to have scores close to the top of our list. If our classifier generates scores that rank these 5 cases as uniformly distributed in the top 15, the ROC graph will look good (Fig. 4a). However, if we had used a threshold such that the top 15 were predicted to be true, 10 of them would be FPs, which is not reflected in the ROC curve. This poor performance is reflected in the PR curve, however.
The authors seem to be aware of this in the supplement and also evaluate performance by a hazard ratio they define:
> We calculated the ratio of the observed cancer incidence in the top 10% of patients over the incidence in the middle 80% and referred to this metric as the top decile hazard ratio. We calculated the ratio of the observed cancer incidence in the bottom 10% of patients over the incidence in the middle 80% and referred to this metric as the bottom decile hazard ratio.
However, binning is a form of p-hacking [2]. And I'm still wondering why they don't just post the Precision-Recall curves.
[1] https://doi.org/10.1038/nmeth.3945
[2] https://doi.org/10.1080/09332480.2006.10722771
[Edit] to add link to [2]
Edit: From the paper:
> A deep learning (DL) mammography-based model identified women at high risk for breast cancer and placed 31% of all patients with future breast cancer in the top risk decile compared with only 18% by the Tyrer-Cuzick model (version 8).
So better than before, but still only detects 31%. If I'm reading correctly, it's 95% correct? I guess that means 5% false positives? That wouldn't be bad.
Due to Bayes' Theorem, this can result in the majority of patients with a positive result being patients without cancer. Depending on what treatment is taken, and the psychological burden of erroneously thinking you have cancer, the negatives of the additional screening could outweigh the positives.
You have to be really careful with FP rates in screening. Of course the answer to this could just be "don't biopsy due to only this result" but at some point that renders the approach pointless, if you can't find a cost effective way of mitigating the FPR.
950 true positives
~50,000 false positives
False positive breast cancer is pretty bad.
1. The cost of a false positive is low. 1a. The cost of a false positive is retesting (what is retesting? Re-mammogramming? Isn't healthcare in USA expensive?) 1b. Observe (so future re-mammogramming and hypervigilance).
2. Additional to the cost of 1a and 1b, cease consuming sugar. Which is a non-insignificant lifestyle change.
What I'm saying is... the statements "Rerun expensive tests and maintain hypervigilance seems like a low cost, in addition to major lifestyle changes that affect every level of one's socioeconomic place (eating out with friends/family, cost of food, food availability, habit breaking, possibly culture clashes, etc)." seems pretty much equivalent to absolute nonsense.
You can be told the risk factors for everybody, but when you are told the risk factors tailored to you it becomes dangerous? Should we tell everyone they're going to live forever so they live in ignorant bliss while healthy?
I can see being unprepared for having that information, but I don't think the solution is not having the information.
There have been studies into exactly this: https://www.ncbi.nlm.nih.gov/pubmed/22859786
Specifically on breast cancer screening, the National Institute for Clinical Evidence in the UK has published management guidance that is quite interesting: https://cks.nice.org.uk/breast-screening
They wouldn't be wold they were positive for cancer. They'd be told that the computer predicts that they are at very high risk for breast cancer. That's knowledge that I think people should have, personally.
I’m not disagreeing that there usefulness to tests, but there is danger too.
https://youtu.be/iUqgTYbkHP8?t=15m37s
The reason most people die from pancreatic cancer, for example, is because we almost always detect it in a late stage.
https://www.sciencedirect.com/science/article/pii/S221353831...
Granted, a small proportion of these cancers are caught an early stage. But I wouldn't say most people diagnosed would live five years longer if they'd caught it early. And that's ignoring lead-time bias [1].
This has to be taken into account when deciding how much screening to do.
https://www.webmd.com/prostate-cancer/prostate-cancer-surviv...
"Pooled data currently demonstrates no significant reduction in prostate cancer-specific and overall mortality. Harms associated with PSA-based screening and subsequent diagnostic evaluations are frequent, and moderate in severity. Overdiagnosis and overtreatment are common and are associated with treatment-related harms."
[...]
"Screening resulted in a range of harms that can be considered minor to major in severity and duration. Common minor harms from screening include bleeding, bruising and short-term anxiety. Common major harms include overdiagnosis and overtreatment, including infection, blood loss requiring transfusion, pneumonia, erectile dysfunction, and incontinence. Harms of screening included false-positive results for the PSA test and overdiagnosis (up to 50% in the ERSPC study). Adverse events associated with transrectal ultrasound (TRUS)-guided biopsies included infection, bleeding and pain."
https://www.cochrane.org/CD004720/PROSTATE_screening-for-pro...
You might be right that many people would be better off not knowing but I’d rather we took a step forwards than backwards.
If this model is equally accurate for black and white women that means that either race is not a factor in predictability, that it is a factor but easily adaptable into a model, or race is a factor and they’ve reduced their ability to diagnose one group in the name of equity.
The linked article suggests accuracy gains are due to better risk models, that use more than age. I’m not sure if that means it’s tied into the image neural net. Would like to see false positive rate too.
Once upon a time, a necessary precondition to call something AI was that it should be something where there is at least the hope that it could one day generalize to pass the Turing test or something along those lines.
Medical diagnostics is one of the primary applications of pattern processing, and since it's pretty damned impressive as it is, it's a bit pointless to try and make it even more impressive by suggesting that you might one day enjoy a chat with your medical diagnostic tool over breakfast, exchanging views on how the Knicks' season is shaping up... (Which both the informed readers, and the people writing this, know pretty damned well is never going to happen, and was never intended to happen).