Doctor's historical success rate exceeds expected success rate on average => Good (or lucky) doctor.
Doctor's historical success rate exceeds expected success rate on average => Good (or lucky) doctor.
Genetics is oversimplified to non-physicians. It's cool that we can diagnose and predict the likelihood of getting Huntington's disease using our knowledge of genetics, but extremely few diseases are this simple. There are huge swaths of the human genome that we don't understand but are likely playing some important role in the regulation of other genes and diseases. We are nowhere close to being able to look at a patient's genome to predict anything useful outside of a handful of exceptions.
Patient histories are honestly often garbage—I say that as a physician. I look through dozens of patients' charts every day, and there are constantly errors, incomplete documentation, and fragmented records across multiple institutions. Just last week I read a chart for a patient who had a documented hysterectomy from years ago. The brand new CT scan I saw showed a perfectly normal uterus. Once something goes in a patient's history it's nearly impossible to correct or remove. If some doctor from ages ago said the patient is allergic to medication X, but the patient denies it, what do I do? Usually, we opt to leave the allergy listed out of fear of the consequences if the patient is wrong.
I had one doc I had only met remotely who entered a height into my online, shared record that was several inches shorter than I was...back when I was an 11-year-old child. It's been 25 years since I was that short.
I pressed them that they could at least have asked me how tall I am, or even consulted the previous entries by other doctors in their same system.
There wasn't even an attempt at an excuse. It was plain, simple negligence.
That doctor had also been insisting I needed to let him do an exploratory surgery, despite never even having had me come to the office in person, so I noped right out of there and started telling that story to everyone I thought might consider seeing him.
All of this talk about how "hard" the statistical analysis is, is strange to me. Maybe "advanced" would be a better term? If you get a patient with a contradictory medical history that somehow also contradicts what they are telling you, simply adjust your expected chance of success appropriately (to zero perhaps). In that extreme case, if you get a good outcome, congrats you got lucky. If you don't, it should have 0 impact on how you are evaluated as a doctor.
Institutions face severe penalties for wrongfully sharing patient data, so most opt to just not share any data. Any research that is performed is done internally on local populations with de-identified data sets. A few brave institutions go well out of their way to create and share de-identified data sets publically, but these data sets still undersample the general population. This is a critical problem because certain diseases are highly prevalent in certain regions (e.g., Lyme disease in New England) but unheard of in other regions (e.g., Lyme disease in Colorado). If your ML model is trained on data largely from New England, it's going to diagnose a patient with the classic "target-shaped" rash with Lyme disease even if the patient is from Colorado (high false positive rate). If the model is trained on data from Colorado, it will underdiagnose Lyme disease in patients from New England (high false negative rate). The only way I know to overcome this problem is to create even larger data sets, but this just isn't possible with data privacy laws.
My understanding is that high dimensionality in the domain is exactly where ML excels (so add location as an input), and that is exactly what a medical diagnosis involves. Perhaps the legislation will get there one day.
https://inews.co.uk/opinion/nhs-data-shared-third-parties-we...
Anyway, its out there now.. so watch this space, I guess!
Most big medical co do a lot of data science (Kaiser and others). Very efficient from a managerial pov. Totally useless, medically speaking.
In that I have several CDC funded projects using machine learning in medical analysis and outcome prediction.
Even on extremely well curated data sets this is a fairly hard problem.
Read “The Alignment Problem”, a very good just above pop sci level book about machine learning. They have one example where ML determined seniors with COPD were at reduced risk from pneumonia, and obviously non-sensical result. Patients with COPD wound up in the hospital at a lower rate than average because doctors know they need careful attention right away.
I’m reluctantly pessimistic about humanity’s near-term capability to appropriately weight input from technology as imperfect but inscrutable as current ml.
This is a little like programmers calling the programming language they invented and write in every day garbage. Separate from patient-reported histories, you medical doctors are the ones documenting these histories and hold the decision making power for how it’s done!
Your first sentence would be the equivalent of the inventor of patient history data keeping calling patient history keeping a garbage tool. Which is actually something that would be totally OK to do and say if suffixed with "and unfortunately so far nobody has come up with a better tool and it's not for lack of trying".
I'm assuming you didn't mean "you medical doctors" in the sense it's easy to read in. In any case, what you are doing here is telling one doctor that he is bad at the medical history writing and reading job when in fact he is the one telling you how he is able to spot other doctor's mistakes and trying to correct them. This is like telling one developer that he's bad at his job, that "you developers are the ones writing bad code and hold the decision making power for how it's done" when that developer is actually someone that tries to make things better both through his own maintainably written code (medical histories) and helping others in code reviews to make their code better and not let bad code get into Prod (finding errors in existing medical histories and trying to correct them).
That can be very discouraging, being thrown in with the bad apples. And even good apples can have a bad day or misunderstand something. But I guess you are perfect and have never produced a bug in your life.
There is certainly noise in healthcare data especially when patient-reported, but is it noise to say that a patient having X procedure later does or doesn’t have serious complications? Analyzes of medical care and their consequences can be evaluated and it’s not noise
And big healthcare data has lagged, partially because privacy concerns trump sharing. There are companies selling anonymized medical records for basically every American now though. Big data is coming
Big _bad_ data... Let's see how we fare in 5y, then. My prediction as a clinician with a special interest in stats: close to zero medical progress. But insurance priced by a ML algorithm, and much greater efficiency in coverage and claim denials.
https://www.healthcare.gov/how-plans-set-your-premiums/
The more likely use case for ML is detecting insurance fraud patterns.
I think that's covered under "much greater efficiency in coverage and claim denials."
The codes we do have are mostly CPT4 and ICD-10 for billing purposes. Those are generally pretty accurate, but not detailed enough to reliably assess whether one surgeon is better than another at a particular procedure.