Machine learning is booming in medicine, but also facing a credibility crisis
statnews.com
statnews.com
For one, it's actually difficult to interpret and find signs in radiologic images. Obvious signs are obvious, but there are others that could be image artifacts, or just indolent variations, or point to something serious. Even with a generally good accuracy, it'll be hard that a general model performs well on those anomalies with low prevalence.
Second, radiologic signs are just signs. Most diseases are diagnosed with more than just radiologic signs. Most signs are compatible with a lot of diseases. If you see a model that pretends to diagnose a certain disease, well, they're looking at it wrong.
Third, you need a way to have responsibility for diagnoses, and a way to find and correct errors. I don't think that's possible with an unsupervised AI, you'll always need a doctor there to check the image and verify the output. There won't be much savings there. Whenever someone says "AI is going to revolutionize medicine", it's really hard to believe them. I mean, you just have to look at EKGs, modern machines can detect anomalies but doctors still learn how to interpret EKGs and double check what the machine says. It's a help, but not a replacement.
ML is capable of doing that if you train it to do that.
I don't get this perfection requirement that gets put on so many systems here. No system is perfect, the question is what failure rate you accept.
So yeah, maybe you get a model with 5% false negatives where humans have 10% false negatives. That's good. However, when you go to apply that model in reality, what happens to those 5% where model and doctor disagrees? First, we should think about the kind of disagreement. In most of those cases I'd bet that the doctors see something wrong but can't say what, not that the model says "this is bad" and the doctor "it's not".
Second, we need to think about what happens when doctor and model disagree. As the model is not 100% accurate, the doctor doesn't know if the model is mistaken or if they're the ones mistaken, so they'll probably order more tests anyways. If it can be something serious, it's worth it to do an extra test to make sure. They'd probably ordered those tests anyways if they weren't sure of what was happening, model or not.
So what did the model for a single test change? Did it really change the diagnostic outcomes? What's the actual benefit of the model? How much it's worth to get from 10% to 5% false negatives with a model for a single test if just adding more tests (say, with the same 10% false negatives) to the mix can give you a 1%, 0.1% false negative rate?
That's my point. Unless accuracy is really high, ML models are not going to remove uncertainty in diagnostics, few diagnostics consist of just a single test. Benchmarking models against human performance in a single test is not a metric that can drive implementation in the real world.
But in general, these are matters of life and death, literally. You're still going to have someone looking at the images and verifying the output. That limits a lot the potential cost benefits of these applications.
My point isn't that you need 100% accuracy. It's that a diagnosis is a process, not a single test. If your model applies to a single step of the process, and if it doesn't remove enough uncertainty about that step, you're still going to continue the diagnostic process and the model is not going to change anything really.
E.g., “it will have no effect on outcomes until pigs fly.”
There will always be uncertainty, and uncertainty isn’t the only relevant parameter.
But it's a really important one and a lot of medical ML research doesn't seem to address in the proper sense. A very simple example: a ML model that classifies lung nodules as benign or malign, with 95% accuracy vs 70% accuracy of regular radiologists. Very good, right? But for actual, real world results, you need to see how the patient outcomes change. If the patients where the model and doctors disagree were going to have extra tests or followup regardless of what the model says, the model is not actually offering anything new despite the increase in accuracy.
So no, it's not absurd to say that models that work for a single type of test need to be very, very accurate to actually bring changes to the procedures that are worth the investment. What matters is whether they actually change patient outcomes, and sometimes it seems like the ML researchers barely consider it.
That's a big if already. Not to mention that a lot of times you don't want to give unnecessary radiation to patients, so they wouldn't be interested in making, say, X-rays more accesible.
> screening tool that less qualified medical professionals (midwives, nurses, GPs) or even the patients themselves can use
But how beneficial would that be? When you need an scan, it's probable that something weird is happening, and you'll probably need more tests than just a scan, and also some specialist that knows which tests to order and what conditions to consider.
Once you think of it in that way, it makes less sense. You'd be investing money, resources, training, time into a solution that might benefit only a small subset of patients (those with something weird but that can easily be discarded by a cheap scanner with high certainty without needing an specialist), without even being sure of what actual clinical benefits you would produce. It's far better to invest and research in other areas.
If it's a regular screening test rather than a response to the patient's complaint, then false negatives wouldn't be as much of a problem because at least it's better than nothing.
> it's better than nothing.
Not necessarily, that's the issue with screening asymptomatic people. You have to balance the consequences and rates of false positives with the benefits of true positives. If ultrasound screening mostly catches indolent diseases, or those where catching them so early doesn't affect outcomes too much; and you have a lot of people going through unnecessary/risky tests, you might end up doing more damage with those screenings. It's not so easy.
Tl;dr: the first diagnosis tends to stick with the patient.
Newborn babies undergo a whole lot of simple and inaccurate screenings. I guess those are cases where the harm due to inappropriate treatments is lower than that due to leaving diseases undetected.
Either way, I'd rather the technology exist and then people work out how best to use it, rather than the decision being forced on us by it not existing.
Uh, TBH no professional so much as glances through the automated EKG summary. It's utterly useless and could be deleted with zero consequences.
I'm curious though. Can you elaborate on the types of issues you have encountered? What brand of equipment were you using?
No chance that a dinky ECG from 10 years ago is running a neural net inside.
Radiologists say this until they get malpractice lawsuits. Seriously, I once injured my wrist and the first radiologist said it was just a sprain, but no surprise the hospital was using a outsourced...overseas provider for that quick analysis. The hospital had a inhouse radiologist review it as a course of business later on in a few days and the hospital panicked with phone calls to get me back in because there were numerous wrist bone fractures and they took the first step towards liability issues
Could an AI do it? Sure. But you sure as hell need a second opinion
That would even drive down the possibility of malpractice.
My question is very simple, at what point Shoddy rules-based engine becomes "AI-driven approach"?
Which makes me wonder how much of the supposed magical 10x value of AI actually just comes from getting your shit together in terms of ETL and data pipelines.
ML folks today are under immense pressure to show small-digit improvements over existing methods and they're using all the wrong techniques (p-hacking, massive hyperparamter runs, press releases that overstate the impact). It's really sad to see all the snake oil.
It becomes very difficult to classify real orgs using hardcode stats from fake ones. When I hear "we are solving x with AI" or "AI driven", it gives me jitters.
Well, it's _something_ different. It's not "code" or "programming", so whatever you want to call it, that's what's driving the investment hype.
Back in the early 20th century, chess was considered to be something only an intelligent agent can do. A few decades later programs became competent enough that it took a grand master to beat them. Today, it's hopeless to try and beat a chess program as a human.
The next barrier was NLP and understanding/generating text and natural speech. Speech recognition worked reasonably well 20 years ago and is close to human level today. Speech generation - given enough processing power - is now perfectly natural. Text understanding and generation is very close as well.
The development of AI bears an interesting resemblance to evolution in that sceptics are busy pointing to "missing links" and as soon as that link is discovered, they chase to the next one.
Intelligence itself is defined poorly enough as it is, and watering the term down by slapping the "AI"-label on everything doesn't help that. On the other hand, intelligence is a spectrum and on some aspects of that spectrum, machines have already surpassed humans decades ago (think of calculators, chess, memory, searching and indexing, etc.).
IMO, the most important distinction to be made is the difference between AGI and AI. A transformer model is AI in every useful definition of the term "intelligence", but it sure isn't "general intelligence" if only for the fact that it cannot actively query its environment for additional information and has no continuous stream of "consciousness".
Before we can dismiss a system as not being AI, we need a sharp enough definition of what we would define as AI first. A "I know it when I see it"-type of definition isn't helpful and always keeps the door open for the ultimate rejection: "but it still hasn't got a SOUL!"
Are people in serious finance bothered that there are a bazzilion of scams every day, including aforementioned blockchains? No. Why should competent AI practitioners care as well?
AI took decades to recover from its first boom-bust cycle. No matter how competent you were back then, getting hired or funded to run your AI business/academic project/whatever was hard because the first wave of AI failed to deliver.
Sure, now AI/ML/Data Science is very commercially/acadamically viable, but the hype has also grown spectacularly. If the general public fails to manage expectations, another AI winter cannot be ruled out.
This is why China will win the AI Age.
Data is the new Oil and USA is still clinging to the old oil while China has AI as number 1 priority.
If just data collection would be enough then all crime should have been eradicated in the US by how much data NSA has.
Is the same true of China, or will China have to rely on its homegrown cadre of scientists? Granted, there might be a lot of talent in a population of 1,4 billion.
Science is about truth. Dictatorships are about appealing to the dear leader. That's why China will never catch up to free and open countries.
Lets go.
Universities get these students to pay full tuition. Professors get graduate workhorses for research. Companies get cheap labor, low environmental standards, etc so they get a larger profit margin on their widget.
China does have a sphere of influence but the total population of china + friends is well below 3 billion people and will continue to drop as a percentage of the world population.
This is exactly my sense of the problem. Not saying that this is the only problem, just that this seems like the biggest immediate problem.
There are a bunch of questionable ML startups that try to do something with, say, a model trained on ImageNet. You can get pretty far starting with ImageNet. ImageNet was made with Mechanical Turk, but there is no Mechanical Turk for radiologists. If you work really hard you might get patient & treatment notes but interpreting those notes presents its own problem.
If there was a system for sharing images that had existed before Covid and enough data had been contributed that reviewers could demand evaluation on some predefined test sets, then a lot of these inconsistent evaluation issues could be weeded out.
The article mentions federated learning, but I feel like that's solving a different problem than reliable evaluation.
There is no incentive in publishing stuff that don't work even though they were reasonable things to try. This indirectly pushes a lot of bad/flawed/incomplete papers and results to be published. Even if your methodology is flawed but you try hard enough, you have a >0 probability of getting your paper published somewhere.
As part of a solution, I think we would benefit from the existence of prestigious journals for negative results, which would incentivise also publishing what doesn't work. This way researchers wouldn't have to try massaging their data and experiments until it looks like it works. They could just publish that it doesn't work, and it would be good for them too.
It seems like the only result of consequence has been Zork, developed on the DM machine at the MIT AI lab in the 70s. No, it was not developed as a medical application but I believe that machine was owned by the medical decision making group.
Stanford’s Knowledge Systems Lab (Ed Feigenbaum’s lab) was next to and affiliated with the Medical School too.
With the rampant lack of statistical rigour it blows my mind they get published. This great paper shows that reported accuracy plummets as dataset size increases: https://hal.inria.fr/hal-01545002/document
Still, I'm thinking that as it improves, it's going to show that doctors are not that good at their job on average, and that's going to be fun to watch.
Just like for wikipedia.
And a high volume if diverse quality data is fundamental: https://en.capillary.io/posts/ai-medical-diagnosis-guide/
Because Americans are trying to solve their healthcare crisis by attacking the doctors and literally not the entire rest of the industry which controls the spiraling healthcare costs. The end of result if we manage to eliminate doctors is still spiraling out of control healthcare.
In other places, e.g. Mexico, India, the radiologist can very much be talking to the patient, and can have information from earlier scans.
Medical AI is trained on labels generated by doctors. Can you explain how it will exceed the performance of doctors on average? Are you assuming that the labels will be generated by the "top x%" of doctors? If so, how will you identify those individuals? Or is there some other mechanism you're expecting to improve the performance?
How about starting with a problem to solve? You don't need anything but a pen and paper. Gotta start somewhere.
In medical AI, you really need datasets. Of the most expensive kind, and lots of it.
You can fund it yourself with basically any income. Of course you will never get anywhere with an attitude like the one you currently have.
It's not the skill or resources that is the problem, it's the mindset.
You confuse results I observed with my "attitude". I am trying quite a few things. I am only observing negative results. That might look like a negative attitude, it is, however, reality.
You could try to start-up and then plan to get acquihired. Or go and do a PhD. It's never too late for more schooling.
I hate to ask why you believe this, because I see VC money being thrown at borderline low-lifes with mediocre ideas. There are so many "angel investors" on Twitter of companies I've never heard of...
The SV bubble can warp your brain, but an outsider's tip is to focus on your idea before you start worrying about funding it...
The chances of succeeding are excessively low, and the funding situation in Germany is very different. You mostly need to have revenue before you get funding. And when you aren't even employable on the normal labor market, then it is a lot harder still.
A friendly suggestion: this attitude may be holding you back more than you realized before now.
What to do instead, is repeatedly ask yourself: "how can I do this?"
Apply this to many specific areas:
- How can I change my CV so it comes across differently?
- How can I begin a small startup without money or other resources?
- How can I move towards my goal without any new degree?
- How is my training in veterinary medicine an asset?
- How can I do medical ML for veterinary medicine?
- Et cetera...
You have a choice now to dig into your heels and say to yourself reasons why none of what I say above will work, and why your situation is truly and utterly hopeless.
OR you can say different things to yourself, maybe get a different result.
The objective reality is that I did not succeed. Somehow people think accepting that means I have a negative attitude. I haven't given up on ever working again, but I'm not delusional enough to think I haven't been unemployed for a long time.