Artificial intelligence versus clinicians: systematic review
bmj.com
bmj.com
This is fine; of course the standards for "is this publishable" are vastly lower than "is this clinically applicable", but it unfortunate how often the lay science press fails to make this distinction.
The other thing is that raw performance numbers, even if properly validated, are only a small part of the story. We've had ML techniques that did better nominal sense/spec than average clinicians for very specific tasks for 20+ years now, but the workflow impact, liability, etc. combined has sometimes limited the implementation, and even if you do roll it out the actual impact can be minimal.
I certainly don't disagree that there is potential for impactful use of AI in some areas of healthcare. Some of them just aren't going to practically happen without a) fundamentally reworking our approaches to data access and sharing and more importantly b) spending a lot of resources on quality labeling.
Regardless of whether or not those preconditions are actually met, we'll get a lot of breathless claims on small retrospective sets, of course :)
That said, the problems are harder than appears to your typical ML grad student, and structurally we aren't there with data. There is no reason currently to think self supervised with "get there" in the near term, but even if you've assumed infinite labeling resources the data sets are for the most part still too disjoint and unavailable. Some strides have been made in some areas recently, but a long way to go. The algorithms part isn't necessarily that interesting really, building up the supporting infrastructure is the real problem.
Well from a machine learning perspective it will probably be in all cases where enough, qualitative data is available and not a lot of contextual knowledge is necessary that cannot be represented in the data. My guess is that this will be the case in a vast majority of cases and it's just a matter of time - but I'm definitely biased as an ML researcher :).
On another note and speaking from personal experience, NLP as well as patient information and note processing is something that might be less glorious than AI analysis of MRIs for example but which would have a far greater impact on physicians’ day to day life. Most residents spend like 50% of their time fighting applications with some of the worst UX I have ever seen just to get some basic information, log something or schedule a simple procedure.
On the other hand there have been ML/AI systems since the mid-late 90s at least, and they haven't for the most part shifted clinical practice much. The reasons for this have not shifted radically in the last decade, although some are moving a bit now.
FDA approval to market doesn't mean anything beyond you have an argument that you haven't introduced new risks (safe) and in at least some cases you can improve something without making others worse (efficacy) but this may in a very narrow indication.
I guess my point is, the purpose of the FDA has never been to ask "does this really work better? should it be the new standard of care?", but rather to balance the risk in new technology or drugs against the potential reward. Even with a PMA it's risk reduction, not proof of how it will play out in the wider world if interactions and practice.
In terms of crisis, the medical system just doesn't scale. You think leaders were worried about running out of ventialtors?
No! You can start pumping them in hundreds of thousands if you really need it.
The problem is that you have only so many doctors and trained nurses to work with. Limited capacity. Training takes 10 years.
No scalability.
It doesn't matter where the tech is right now. The world will push 100% for whatever can be done via apps, sensors and unskilled workers.
Check out Jonathan Rothbergs vision for a diagnostic kit by everyones toothbrush
However, across a population, nobody is being harmed by the ML analysis, and it can pick out patterns a person could not at speed and at scale. Spending research to get your individual diagnostic ROC curve a few percent higher is a waste of time if the consequences of it being wrong are greater than the time savings it provides.
It's fine for veterinary medicine, and maybe prisoners and keeping people with hidden symptoms out of crowded places, but for regular people, using ML has got a host of ethical problems that aren't being dealt with at the policy level. You can use rules engines (literally, DROOLS) instead of non-deterministic ML schemes to do diagnoses and prescriptions.
Have said before, ML is only useful for problems with asymmetric upside, and it is worse than bad for ones with asymmetric downside.
Excuse me but did you just compare the application on prisoners and animals versus whatever "regular people" are supposed to be? What the hell?