Google Med-Palm M: Towards Generalist Biomedical AI
arxiv.org
arxiv.org
Ah, but copyright/patents/IP. Well, IP was created to foster production of useful immaterial stuff. If you now want to use it to hinder production of useful immaterial stuff, you can go to f*ck yourself if you ask me.
Ah, but lawyers and liability. I propose only that the doctor is required to consult the LLM. Easy to log and verify. All liability stays at the doctor who makes the final diagnosis.
With 60-70% correct rates on most training sets and 0.63 critical errors per report, for any physician not very well-versed with the limitations of LLMs, this is more of a liability than an asset. Some of the biggest barriers to care are cognitive, such as anchoring or availability biases. LLMs in their current state will only muddy the water.
Good physicians already know and do use these tools, bad ones will only get worse. A legal mandate will not benefit care.
Doubtless these models will progress to where this calculus will change. The only benefit from a mandate now that I can foresee is to accelerate fine-tuning by forcing widespread reinforcement learning by physicians, but that is a different discussion.
> All liability stays at the doctor who makes the final diagnosis
Sorry, what? You’d force clinicians to use a specific technology (a specific “how” for finding their answer) and also make them liable for the correctness of that answer?
You seem to have a strange idea about the law if in the same breath as you make something illegal, you sigh with exasperation at lawyers and liability.
I did not say the doctor would need to trust the answer. Only that they should be required to ask.
What if the case is so straightforward (you've see thousands of these, hundreds a year, the entire system is built around them) that you know the diagnosis in less than the blink of an eye?
What if it's emergent, and you have no time to think, like a major hemorrhage? Not only is it obvious, but you must act now, right now?
What if there is a highly studied, routinized process (e.g. cardiac arrest) where you're managing a team going through the diagnostic procedure and treatment, which, over decades, have become a carefully interleaved dance performed at stacatto pace, and, again, there is no time to consult an LLM?
What if? What if? What if?
Are you so certain?
That seems to cover all your cases?
It doesn't. LLMs are trained on literature. Women and people of color are severely underrepresented in medical literature.
E.g., patients of color are rarely selected for clinical trials etc.
Maybe we can legislate this into existence as well?
Probably most of the time when they ignore it, they'd be right. After all, neither the NN nor themselves are right 100% of the time.
But is that other case a lawsuit risk? And we've developed into a very risk averse society.
I’m a data scientist, and it stuns me the way in which diagnoses are made compared to how they could be made if we had a large worldwide dataset of symptoms and other observations to draw correlations from. Especially with regard to preventative medicine.
Not only that, but there are a lot of bad doctors out there. If you go to four different doctors for an even slightly obscure problem, there’s a good chance you will get four different diagnoses. If we applied rigorous statistical tests to the assessments made in the medical industry, I think everyone would be unsettled at how inconsistent and irreproducible everything is (as applied to the medical practice—not necessarily academic medical research).
I'm concerned that misusing LLMs would allow even shorter consultations, and that would become enshrined in reimbursement rates.
Because the parent commenter "spent some time" with ChatGPT and doctors, we should change our entire paradigm of modern medical care carefully refined and honed over 500 years.
Yeah bro, you know better than all of modern medical science because you played around with chat gpt for an afternoon.
Advertising your medical practice as an “LLM-consulting” (better name needed) one in the same way there’s “Montessori” schools could be interesting.
Another option I think that would be interesting would be making the LLM patient-facing and required as one of the check-in “docs”. And then attach the results to the patient’s file for the medical professional to view.
Could also be great for pre-screening and/or suggesting a virtual visit if appropriate.
Generalists tend to lack the specialized knowledge to diagnose certain things even when the evidence would be clear to a specialist.
NN based approaches have the potential to bridge a real gap here.
We've sure come a long way from "a computer cannot be held accountable, therefore a computer must never make a management decision".
The evaluation was done on 246 X-rays, which is good; it would be better if they were not all chest X-rays (to see generality), and if there were more than only four radiologists from the same country.
The achieved 0.25 clinically-significant errors is impressive, although it is only for the best of 3 models, which can incur bias; averaged across models, it is 0.27, a bit worse than human error. Additionally, I am wondering where they get the human baseline; they state:
> These results are on par with human baselines from prior work [14]
but the citation[1] doesn’t give data in the same format (and its format is honestly better: it indicates that humans make no urgent errors or worse in 64% of reports).
Surprisingly, there is no improvement with model size; the largest model performs the worst.
Basically no one could figure out what was wrong with the patient despite 2 years of seeing GP + Specialist + strong drugs. In about 5 minutes GPT3 gave my wife 10 possible diagnosis, she looked up the ones she wasnt familiar with, found one she thought matched, and later did the confirmation test: Yes, he had it. No more medicine needed...(Surgery was needed).
The insane part to me was that I later changed the prompt to say: This is the most likely diagnosis, and it got it correct.
The update on this story, we had another patient with the same issue come in, wife knew the symptoms at this point and knocked out another diagnosis.
The GP should be diagnosis this, my wife is in a specialty that doesnt deal with this. This means multiple physicians are missing this relatively common diagnosis. I hope that medical software takes note pages and places possible diagnosis at the top of the page automatically. Given how many of my tech friends are resistant to AI, I imagine healthcare field is way worse.
For example, a doctor googling your symptoms doesn’t require a BAA with google.
What happened to "you just train it on more data and performance goes up exponentially"? There were so many charts proving this and we could even project at what level of data/compute world-conquering superintelligence would inevitably emerge.
Come on. This is Google. They have unlimited compute resource, biomedical AI is a core strategic objective, they've spent untold $$$ getting data and working on this for years. There are news articles on their efforts going back a decade. And all they could do is wave their hands at "maybe not enough training".
This type of PR research is what is really holding back AI for medical images.
I guess that at least in France, most people living elsewhere than large towns do not need that.
What they need (IMO) is competent radiologists in a time of medical deserts.
Because even when there is staff, this staff do not spent enough time analyzing images while being quite expensive. Their main goal is to make money quickly.
It could be a kind of force-multiplier for the few health professionals they have on hand.
They may still lose out in time, but not without a battle over explainability.
looking forward on updates
Medical practice is very knowledge intensive, especially general practice. Very prone to human error, have you seen the stats on accidental deaths in hospitals? Excellent high-value use for (future) learning systems which can ingest and apply a large body of knowledge and reason in the face of a huge poorly understood graph of causal factors.
No, the last people to be replaced will be those doing unpaid labour, such as parenting, household chores, unpaid elder care: there's the least economic incentive to do so. So I do agree in large part with your final sentence.
I choose my GP based on my ability to trust and empathize with him, not because of his grades in medical school.
Medicine is an intrinsically human activity. It may be augmented and improved by AI, but people are still going to want human medics around, it’s just the nature of being sick.
I agree the direct interface to a lot of medicine needs to be a human as we need human assurance when we are sick and scared. Bed side manner can’t be replace to devalued. But we suffer from critical shortages of medical workers, many of whom can be augmented powerfully with knowledge machines.
My understanding is the real barrier is providers themselves find the user interfaces counter intuitive, aren’t afforded time to train on new systems, and are under so much pressure to produce with such a limited staff they can’t spend the time and energy to work these tools into their workflows. Add on to that the regulatory hurdles to innovation, general risk aversion in the development field, complexities of selling unproven tech into unsophisticated hospital network administration, etc, it’s no wonder it’s a slow slog. It’ll take someone like Kaiser really committing material capital, time, resources, and mandating adoption with a significant training program to see any success.
That's a matter of personal preference. I prefer self-service checkout in the grocery store even if there's an unused staffed checkout. Generally, I prefer robots and AIs wherever they are available.
Empathy isn't going to help diagnose.
Regulatory arbitrage.
You'd think Pilot would be the last type of 'vehicle driver' that would be automated. On the contrary, they were the first to adopt almost 100% self-driving.
> mental
We are finding curious things like 'playing tetris is the most effective way to avoid trauma'. Doctors will help in lab and diagnostic work, but the inference of what solution to administer might actually come from an AI. Ofc, similar to flight, the most challenging situations will still be 'manned' (like fighter jets or non-routine surgery)
Healthcare is inaccessible and understaffed in most parts of the world. If AI doctors can be better than the bottom-10-percentile doctor coming out of a US med-school, then they should be adopted wholesale.
So far WebMD with its entire algorithm being if(symptom) return "cancer"; has somewhat destroyed all public faith in these sorts of systems, but that doesn't mean it's impossible to get done to a useful degree.
That's pretty much all we can expect from Google being involved in health care. Their entire business is built upon not caring about human lives.
> Generalist biomedical artificial intelligence (AI) systems that flexibly encode, integrate, and interpret this data at scale can potentially enable impactful applications ranging from scientific discovery to care delivery.
The problem with AI models is it may bring us medical innovations that are actually harmful to humanity, such as increased lifespans, which in turn will drive further overpopulation.
If we actually reach a point of technological singularity, it may well be the case that we are living amongst the stars within 10 years. Nobody knows. I wouldn't worry too much.
No, new developments in transportation are unlikely to solve anything, except give some people more space in the short-term, and cause even more habitat destruction. New efficiencies in agriculture will help in the short term, but then make producing new humans even more efficient, causing even more destruction.
As for your predictions that we will be living in the stars, that's depressing. On spaceships and barren planets? Think terraforming will help? We can't even take care of one of the easiest planets on the solar system, earth.
And what you managed to come up with is... people will live longer! That's just hilarious :-D