Meditron: A suite of open-source medical Large Language Models
github.com
github.com
I’ve run into people on this very site who use LLMs as a doctor, asking it medical questions and following its advice.
The same LLMs that hallucinate court cases when asked about law.
The same LLMs that can’t perform basic arithmetic in a reliable fashion.
The same LLMs that can’t process internally consistent logic.
People are following the medical “advice” that comes out of these things. It will lead to deaths, no questions asked.
I tried unsuccessfully to search for an ECG analysis term (EAR or EA Run) using Google, DDG, etc. There was no magic set of quoting, search terms, etc. that could explain what those terms were. Ear is just too common for a word.
ChatGPT however was able to take the context of the question I had (an ECG analysis) and lead me to the answer right away of what EAR meant.
I wasn’t seeking medical advice though, just a better search engine with context. So there are clearly benefits here too.
That said if its correct, LLMs are less work for this kind of thing.
But often when people explain to another person what they are looking for, as done in their comment, they will do a better job explaining what they need than when they are in their own head trying to google it. Which is why I just snipped their words to search for it.
Using ChatGPT as a starting point? Sounds really good to me, been there, done that.
You can always check information before believing or acting on it.
However it’s often super difficult to even get started and know what it is that you should be reading more about.
I appreciate the advisory notice in the README and the recommendation against using this in settings that may impact people. I sincerely hope that it's used ethically and responsibly.
I don't think people should trust LLMs completely, but let's be real, they shouldn't trust humans completely either.
AI doesn't necessarily make that risk higher or lower a priori.
Plus if you knew how much of current medical practice exists without evidence you wouldn't be worrying about AI.
"Six patients 65 years or older (2 women and 4 men) were included in the analysis. The accuracy of the primary diagnoses made by GPT-4, clinicians, and Isabel DDx Companion was 4 of 6 patients (66.7%), 2 of 6 patients (33.3%), and 0 patients, respectively. If including differential diagnoses, the accuracy was 5 of 6 (83.3%) for GPT-4, 3 of 6 (50.0%) for clinicians, and 2 of 6 (33.3%) for Isabel DDx Companion"
Yes a patient could be at risk - they're at risk from everything, including a poorly trained/outdated doctor. And even more at risk from just not having access to a doctor. That's the point: it's a risk on both sides; weighing competing risks is not whataboutism.
If someone is going to make that mistake, there were other mistakes happening, not just using a LLM.
On the positive note. LLMs offer the chance to potentially do much better diagnosing on hard to figure out cases.
If this can help with that, I am all for it.
But they could be presented 99% probability for flu, 1% or wazalla, and that testing for wazalla means pinching your ear tout may actually be correctly diagnosed sometimes.
It is not that MDs are incompetent, it is just that when wazalla was briefly mentioned during their studies, they happened to be in the toilets and missed it. Flu was mentioned 76 times because it is common.
Disclaimer: I know medicine from "House, MD" but also witnessed a miraculous diagnosis on my father just because his MD happened to read an obscure article
(for the story, he was diagnosed with a worm-induced illness that happened one or twice a year in France in the 80's. The worm was from a beach in Brazil, and my dad never travelled to Americas. He was kindly asked to provide a sample of blood to help research in France, which he did. Finally the drug to heal him was available in one pharmacy in Paris and in Lyon. We expected a hefty cost (though it is all covered in France), it costed 5 franks or so. But we were told with my brother to keep an eye on him as he may become delusional and try to jump through the window. The poor man cold hardly blink before we were on him:) Ah, and the pills were 2cm wide, looked like they were for an elephant. And he had 5 or so to swallow)
I'd take my chances with a "properly trained" AI any day. Problem is, most medical corpus is full of bogus studies that have never been replicated, so it might be close to junk at this stage.
> It will lead to deaths,
regular doctors kill people everyday and get away with it because you accept the risks. What's different?
Seems like... there are lots of opportunities these days to clear up what open source means?
Results: @70B: Better than GPT3.5, better than non-fine tuned Llama, worse than GPT-4.
70B gets a human passing score on MedQA. (Passing: 60, Medtron: 64.4, GPT-3.5: 47, GPT-4: 78.6).
TLDR: Interesting, not crazy revolutionary, almost certainly needs more training, stick with GPT-4 for your free unlicensed dangerous AI doctor needs