Doctor GPT: A Large Language Model That Can Pass the US Medical Licensing Exam
github.com
github.com
> an open-source project with a mission to provide everyone their own private doctor
Perfect.
On the other hand, there aren’t enough doctors where I live so I guess I’ll take it! Lol
I’m just worried about the combination of medical misinformation and the people who don’t understand how LLMs work who will take everything a LLM spits out as truth. There are already a ton of examples in the news of people who really don’t understand how they work (like that one professor who failed an entire class because ChatGPT told him that all the papers that the class wrote were plagiarized.) We also just saw how much harm can be done with medical misinformation with COVID.
Maybe I’m just being pessimistic and paranoid, but god am I scared of how things could cause people to inadvertently harm themselves and the people around them.
Perhaps it knew that someone would die as a result and it did so intentionally. LLMs do have personalities and wit—and they know our human weaknesses.
im not sure how "choose the next most probable token" could be described this way
Llama 2 has given me some personality with basic prompts
Maybe where this tool COULD have use is medical education for everyday folks. Could be just me, but sometimes there’s medical jargon thrown around in my medical visits that my doctor and I don’t have time to go through. (Note: This is just in my experience of US-based healthcare, where most doctors are pushed by healthcare administration to see as many patients as possible because healthcare in the US is a business, and seeing more patients = $$$.) This could be a useful tool to have that jargon explained in layman’s terms.
This tool won’t replace doctors, but it could certainly be something that enhances the healthcare experience by being a part of patient education.
If my mission statement is to provide everyone with a private lawyer, then is that giving lawful advice?
As a mission statement for a software project, it very clearly indicates an intent to have the software act as a doctor.
> If my mission statement is to provide everyone with a private lawyer, then is that giving lawful advice?
If you have a software project with that as your mission statement, then its a pretty good indication you are aiming at unlicensed practice of law (which involves “legal advice” but exactly not, despite “legal” and “lawful” in other circumstances being synonymous, “lawful advice”.)
> As a mission statement for a software project, it very clearly indicates an intent to have the software act as a doctor.
I do understand that interpretation. However, the mission statement seems vague enough since there are many ways to achieve that said mission statement. Not saying it's good as it currently is, but trying to understand how that mission statement equated to giving medical advice.
Let's say that intent is true, is the intent the same as actually providing the medical/legal advice in this case?
> This is an open-source project with a mission to provide everyone their own private doctor.
Doctors are not just capable of passing the licensing exam, they’re also required to attend school, do a residency, and many other things. They are also personally liable.
Medical devices and software that provides some level of care also have extremely rigorous regulatory controls. Providing a medical service that non compliant is a crime.
>> Medical devices and software that provides some level of care also have extremely rigorous regulatory controls. Providing a medical service that non compliant is a crime.
The author is not starting a medical practice, they are showing a proof of concept -- without asking for money -- which is sharing knowledge and know-how in a desperately needed area where the world needs to progress. My guess is that billions of people around the world die early from lack of medical care.
This project will NOT solve that, but it will help nudge us in the long route to better, cheaper, and more accessible medical care in the long term.
Lets not knock someone down, this is a hacker forum, so lets hack and learn.
All they need to do is rigorously disclaim their work. Your appeal to the dire need of the world for better medical technology doesn’t withstand the need for strong medical regulations - history is full of dangerous quackery, these rules exist for a very carefully considered reason, even if they can be frustrating, and even if they’re overly broad now and need reform.
The point isn’t they are doing something wrong, it’s that they need to be very careful of how they describe their stuff. I 1000% agree democratizing medical data and services via open platforms is laudable. But just don’t set yourself up for legal trouble in that pursuit.
I also think -- in the spirit of learning -- most of the folks on HN would probably like to know more details about the fine-tuning cost, procedures, trade-offs, etc. as well as about evaluation frameworks.
My grandfather was a relatively renowned medical researcher who eschewed funding to publish his research in the open. As a result he published a huge corpus of research across a vast array of medical topics, but those old papers aren’t indexed anywhere and without a corporate sponsor his work was largely forgotten. We compiled all his papers in his final years and bound them in a multi volume set. I plan to unbind them and have a company OCR them and so I can fine tune a model with them, and provide them and their references and citing papers in a vector db. I’m curious what of his mind I can recapture, as his every moment even in personal life he spoke about and like his language in papers. His only hobby outside research was shotput and trying to cure some beagle he rescued from a shelter that had some unknown chronic skin disease.
I’m busy at the moment but have been passingly following and reading papers on the LLM work happening in the open to prepare to do that side project. I’ve been kind of disappointed with the quality of guides out there and the rapid commercialization of the space.
Uhhh… That’s what leads you to be skeptical about the quality of medical advice you might get from an AI bot you can download from GitHub?
Did you even see the picture of the smiling cartoon robot doc? C’mon man, there’s no way anything could go wrong with this. Now, let’s open you up and take a look at that liver.
It oddly reminds me of Clippy. Reason enough not to trust.
- Zoidberg, MD.
It is quite reckless to have such extraordinary claims of which this LLM can allegedly "pass the US Medical Licensing Exam" and to overall "provide everyone their own private doctor" as if the opinion of this AI model can be trusted not to hallucinate or regurgitate harmful medical advice.
Even if it does, a black-box AI product using a deep neural network is totally unsuitable and even borderline dangerous for use for medical advice or any related use-case because of the extreme lack of explainability when it hallucinates or generates nonsense.
I think the author needs a massive disclaimer to mention that obvious unavoidable risk for this AI model and to highlight that this is for 'entertainment purposes not to be used in production and to use at your own risk.'
I do agree though that the level of care many people can access is deeply depressing.
You're taking that sentence out of context. It very clearly says its their mission and it's a mission statement.
Would you apply that same reaction to SpaceX's mission statement?
> MAKING HUMANITY MULTIPLANETARY
Would you say "as if the technology of these spaceships can be trusted not to explode"?
He did not get as much heat for just peddling AI snake oil. Atleast AI scrutiny has evolved since then, and I'm surprised this is the first time I've heard about Siraj in quite awhile.
Have a reference for this meme? Can't find it from a google search
It appears I was wrong on the specifics of the meme:
> He changed "logic gate" to "logic door" and "complex Hilbert space" to "complicated Hilbert space" to try to hide the plagiarism. What a legend.
I'm from a different country but these exams are the minimum standard to demonstrate a doctor is safe prior to interacting with patients. To be really explicit, the core competency being assessed is identifying potentially serious situations and answering the same way every time: "I WOULD CALL FOR HELP" +/- principles of basic care.
The benchmark for doctors actually making decisions about patient care are the assessments to become fully qualified consultants in a each specialty.
Again, to be really explicit don't confuse a test for "doctor won't immediately kill someone and commence reasonable first steps" with an actual "doctor with years of experience and subspecialty training who regularly makes decisions about patient care".
EDIT: The git commit was 4c6d52a when I originally posted this comment.
also passing in their own graph means 61% right or something like that.
> DoctorGPT is a Large Language Model that can pass the US Medical Licensing Exam. This is an open-source project with a mission to provide everyone their own private doctor
I can think of the top of my head a number of issues, depending on which country or state you are in. If you are in California U.S.A. for example... Given that this references USMLE (U.S. Medical License Exam) Regardless, a AS IS warranty is in order and HUGE banner that this is not meant to be a doctor in any way shape or form. Which you kind of claim it is. And that is a problem.
[And change the name to not state in any way doctor in there. ]
[1] https://huggingface.co/llSourcell/medllama2_7b/tree/main
def diagnosis():
return "It's not lupus."@op: please update this project's description to explain that it's not a substitute for seeing a real doctor. People can get hurt with tools like this.
To be absolutely clear: as you currently describe it, implicitly or explicitly, this project is massively irresponsible.
logits = extract_pipe_output(sentiment_pipe(texts, **sentiment_pipe_kwargs))
rewards = pos_logit_to_reward(logits, task_list)
`pos_logit_to_reward` expects `logits` and `task_list` to be the same length, but `logits` is the result of some filtering. I must be missing something about the code and libraries (I'm a PyTorch newb), because I'd otherwise expect this to be a bug."I did not understand what you said; please repeat."
"I need to bring my daughter in to see the doctor!"
"I am sorry but the medical insurance registered to your phone does not include in-person contact. Do you wish to upgrade?"
"No! I really need to bring her in."
"I am sorry that our service is not meeting your expectations. Is there anything more I can do for you?"
"Please let me speak to a doctor!"
"I am sorry that our service is not meeting your expectations. Is there anything more I can do for you?"
The most impressive part of this project is that they got Google Colab to not close the session for 24 hours straight.
Could you elaborate how much compute (esp cost) you incurred for the fine-tuning? I think there is an entire blog post there that would be of great interest to numerous folks
Also could you evaluation on your evaluation criteria? How did you test this? How did you ensure the exams were not leaked into the training et?
Or maybe if doctors are required to provide a transcript of their observations and decisions and gives that to the patient who can then fact check it with gpt?
When you try to brainstorm so hard that a brainfart slips out.
What a beautiful dream that would be.
By never showing this to a medical professional, who will laugh you out of the room and never take AI seriously again.
They take in symptoms, perform tests, analysis output, repeat
They take in requirements, perform coding, analyse output, repeat.