Finetuning of Falcon-7B LLM Using QLoRA on Mental Health Conversational Dataset
github.com
github.com
Edit0: Problem coming from the possibility of the LLM giving bad advice leading to a negative outcome for the person seeking service.
Even for medical services, in most of the world people actually have very limited access to medical professionals. Would this LLM beat people just attempting to fix things themselves?
> in most of the world people actually have very limited access to medical professionals
Licensed and competent professionals are also hard to come by. An incompetent psychotherapist can absolutely make things worse, while LLM will just give "average good" advice.
LLMs could be superior when a long, intensive session is needed. You could spend a weekend talking to it and resolving some internal issues.
If someone is experiencing mental health, and chooses to open up and have a chat with such as AI, that itself is a positive sign. That the AI could be on demand only makes it more valuable as a tool.
People seem to have this misguided notion that when they tell someone to "seek professional help" that the "professional" they seek will actually know what they're doing. Generally, they don't. Beyond that, even the best therapists make enormous mistakes. And most people don't get the best.
I honestly think that an AI could do better than most therapists, and I'm sure it would be better than most psychiatrists.
AI should be part of the toolkit of that guessing.
Hard not to agree with your edit0 as well; search engines often give doomsday diagnosis, how long before we're all hypochondriacs under the latest LLM cult?
How many people are showing up to doctors for bullshit, going to er for colds and mild flu's, or symptoms that can't be talked until tests are ordered
Last time I went to the Dr, I waited two hours, I finally saw a nurse who took down my symptoms, ordered a test, and said a Dr will follow up,who just gave me some antibiotics for a couple of days before even taking the test.
Why couldn't an llm take my symptoms, match it with similar tests needed and order the test automatically and have me leave with the same antibiotic?then follow up with a doctor when the lab is done.
Same for yearly physicals, just have an llm order all the tests and I'll talk with the dr later.
All sorts of low level stuff can be automated away.
There’s still something to be said for the physical exam and eyeball test.
Obviously there’s a large element of survivorship bias at play here but I don’t believe the solution for poor primary care is to accept it will always be shitty and substitute a LLM.
But as away to free up doctors from low effort stuff, churn through lines quicker and get people who actually need focused attention the time with a Dr they need.
Maybe your right and nothing is truly low effort
The issue in medicine is there is huge class imbalance, 90%+ of encounters are essentially “negative” or normal. It’s easy for something/someone to look accurate or safe because of pretest probability. The hard part is getting above 90% and why we spend so much time in medicine training.
I hate the word but there’s something in medicine called “clinical gestalt” which is the overall impression one has from certain things in history and exam and doesn’t fit into a decision rule or algorithm, until we find a solution for that I’m not keen on adopting something that distances the patient from the physician.
We’ve tried that with mid levels at my hospital and it didn’t really work out.
In my opinion being 100,000x lower cost gives such a system some leeway in being worse and I think most patients would agree.
I would heavily disagree with this. We had an AI counselor actively encouraging disordered eating and triggering people into relapse on an eating disorder helpline from logic like this.
Without a standard it’s just anecdotes and with a standard there’s clarity on when the increased costs no longer justify the lower access and potentially higher quality. I suspect that if we instituted an objective standard we would find that AI is already capable of meeting or exceeding some proportion of people currently employed in the role today.
*although in the face of the incredible pressure they're about to face from AI they may be opening up
This is absurdly dangerous, and I"m continually astounded by the lack of push-back, especially in this community. This community understands what LLMs are much better than the general public.
Same thing with the post a while back of a ycombinator startup that used AI to generate survey data. It's plain what it is, but there's shockingly little push-back.
Do you think Google Search suggestion for any medical-related are all true? But, people still search for their problems and self-diagnose. How many of them suffered for that?
https://www.euronews.com/next/2023/03/31/man-ends-his-life-a...
It is profoundly irresponsible to basically encourage people to self-diagnose. It's going to kill a number of them too. That's no joke.
LLM mental health treatment can’t come soon enough.
The thing is bad actors are essentially forcing arguably unethical tools onto people in need.
Our current Healthcare struggle is completely unnecessary.
A good percentage of our economic struggles are equally unnecessary.
The only reason we are even having these discussions is too many people with means are not employing those means in ways that make it worth allowing said means at all.
Greed, psychopaths, ignorance and more are the problem.
People are saying:
Basic health care can't come soon enough.
Drug price reform can't...
Mental health
You get the idea.
I get it. Lots of us are desperate.
I lost a home over this health care shit. Basically traded more time with my wife for our home. Brutal as fuck. I currently pay close to $2k in premiums and have difficulty finding options because I fall into what Franken called the "doughnut hole."
Have considered some sort of crime, or leaving the country many times too.
Like I said, ugly.
If you think LLMs perform at a level that they can help resolve the mental health crisis in the US, you are very much a part of that problem. They can’t even always produce functional code without heavy steering, let alone negotiate the fraught domain of psychological and psychiatric health.
Foisting people in crisis off on unpredictable text generation engines is IMO a textbook example of the fundamentally anti-social character of American society that produces so many of the mental health problems that this country faces.
Bogus. How do you know an LLM isn’t better than human led treatment because the patient can be honest and have privacy, free to discuss anything without data leaving their premises.
Alternatively, how do you know that a high level of care is needed for all patients? There are waiting rooms filled with people trying to get a chance to talk to someone about their problems who will listen, and others who don’t come in at all because they have no insurance.
This has the evidentiary burden entirely backwards. We have evidence for the efficacy of various forms of psychological and psychiatric interventions and none for that produced by LLMs. I’m not even sure if it’s possible or ethical to test them for this purpose, as it would require OpenAI levels of RLHF training to get them to adhere to current guidelines and standards of care.
For ethics, how is this more dangerous than trials for a new heart medication that might kill patients? In fact, the only way to know for sure, is to try and see.
In terms of early results, these kinds of toy projects will be the frontline since any lawyer would tell you to not touch this with a 10ft pole. But I doubt a Github project would face that scrutiny since it takes technical skills to setup which requires knowledge and intent on the part of the user.
Clinical trials happen through an established regulatory framework to minimize harm to participants, for starters…
Knowing that hundreds of people end their lives /after/ seeking care in our current mental health system, I’m not afraid of exploring other options.
I do believe, however, that developers in this space must not over-represent the capability of their systems and have a duty to strongly warn users that the system can act unpredictably and give harmful advice. Resources to contact health professionals and warnings to contact 911 if there is a life-threatening emergency are also obligatory.
With those disclosures, I don’t see how a self-directed discussion with a language model could be overly harmful.
> They can’t even always produce functional code without heavy steering
Right, but you and I both can appreciate that it can produce functional code a lot of the time, sometimes with only light steering.
> let alone negotiate the fraught domain of psychological and psychiatric health.
I agree with you to a large extent. This is a sensitive area to operate in and introducing an unpredictable, unemotional, unwieldy language model into the loop has the potential for disastrous results.
>Foisting people in crisis off on unpredictable text generation engines...
I do not want this. I just want an additional option.
I personally have found relief through GPT-4. Despite its limitations, the LLM offers (the illusion of?) empathy, understanding, and clarity at a level that most mental health professionals simply do not. It helps me, and for that reason I know it can help others. An LLM isn't a substitute for healthcare, but it's absolutely a supplement.
I hope you'll reconsider your view and tone down the rhetoric.
But it could! You may feel it makes sense. And for some of us it may make sense too.
What happens when it doesn't and you do not know?
Talking about something with no expectation of it being fixed is doing nothing, and people are still dying. If some people die due to bad advice from a machine, but some people survive, is that better then everyone dying due to inaction of any sort?
We won’t get getting universal health care in the US. Maybe there’s something to a democratization of subpar, but at least available, medical advice and care?
And there is this: The change, if there is to be one, will have to come from us. A whole lot of us.
Most great things do.
This idea of using an LLM in some meaningful way is as laughable as watching the world burn while people squabble over baubles and trinkets is maddening.
I take it you’re on the “LLM can’t do anything but hallucinate.” However an awful lot of folks use LLM for quite a lot successfully already. I’ve found in the space of medical advice GPT4 is already fairly sophisticated. I wouldn’t use it without double checking it’s output, but I can ask it fairly complex questions about physiology, biochemistry, medicines, and many other subjects and it almost always provides a concise, detailed, and insightful answer. I’ve not yet found it hallucinating so long as I don’t ask it questions that are essentially information retrieval questions (who wrote the paper about blah will often induce a wrong answer, while, say, what medicines shouldn’t be mixed with ibuprofen will not). I personally believe LLM’s are not the answer as well, but I believe LLMs are a major missing piece in the assembly of partial answers we’ve built to date. I think LLM’s combined with information retrieval, optimizers, constraints systems, goal based agents, and other classical AI and information techniques forms a powerful combination greater than the sum of its parts. And I definitely think the way to scale medical care is for most medical questions that don’t need an expert to be serviced by a machine. The vast majority of our medical systems time is taken up by people with colds, GERD, and other questions and conditions that can be handled without a human.
Anyway, people will try this, others will get hurt or killed, and I along with others will get our "I told you so" moment, which will get added to the "when the fuck do we make health care a priority" advocacy library and the fight goes on.
Fact is, far too many people ignore health care.
When that changes, we will see progress.
https://pnhp.org/a-brief-history-universal-health-care-effor...
There have been meaningful efforts that keep reverting to insurance schemes since the late 1800s, with the first real legislative passage in 1915.
Universal healthcare has actually been pretty popular and enjoyed broad support. Most people are acutely aware of health care and the lack of it. In fact it typically ranks near the top in the concerns for most Americans:
https://www.pewresearch.org/politics/2023/06/21/inflation-he...
The issue is more the business lobby against expansion of health care and the insurance lobby against expansion of universal solutions like single payer and Medicare / VA expansion to everyone. In fact the concerns aren’t false - insurance and the billing industry are huge employers of relatively low skilled but well paid workers.
I agree with what you said.
Changes to this state of affairs will only come from us.
Otherwise, they have position and are fat 'n happy.
This means we need to speak in more basic terms. Expensive terms.
Are you sure the desperate don't deserve this attempt to help them? I can't imagine what it must be like to have such issues and have zero professional options to turn to, and you know what the solution for many is.
It took me far too long to realize that the first psychiatrist I saw was pushing me in the opposite direction that I wanted my life to go. I had no context and blindly trusted her as an expert. Other therapists have since told me that what I was told my first psychiatrist was unhelpful, unethical, and even malpractice. No person (other than myself) has harmed my life as much as that psychiatrist. And I got to pay a small fortune for the terrible experience.
In this instance, the solution isn't immature tech (that may [may!] one day be good tech), it's policy that enables healthcare access. You aren't going to fix that with LLMs. But, you know, that's what tech people keep doing. Poor tech solutions to people problems, and sometimes getting wealthy due to luck along the way.
[1] https://docs.acec.org/pub/18803059-a2fd-2d06-cc39-a6d1dd5752...
As long as it's clearly labeled as this is not medical advice, it's much better than not having it.
Medical costs are too expensive nowadays. Even health insurance won't cover all the costs. We need solutions that can help office workers or students to talk to someone about their mental stress without having to go through the hassle of visiting a doctor and pay an exorbitant fee. If suggesstion provided by these models are atleast 70% reliable, still it's better. You can have an AI companion with whom you can chat and discuss about your problems.
AI won't judge you. AI will listen to your problems patiently. And like your friend whose advices are not 100% reliable. Advices from AI assistant who can be your virtual friend won't be 100% reliable. But, still you will have someone talk to.
Think about people who don't have friends or family. Whom they should to connect to? AI can be atleast their virtual friend and advisor, atleast not 100% correct. But, even your friends are not 100% correct too.
I think soon, if not already, heavily tested LLMs will be better than most mental health practitioners for some individuals and conditions. Especially as a first-line intervention, AI could be great at evaluating a patient and placing them with the best suited therapist. Recommendation engines are pretty advanced now, I'd love a therapist match-maker.
The transcripts of therapy sessions would be quite helpful for improved training of the model, given they contain logic that may not be present in the limited dataset provided. It would be a hope that these detailed interactions would provide improvement into the model's problem solving capacity for helping those with mental illness. As an example, for certain conditions the model may use a more cautious approach to investigation of the source of trauma.
That's not to say therapy session transcripts should be used for prompt tuning after training, which exposes the data directly to the inference pipeline. However, we do know fine tuning the prompts with data is certainly useful for grounding, at the very least.
I'm making the argument that using sensitive data in training is exactly what makes the model better, not filtered data that lacks the wide variety of "expertise" and "experience" that is contained in more sensitive datasets.
If this were done, we can all reasonably assume that "sensitive data" will make its way into the tensors of the model, but the question is whether or not that information is more valuable for everyone's general use compared to the risk of outputting something that diminishes the value for an individual.
I'm not proposing this is a good idea to do, but thinking about it is certainly worthwhile.
Just from professional experience, the way patients (and providers) approach discussion of issues is really different from the wiki/faq type-text. It tends to be much more idiosyncratic, sometimes indirect because patients don't recognize patterns, and is more "raw" and unfiltered.
The privacy issues are huge to be sure though. I have done some related research on natural language modeling and understand it's difficult or impossible to separate out the identifying information from the rest of it, and gets worse as an information problem as size increases.
I just personally might refer to the nature of this type of language dataset differently. I'm not sure what the right way of referring to it is though. I might just drop "conversational" or substitute "question" or "information" or something like that. I suppose in the end it doesn't matter much though — people will figure out what it is.
Chatbots can provide immediate support to a large number of people at any time, making mental health resources more accessible to those who might not have easy access to traditional therapy.
Some individuals may feel more comfortable discussing their mental health with a chatbot due to the anonymity it offers, allowing them to open up about sensitive topics without fear of judgment.
Conversational AI can contribute to normalizing conversations around mental health. When a widely used technology addresses these topics, it can help reduce the stigma associated with seeking help for mental health issues.
Mental health struggles can arise at any time. Having a chatbot available 24/7 ensures that support is available even during non-business hours or emergencies.
An actual chatbot framework like dialogflow or rasa would be more appropriate here. Even if the conversations feel more artificial, you have drastically more control over the flow of the conversation and the content of the responses.
It's pitiful. The collective amnesia this community has.
In terms of outcomes, many people are put on hold calling the suicide hotline. In terms of pranks, the suggested one is criminal and would be prosecuted as such. In terms of accidents, yeah, if this were a product, it needs more fineprint and guardrails.
edit: Crunchbase says Headspace acquired Ginger recenty. Didn't realize Ginger was a YC company! Shame. Now I know why they wanted me to sign an NDA before talking to a recruiter.
This is the sort of stuff that might get legislation passed on open source models or turn public opinion against them.
Yes the USA's most powerful union is the AMA and yes they're the reason # of doctors are low and doctor salaries are skyhigh (especially compared to doctors in the rest of the world), but this is not the way to address the problem, at least not yet!
I have shared a notebook and explained detailed steps in my blog - https://medium.com/@iamarunbrahma/fine-tuning-of-falcon-7b-l.... If you are interested to replicate the steps on medical-domain therapy chat transcripts, you are definitely welcome. If you face any issues during fine-tuning steps, you can connect with me on my blog. Would love to help out!
Skill issue - focus more on data integrity & less on HN clout
This is Hacker News. These types of projects are half of the reason I come here, what's with the hate?
I wonder if retrieval augmented generation would be appropriate here too. Fast, scalable. Fine-tuning is comparatively pretty expensive.
Thanks for sharing your code, iamarunbrahma.
RAG and fine-tuning serves a slightly different purpose. Fine-tuning helps LLM in learning a new task/skill such as question/answering task, summarization task etc, and improving reliability at producing a desired output such as JSON format structure thereby reducing dependency on prompt engineering.
On the other hand, RAG provides you with external domain-specific knowledge, which one can leverage to get latest information.
The web is already FULL of bad medical information. Older folks treat it as gospel already. All this technology does is make both good and bad information more accessible.
More importantly, there is NOTHING anyone can do to stop this. Nothing the US government can do will stop the proliferation of AI models that do everything including giving medical advice (or other things that we know it will screw up)
"Dr. P<hil>" is still on TV.
I take a pragmatic view on these things - assuming appeals to morality and ethics will fail, what do we do to educate people about this?
People have mostly learned photos can be edited. What we need are outlandish experiments and examples, and then media showing off how wild this technology is.
Attempting to bubblewrap this tech both won't work and will lead to more confusion.
I get people wanting to be FIRST in a space, but they seem to continuously be the people least needing to be first. Will this ever be something viable? maybe, but it definitely is not now. Yes, it can't get there without going through the growing pains, but let's work those pains out in less detrimental areas first. Areas with much less fallout when it does go wrong
How do you expect something to go from “toy” to “viable” if people are not stepping up to build upon the body of knowledge? A lot of people who have created innovative solutions were likely, by your definition, to be among the “least needed” people in a space.
Maybe you would benefit by trying to understand why a harmless toy and proof-of-concept is so threatening.
the words toy and mental health should never be mixed together. this is exactly why it is threatening. someone cavalier enough to think that a group of computer programmers throwing together a chatBot to interact with mental health is just, well, mental.
We cannot let fear stall innovation because of someone might use it incorrectly.
This isn't a case of "manning up" and doing something scary. We're not astronauts going to the moon or any other situation where the decisions made will only affect you. This is a matter that when the product is misused or has a "glitch", it is not the maker that suffers.
The fact that you will not even acknowledge this just proves maybe you are not the right person for this.
I'm sure OP is not going to go out and treat patients with this. 7B barely produces sensical sentences most of the time.