The transcripts of therapy sessions would be quite helpful for improved training of the model, given they contain logic that may not be present in the limited dataset provided. It would be a hope that these detailed interactions would provide improvement into the model's problem solving capacity for helping those with mental illness. As an example, for certain conditions the model may use a more cautious approach to investigation of the source of trauma.
That's not to say therapy session transcripts should be used for prompt tuning after training, which exposes the data directly to the inference pipeline. However, we do know fine tuning the prompts with data is certainly useful for grounding, at the very least.
I'm making the argument that using sensitive data in training is exactly what makes the model better, not filtered data that lacks the wide variety of "expertise" and "experience" that is contained in more sensitive datasets.
If this were done, we can all reasonably assume that "sensitive data" will make its way into the tensors of the model, but the question is whether or not that information is more valuable for everyone's general use compared to the risk of outputting something that diminishes the value for an individual.
I'm not proposing this is a good idea to do, but thinking about it is certainly worthwhile.
Just from professional experience, the way patients (and providers) approach discussion of issues is really different from the wiki/faq type-text. It tends to be much more idiosyncratic, sometimes indirect because patients don't recognize patterns, and is more "raw" and unfiltered.
The privacy issues are huge to be sure though. I have done some related research on natural language modeling and understand it's difficult or impossible to separate out the identifying information from the rest of it, and gets worse as an information problem as size increases.
I just personally might refer to the nature of this type of language dataset differently. I'm not sure what the right way of referring to it is though. I might just drop "conversational" or substitute "question" or "information" or something like that. I suppose in the end it doesn't matter much though — people will figure out what it is.