So first step in the pipeline is cleaning up the transcripts for incorrectly transcribed words or sentences. Using the context of the rest of the transcript, it is able to fix the vast majority of them. Then we add punctuation and format it with paragraphs. Then I have another check over the whole transcript for any remaining issues.
After all of that, I have a relatively clean transcript that represents the original audio very closely. From there, I am doing things like: 1. creating question/answer pairs from the transcript 2. creating a document of additional context that fills in details about what the speaker is talking about but may not have explicitly said 3. creating a summary of the transcript identifying the main purpose 4. creating a knowledge graph from the transcript with nodes and edges 5. creating an annotated version of the transcript using that knowledge graph
I plan on putting some of this data into a vector database, and some of it will be used for fine tuning LLaMA2 models on specific tasks (like knowledge graph creation, annotation using a knowledge graph, and writing using a knowledge graph to keep track of events)
My only experience with transcripts is in the context of transcribing short interviews. I used Whisper and it was pretty good. I mostly work with quantitative data, though.
In terms of the disambiguation of speakers, I haven't done it, but I remember blind signal separation discussed in a signal processing seminar I attended. There is also this paper, in case you haven't seen it already: https://enk100.github.io/speaker_separation/
Thanks again!
Also, I haven't tried using Whisper for getting a transcription from the audio. I went the route of downloading the automatically generated transcripts from Youtube for a set of videos. An audio processing pipeline is definitely something I could add later though as an additional input channel for the overall pipeline.
Here is an example: https://gist.github.com/Tostino/f6f19e88e39176452c1a765cb7c2...
Here is the transcript that I created that knowledge graph from, and then annotated with the knowledge graph for training purposes: https://gist.github.com/Tostino/e64524437848fbb3aebe52056df8...
Edit: I am using symbolic IDs intentionally. Reason for that was this paper: https://ai.googleblog.com/2023/07/symbol-tuning-improves-in-...