Youtube transcripts. It only works with single-person channels at the moment, as I haven't worked on disambiguating multiple speakers. They are very messy if they are just an auto transcription from Google. Practically unusable in most cases.
So first step in the pipeline is cleaning up the transcripts for incorrectly transcribed words or sentences. Using the context of the rest of the transcript, it is able to fix the vast majority of them.
Then we add punctuation and format it with paragraphs.
Then I have another check over the whole transcript for any remaining issues.
After all of that, I have a relatively clean transcript that represents the original audio very closely. From there, I am doing things like:
1. creating question/answer pairs from the transcript
2. creating a document of additional context that fills in details about what the speaker is talking about but may not have explicitly said
3. creating a summary of the transcript identifying the main purpose
4. creating a knowledge graph from the transcript with nodes and edges
5. creating an annotated version of the transcript using that knowledge graph
I plan on putting some of this data into a vector database, and some of it will be used for fine tuning LLaMA2 models on specific tasks (like knowledge graph creation, annotation using a knowledge graph, and writing using a knowledge graph to keep track of events)