Self-hosted offline transcription and diarization service with LLM summary
github.com
github.com
https://github.com/bugbakery/transcribee
It's noticeably work-in-progress but it does the job and has a nice UI to edit transcriptions and speakers etc.
It's running on the CPU for me, would be nice to have something that can make use of a 4GB Nvidia GPU, which faster-whisper is actually able to [1]
https://github.com/SYSTRAN/faster-whisper?tab=readme-ov-file...
https://github.com/bugbakery/transcribee/issues/427#issuecom...
https://news.ycombinator.com/item?id=40270219
Not sure what you mean by source/platforms, you might find what you need for your operating system on that discussion 3 weeks ago with links to options for macOS, Linux, and Windows.
https://vb-audio.com/Cable/index.htm
So, if you wanted to transcribe the audio from one of your physical output devices (say you had your meeting software outputting on external headphones), you could set the virtual audio device as the monitoring device on the physical external headphones. Therefore, you end up with a virtual audio input device containing the audio from your meeting software. I also do this to apply a chain of filters to my condenser mic in OBS because it picks up everything.
I wanted to be able to transcribe and diarize in realtime though, which is much harder. Didn't manage to make that happen.
30 years of audio that needs transcribing, summaries, and worksheets made out of them.
There are a few whisper diarization "projects" but i've never been able to get it to work. Whisper does have word-level timestamps, so it should be simple to "plug in" diarization.
I don't need an LLM or whatever this project has, but i will see if it's runnable and if it's any better than what a couple podcasts i listen to use.
edit: see some people mentioning whisperx, which is one of those things that was cool until moving fast broke things:
>As of Oct 11, 2023, there is a known issue regarding slow performance with pyannote/Speaker-Diarization-3.0 in whisperX. It is due to dependency conflicts between faster-whisper and pyannote-audio 3.0.0. Please see this issue for more details and potential workarounds.
which means that what i gain is a ~3x increase in large-v2 speeds but i instantly lose those gains with diarization, unless i track down 8 month old bug workarounds.
I'll stick with the py venv whisper install i've been using for the last 16 months, tyvm
https://github.com/MahmoudAshraf97/whisper-diarization
I remember having the usual python package hell when NeMo was updated somewhere, but it seems to be decently well maintained so give it a go.
*Edit, I remember reading somewhere that pyannote was a weak link in other repos, that might be why your other tests were not great.
It's what we were doing at our company until Anthropic and others released larger context window LLMs. We do the TTS locally (whisperX) and the summarization via API. Though we've tried with local LLMs, too.
mistral which clocks at 32k context
I may be wrong, but my understanding was/is:- Mistral can handle 32k context, but only using sliding window attention. So it can't really process all 32k tokens at once.
- Mixtral (note the 'x') 8x7B can handle 32k context without resorting to sliding window attention.
I wonder whether Mistral would do a better job summarizing a long (32k token) doc all at once, or using recursive summarization.
Maybe a neat eval to try.
You can short circuit that time to build up a node a bit with a prebaked AMI on AWS, but there's still some amount of time before a new node can start running at speed, around 10 minutes in my experience.
I haven't looked at this particular solution yet, but I really find the LLMs to be hit or miss at summarizing transcripts. Sometimes it's impressive, sometimes it's literally "informal conversation between multiple people about various topics"
They give $200 of credit.