OTranscribe: A free and open tool for transcribing audio interviews
otranscribe.com
otranscribe.com
Worked excellent.
It generates both a file that just contains a line per uninterrupted speaker speech prefixed with the speaker number, as well as a file with timestamps which I believe would be used as subtitles.
I’ll be honest, I haven’t dived much into this as I just needed something transcribed quickly, but when I was looking at WhisperX I couldn’t find a CLI that would just out of the box give me a text file with a line per speaker statement (not per word).
whisperx $file int8 --min_speakers 3 --max_speakers 3 --language de --hf_token $token --diarize
It seems like it does:
https://github.com/MahmoudAshraf97/whisper-diarization/blob/...
disclaimer : I am not affiliated in any way, just a happy customer! I had some nice mail exchanges after bug reports with the (I believe solo-)developer behind these tools.
---
Thomas here, maker of Spectropic and Audiogest. I am indeed focused on building a simple and reliable Whisper + diarization API. Also working on providing fine-tuned versions of Whisper of non-English languages through the API.
Feel free to reach out to me if anyone is interested in this!
I am also experimenting with post-processing transcripts with LLMs to infer speaker names from a transcript. It works pretty decent already but it's still a bit expensive. I have this feature available under the 'enhanced' model if you want to check it out: https://docs.spectropic.ai/models/transcribe/enhanced
You can do this locally if you have enough (V)RAM, but I prefer the OpenAI API, as usually I don't have enough at hand. And the various Llamas aren't really quality on par with GPT-4. If you only need Whisper, and no translation, then local execution is indeed very viable. High quality Whisper fits in 4GB of (V)RAM.
vs
https://github.com/ggerganov/whisper.cpp
They are two inference engines for running the whisper ASR model, each with their own API AFAIK.
- transcription
- machine translation
- OCR
- image recognition
So no AI here, folks.
- Transcribe word-by-word in real time as audio is recorded
- Work entirely locally
- Use relatively recent open-source local models?
I've been using otter.ai for real-time meeting transcriptions - letting me multitask and instantly catch up if I'm asked a question by skimming the most recent few seconds worth of the transcript - but it's far from perfect and occasionally their real-time service has significant transcription delays, not to mention it requires internet connectivity.
Most of the Whisper-based apps out there, though, as well as (when I last checked) the whisper.cpp demo code, require an entire recording to be ingested at once. There are others that rely on e.g. Apple's dictation frameworks, which is a bit dated in capability at the moment.
Anything folks are using out there?
I'd be in favor of some startup pulling an Uber or AirBnB and blatantly violating those laws to the benefit of the deaf or elderly if it meant we could get something better on the books.
Google Pixel phones have this feature and it works _very_ well.
New Microsoft Surfaces have this feature but just works for English
https://support.google.com/accessibility/android/answer/9350...
This was for English. One problem it took me a while to realize: when I switched it to transcribe a secondary language, it was not doing it on-device anymore. You can tell the difference by setting airplane mode.
If you elect to, the audio + transcription can be access/searched via https://recorder.google.com
Not sure if that helps for your specific usecase.
It works well enough that I use it with Telegram to shove messages over to my Linux machine when I don't feel like typing them out, which is such an unsophisticated hack, but is getting the job done. I spent a couple hours trying to find a Linux-native alternative, or even get this running in Waydroid, and couldn't find anything that worked as well, so I decided not to let the "smooth" become the enemy of the "good enough to get the job done."
Language models were those from BSC (Barcelona Supercomputing Center) at the time. The transcription is done via WASM, using Vosk [2] as base.
I hope it fits.
[0] https://github.com/projecte-aina/oTranscribe-plus [1] https://otranscribe.bsc.es/ [2] https://github.com/alphacep/vosk-api
I wish they had Speaker Diarization, but they are waiting for upstream Whisper to add it: https://github.com/argmaxinc/WhisperKit/issues/31
You do still need to proof and QA even AI results, if you want a publication quality result, and do things like attribute who is speaking when (at least Whisper can't do that), and correct "unusual" last names and things. So I feel like people using AI still need good tools for the correcting/finishing/proofing too, that would be similar to the tools for non-assisted transcription.
It is now operated by Muckrock and hasn't seen changes made to it in a while.
That's why it doesn't have any of these integrations, the technology just didn't exist.
Does oTranscribe automatically convert audio into text?
Sorry! It doesn’t. oTranscribe makes the manual task of transcribing
audio a lot less painful. But you still have to do the transcription.https://www.gally.net/temp/20240809geminitranscription/index...
Aside from some minor punctuation and capitalization issues, Gemini’s transcription looks nearly perfect to me. There were only one or two words that I think it misheard. If I had transcribed the audio myself, I would have made more mistakes than that.
One passage struck me in particular:
And then he comes up with "weird," which becomes viral and the rest, and here he is.
How did Gemini know to put “weird” in quotation marks, to indicate—correctly—that the speaker was referring to Walz’s use of the word as a word? According to Politico, Walz first used the word in that context in the media on July 23.https://www.politico.com/news/2024/07/26/trump-vance-weird-0...
- auditory cues
- the sentence would be gramatically incorrect and make no sense without them
Just guessing out of the blue.
But I think it's likely that LLMs (and other speech recognition systems) need to exploit sentence context to recognize individual words and punctuation, and this is an example were it went well.
Human listening is similar in a way, we can recognize words even when spoken very mumbly or fast, if we have context.
So we always hear phrased rather than words.
Do you have the audio or video file?
I'd like to run it through our AI video editor and see how it punctuates the transcript.
https://www.gally.net/temp/20240809geminitranscription/inter...
The full interview including video is on the New York Times website, though you might need a subscription to view it:
https://www.nytimes.com/2024/08/09/opinion/ezra-klein-podcas...
The NYT’s closed captions do not put “weird” in quotation marks; they also divide sentences weirdly and have other mistakes. But they get some things better than Gemini, such as capitalizing “House” when it means the House of Representatives.
I haven’t compared the audio-only podcast version and the video version carefully; it’s possible that parts of the audio were edited or re-recorded for one or the other.
Let us know how your AI video editor does!
We have a grammar correction step which "fixed" the sentence to "...he comes up with something weird, which becomes viral...".
https://video2srt.ccextractor.org/
Disclaimer: Working on this project.
I'm always surprised at the amazed look of my friends when they see me concretely use the tool. They just didn't picture it until they saw it in action.
It is functional but a bit slow. I think using whisper directly instead of swift bindings will help a lot.
Really interested in adding diarisation but having a lot of trouble converting Pyannote to CoreML. Pyannote runs so slowly with torch on CPU. Haven’t gotten around putting my latest work for that on GitHub yet.
Happy to accept contributions —
Some priorities right now:
* Fixing signing for local builds
* Replace swift whisper with whisper cpp
* Allowing users to provide their own models
Current features: 1. Download from YT 2. Transcribe using Vosk (output has time codes included) 3. Speaker diarization using pyannote - this isn't perfect and needs a bit more ironing out.
What needs to be done: 4. Store the transcription in a search engine (can include vectors) 5. Implement a webapp
If anyone here is interested to join forces, let me know.
> A free web app to take the pain out of transcribing recorded interviews
How did you use a web app on the plane with no internet?
You don't need the internet to use a web browser
Not developing it actively after I created tables of contents for the several videos I needed, years ago. If I ever need it again, I will probably work on mobile UI (aka responsive)
Nowadays, I use libretranslate/libretranslate and pluja/whishper to do this, but not at real time.
Make sure they are recent tutorials because they will probably mention how to use the automated generation tools/plugins that wasn't available years ago.