Whisper.cpp v1.4.0
github.com
github.com
make stream
./stream
FYI it seems to work on mobile sometimes but sometimes not at all. You might want to add a warning about when it might work or not work on mobile.
Any plans to make the source available for the client side part?
Thanks
Mobile is an issue, some of the web APIs only work on desktop for now.
Thanks for the feedback
Someone else mentioned “Hello Transcribe”, which is a cool demo due to how real time the transcription is, but it’s effectively unusable for anything practical because it seems to split the audio on 30 second chunks (and not in the contextual way that Whisper normally does it).
Whisperboard is a free, open source iOS app (available on the App Store) that seems fairly user friendly, but there are a few crashes and bugs that the author is hopefully going to work out soon.
Aiko is extremely simple to use, but it only supports the “medium” model, so it isn’t very fast, but the results are good quality. Aiko is really designed to be paired with another app, like the Apple Voice Memos app, where you can share a voice memo directly to Aiko.
does this count?
It supports wav, mp3, m4a, and mp4. All the demuxers were written in JavaScript which was an interesting exercise. I couldn’t figure out a way using native JavaScript audio APIs to load small sections of long audio/video files without loading the entire file.
a macos menu bar app that uses whisper.cpp under the hood.
Uses OpenAI API so really fast
As a Flemish/Dutch speaker, not a single assistant is able to understand me in my native tongue. I can use Siri in English, but then she doesn't understand things like contacts or addresses. Also I cannot dictate text messages or people think I'm a spammer.
https://github.com/ggerganov/whisper.cpp/releases/tag/v1.4.1
I don't know much practically about how hard it would be to take the Whisper PyTorch (1 or 2?) trained models & to make good use of them elsewhere. I expect Whisper.cpp probably better caters to users, is more readily consumable. Doing the same integer quantization that 1.4.0 whisper.cpp does would be another ask: how hard would that be?
Fwiw, Whisper.cpp uses Nvidia's cuBLAS. There does appear to be an AMD rocm port. https://github.com/ROCmSoftwarePlatform/rocBLAS
What do you mean by "elsewhere"?
The whisper models can be used without much effort. The tooling provided by OpenAI allows for them to be used and installed in python with ease. People use them in companies to transcribe internal media datasets. No need for CPP implementation for internal uses.
Main problem is when _on device_ is required and the OpenAI model is not quantised and so use on say a smartphone becomes trickier.
Until then... https://superwhisper.com
Check out on https://twitter.com/superwhisperapp
Particularly you can talk to it with a “international” English accent and it will work. Or French or Spanish or Norwegian… and it will transcribe English.
"Adding --task translate will translate the speech into English."
I tried translating a conference with a German speaker. The transcription was superb, but the translation no so much.