Whisper.api: Open-source, self-hosted speech-to-text with fast transcription
github.com
github.com
For anyone confused about the project, it is using whisper.cpp, a C-based runner and translation of the open whisper model from OpenAI. It is built by the team behind GGML and llama.cpp. https://github.com/ggerganov
You can fork this code, run it on your own server, and hit the API. The server itself will use FFmpeg to convert the audio file into the required format and run the C translation of the whisper model against the file.
By doing this you can separate yourself from the requirement of paying the fee that OpenAI charges for their Whisper service and fully own your translations. The models that the author has supplied here are rather small but should run decent on a CPU. If you want to go to larger model sizes you would likely need to change the compilation options and use a server with a GPU.
Similar to this project, my product https://superwhisper.com is using these whisper.cpp models to provide really good Dictation on macOS.
Its runs really fast on the M series chips. Most of this message was dictated using superwhisper.
Congrats to the author of this project. Seems like a useful implementation of the whisper.cpp project.
I wonder if they would accept it upstream in the examples.
If you have Nvidia hardware the ctranslate2 based faster-whisper is very very fast: https://github.com/guillaumekln/faster-whisper
We use it for our Willow Inference Server which has an API that can be used directly like OP project and supports all Whisper models, TTS, etc:
https://github.com/toverainc/willow-inference-server
The benchmarks are pretty incredible (largely thanks to ctranslate2).
Fwiw decent acceleration works on any avx2 compatible chipset. I get realtime speed for everything but the large models with a recent Ryzen system. The apple silicon is good but not as special as folks think!
"This project provides an API with user level access support to transcribe speech to text using a finetuned and processed Whisper ASR model."
Why is this a service at all? Why not just a library? Or a subprocess?
This open source project provides a self-hostable API for speech to text transcription using a finetuned Whisper ASR model. The API allows you to easily convert audio files to text through HTTP requests. Ideal for adding speech recognition capabilities to your applications.
Key features:
- Uses a finetuned Whisper model for accurate speech recognition - Simple HTTP API for audio file transcription - User level access with API keys for managing usage - Self-hostable code for your own speech transcription service - Quantized model optimization for fast and efficient inference - Open source implementation for customization and transparency
How does this compare to what is possible using https://goodsnooze.gumroad.com/l/macwhisper for example?
Thanks!
Whisper – open source speech recognition by OpenAI https://news.ycombinator.com/item?id=34985848
One issue I faced with all the whisper based transcript generators is that there seems to be no good way to make editing/correcting the generated text with word level timestamp. I created a small web based tool[0] for that.
By any chance if anyone is looking to edit transcripts generated using whisper, you'd probably find it useful.
An iPhone app that could do this from the microphone would also be amazing. Google Translate and it's various competitors from Microsoft/Apple are nearly there, but they all stop listening inbetween sentences. Something that just listened constantly, printing translated text onto the screen, would be amazing.
Just visit the swagger create account and then gettoken to grab a token
> curl -X 'POST' 'https://innovatorved-whisper-api.hf.space/api/v1/users/get_t...'
while later saying
> curl -X 'POST' 'http://localhost:8000/api/v1/transcribe/?model=tiny.en.q5'
If you just change the first "curl" to also be localhost:8000, this will be cleared up.
Even if it meant delaying the broadcast for a second while transcribing the accessibility value could be immense.
If it's completely self-hosted why do I need to get a token? Where does the actual model run?