OpenAI releases Whisper v3, new generation open source ASR model
github.com
github.com
https://github.com/openai/whisper/blob/main/language-breakdo...
Having extensively tested Whisper v2 large against other 'lower WER' models and found them wanting (because of differences in their methodology for generating output), I'm super curious to get a feel for how v3 holistically behaves.
Will probably test it right now. :)
And I can confirm - my app Whisper Memos (https://whispermemos.com) is very popular in Czech Republic.
It makes perfect sense. Whisper is almost as good as transcribing Czech as English!
A funny note, if Siri is set in Korean mode and reads your texts that come in as English, they sound like a racist imitation of a Korean accent. It is absolutely hilarious.
it does works amazing in PT-BR Whisper V2, I can't even imagine it being better, and turns out, V3 promises it to be better...
A great example is that — for most words from any language that uses a subset of the Czech alphabet — a Czech speaker can just pronounce the word instead of spelling it and another Czech speaker will be able to write it down.
e.g. "messerschmitt", "nešamas", "cadeira", "philosophy", "tastaturi", "nicchia", "kaupunki", "abordagem", "povjerilac", "primauté" are all foreign words with very unambiguous pronunciation in Czech.
New models and developer products - https://news.ycombinator.com/item?id=38166420
OpenAI DevDay, Opening Keynote Livestream [video] - https://news.ycombinator.com/item?id=38165090
I need to write a lot of long texts for work and some good dictation software would be great. I know there's Dragon, but somehow I have not been able to find something that fits my need and is free.
https://github.com/foges/whisper-dictation
All Platform (hotkey)
https://github.com/doctorguile/faster-whisper-dictation
and if you use emacs
Is there a model that is the best at wake word detection? The last that I looked, it seemed like this was fairly lacking.
Balancing wake reliability vs false wake activation is a tricky balance. OWW is decent but could certainly be better.
It's used with Home Assistant now so I expect the training data and implementation overall to get significantly better fairly soon.
Edit: I understand that you can use small samples and approximate something like streaming, but the limitation here is you wind up without context for the samples, increasing WER. It would be nice if there was some streaming option.
The zeroes + faster model is how Gladia (mentioned in my other comment here) achieves live transcription by simply transcribing really short chunks one after the other, I believe.
For more advanced stuff you kinda have to get your hands dirty, which I've done for my own product (not linked).
There are other ways too with different trade-offs, can e-mail me at the link in my profile if you'd like to talk about how.
(I am definitely inspired as we don't currently provide such a straightforward experience in our own signup flow)
I tried it in English, French, and my broken Spanish, and all 3 came out great. One surprising thing is that if you switch languages mid-transcription with the "single language" model, it will transcribe the second language and translate it at the same time, so the entire transcription is in a single language, but the meaning is preserved.
It's inefficient, but even older gaming GPUs are fast enough for real time performance, and accuracy is good. If you were going to train a model from scratch for real time you could do something more efficient, but it works as is.
Edit: I'm not sure what you mean by "you wind up without context for the samples". You can supply context to Whisper.
>> Video Memory: 6 GB >> Graphics Processor: Nvidia GPU with 12 GB VRAM or greater
Would it work on an x64 Intel i7 machine with 32 GB RAM and with Nvidia RTX3060 with 6 GB RAM.
I'd like to support less VRAM. Maybe a future version will offload some of the processing to the cloud.
The idea was to provide transcription results as fast as you can, and you can refine it along the way by providing more and more context.
Here's a thread about running realtime on raspberry pi devices (some tweaks required): https://github.com/ggerganov/whisper.cpp/discussions/166
https://openai.com/blog/introducing-chatgpt-and-whisper-apis
> The Whisper v2-large model is currently available through our API with the whisper-1 model name.
Thanks for pointing that out, it's appreciated.
from openai import OpenAI
Traceback (most recent call last): File "<stdin>", line 1, in <module> ImportError: cannot import name 'OpenAI' from 'openai'
If so where is the current documentation?