I've been meaning to pull all the audio for a while and this has inspired me.
A fun thing to do would be to pass through Whisper, a great corpus to play with.
I've been meaning to pull all the audio for a while and this has inspired me.
A fun thing to do would be to pass through Whisper, a great corpus to play with.
I wonder... is there a way to "prime" Whisper (e.g. with the embedding of the episode synopsis) so that it "listens out" for words related to a particular topic? I haven't looking but this would be neat!
prompt
string
Optional
An optional text to guide the model's style or continue a previous audio segment. The prompt should match the audio language.I will also say that I've personally been unimpressed with the new "large-v2" model, even though it supposedly scores better. The original "large-v1" model seems to work better than the "large-v2" model in the audio clips I've been testing Whisper against, but results will vary. In general, I find I'm really happy with what "small.en" and "medium.en" will emit, and they're much faster than the large models. (The ".en" models are specialized to English, and usually perform better for strictly English input, whereas the non-".en" models are trained on multiple languages.)
Update: turns out this exists already: https://platform.openai.com/docs/guides/speech-to-text/promp...
This was the same audio sample fed through each of those four models, with no "prompting" as you're discussing.
If I add "--initial_prompt ChatGPT", then all four models are able to get the spelling correct.
Regardless, I don't think "chat GPT" versus "ChatGPT" is a huge deal. There will always be some level of uncertainty and ambiguity in the transcript, and even books written by humans always have a few typos get past multiple stages of copy editing. Perfection is virtually unachievable, but you can always scroll through the transcript and make some edits after the fact, if desired. Maybe some future model will magically eliminate all typos.
I cleaned that bit up with a bulk replace of "chat GPT" with "ChatGPT" in VS Code.
API used and results can be found here: https://techtldr.com/transcribing-speech-to-text-with-python...