Totally free.
This needs to be much better to make sense and their own graphs show only marginal improvements in specific scenarios.
Totally free.
This needs to be much better to make sense and their own graphs show only marginal improvements in specific scenarios.
The results from Whisper are incredible, with very few mistakes. Though it did get Nelson Mandela's first name wrong (transcribed as Nesson). What's more, Whisper finished transcribing a 60-minute audio stream in 20 minutes on commodity hardware (T1000 G8 NVIDIA GPU). Broadly, here are the steps I used:
* Download and install podman.
* Download and install git.
* Download and install curl.
* Open a command prompt.
* Run the following commands to containerize Whisper:
git clone https://github.com/lablab-ai/whisper-api-flask whisper
cd whisper
mv Dockerfile Containerfile
podman build --network="host" -t whisper .
podman run --network="host" -p 5000:5000 whisper
* Download MP3 file (e.g., filename.mp3).* Run the following command to produce a transcription:
curl -F "file=@filename.mp3" http://localhost:5000/whisper 33,53s user 2,05s system 443% cpu 8,023 total
with the 'tiny.en' model whereas whisper.cpp gives me 22,71s user 0,12s system 745% cpu 3,062 total
with the 'base.en' model for a 15s audio clip on an i7-3770 (8 threads).In my workflows I've found rare but noticeable quality differences between the model sizes. So when practical I try to use the larger ones.
Now installing the dependencies of every git repo I want to try on my host system, that's how an environment becoming needlessly complicated
For instance, this was added to the transcription of a silent section of audio:
> Hello everyone welcome to my channel. Today Im going to show you how to make a very simple and easy recipe. I hope you enjoy the video. If you like my recipe dont forget to subscribe to my channel
It makes me wonder how much of Whisper is trained on audio from Youtube, which was transcribed by this model.
Now if only the timestamp timings were correct for noisy audio... I've tried stable whisper and another fork I forget the name, but I need to run the audio through RTX voice if I want consistent timestamps...
I'm imagining the near future will see a portable fast streaming model for real-time voice translation, piped into text to speech. Hook it up to an earpiece and you've got a real-life Babelfish
Source, me with my Eastern Europe accent. There are engines out there that do far better with my accent. FAR better.
I really don't care about closed API models of anything that has a good/usable open source version. Whisper works well enough, I'm never going to follow up on this USM research or use it. The only reason to pay for the API access would be for some super niche language. And if Google is paywalling this only for the few customers who need it for use in terribly under-represented communities ... that's a kind of douchebaggery all its own.
The only reason people are paying for OpenAI's GPT-4 is because there's literally no usable open-source LLM. The instant a "good enough" one exists, OpenAI's revenues will drop by >95%.
Hopefully Google will at least use this in Google Home because it's still bad enough to notice.
Regardless, I don't want to use the API, but I'm working with public information anyway; and so, while I have considered moving to Whisper now that that's an option, it hasn't been a priority and it isn't clear to me that Whisper is good at random non-English languages anyway.
Which probably works out to one less error per thousand words or something crazy like that.
The state of the art is pretty much better than humans at this point iirc.
I know that's not the point of this model (the point is that for a lot of languages, its the only model available). But paywalling it seems greedy, you'll only extract money from those under-represented communities. On the other hand, maybe this never would have been built without the profit motive. Idk. I wish we could fund these things as "basic science research" without a need for direct profit. Let positive externalities pay us back down the road.