OpenAI quietly launched Whisper V2 in a GitHub commit
github.com
github.com
* only an update to the `large` model (and therefore irrelevant for hobbyist, low CPU use cases)
* architecturally identical to the original whisper model
* just trained on more epochs
This is maybe the most important thing there is for a lot of cases since additional training is the thing that is just about impossible to do for most people.
> The major version number is incremented when the number after the dot starts looking "too big." There is literally no other reason. [1]
How come you don't support audio files longer than 1hr? Is it because of $$ cost?
The above demo app gets faster transcription by chunking audio and parellelizing over dozens of CPUs, so you can transcribe a about 1hr of audio for $0.10.
Interesting, which model are you using? We use the medium model which is the sweet spot between time/performance ratio. We also chunk, We try to detect words and silences to do better chunking at word boundaries but if you do more chunking and you don't get the word boundaries right it seems like whisper loses some context and the accuracy suffers. We will soon support longer hours. We just want to make sure the wait time for transcription doesn't suffer for most users. But great demo, reach out to me if you want to collaborate
For example, instead of:
Hello there!
Hi
How are you?
Good, and you?
You get something like: Voice A: Hello there!
Voice B: Hi
Voice A: How are you?
Voice B: Good, and you?would be lovely to see this feature open sourced.
This is the best AI diarization and transcription I’ve been able to get so far: https://github.com/zachlatta/openai-whisper-speaker-identifi...
> The Google Recorder app (...) transcribes meetings and interviews to text, instantly giving you a searchable transcription that is synced to the recorded audio (...) With this new update, the recorder can now identify and label each speaker automatically—an impressive feat. Google Recorder is exclusive to the Pixel 6 and newer Pixel devices.
https://arstechnica.com/gadgets/2022/12/pixel-7-update-adds-...
It’s ok, but the quality of speaker identification is nowhere near as good as the transcription itself.
I’d love to see models which try and use stereo information in recordings to solve the problem. Or, given a fixed camera and static speakers, I thought it should even be possible to use video to add information about who is speaking. There doesn’t seem to be anything like that right now tho.
You can do this pretty conveniently using pyannote-audio[0].
Coincidentally I did a small presentation on this at a university seminar yesterday :). I could post a Jupyter notebook if you're interested.
PS: Bai & Zhang (2020) is a great review on the literature [1]
If Whisper achieves 85% or higher accuracy on this audio, it would be a miracle. Garbage in, garbage out tbh. Project 25 needs to move to a modern codec, ideally not one seeing little development done by one small company.
Speaking of the telephone, they definitely could put improved audio quality as part of the 4G and 5G specs, but they don't. A modern telephone network is all IP anyways and backwards compatibility can be maintained.
* Better than anything else I've seen, but certainly far from perfect.
https://github.com/ggerganov/whisper.cpp/tree/master/example...
$ make base.en
$ ./examples/livestream.sh http://a.files.bbci.co.uk/media/live/manifesto/audio/simulcast/hls/nonuk/sbr_low/ak/bbc_world_service.m3u8 10
[+] Transcribing stream with model 'base.en', step_s 10 (press Ctrl+C to stop):
Buffering audio. Please wait...
here at the BBC in London. This is Gordon Brown, a former British Prime Minister, who since 2012 has been UN Special Envoy-
Lemboy for global education. We were speaking just after he'd issued a rallying cry on the eve of the 2022 Football World Cup. For
Governments around the world to pressure Afghanistan to let girls go to school. What human rights abuses are what are being discussed as we...
run up and start and have happened and have seen the World Cup matches begin. And it's important to draw attention to one human rights abuse.
that everyone and that includes Qatar, the UAE, the Islamic organization of countries, the Gulf Corps...
[0] https://github.com/ggerganov/whisper.cpp/blob/master/example...News, betting games and some shows HAVE to happen exactly at a certain time and are very rarely late so you can use these known times as checkpoints with some bias. So say I run this command `ffmpeg -i http://someexamplesite.fm/8b0hqm93yceuv -c copy -f segment -segment_time 60 -reset_timestamps 1 zip-%03d.mp3` that automatically chunks the stream into 1 minute files and I know news gets read at 1pm, I can merge everything from 10am to 1pm into say "Segment B" and then process that, you get the idea.
I also have a step in the pipeline after on the transcript after that tries to summarize what it can so any small gaps would likely be inferred as I have not noticed anything too wonky and the summaries and text I get so far have been pretty clean and good enough for my needs. Once I bench this some more however I'm sure I will have this and other interesting problems to solve, the one I'm fighting now a bit is ads that have music and fast talking.
Does anyone have solutions for clearing out "silence" from an audio file that works off something a bit more accurate than just "<= decibel x"?
Edited for grammar.
-------
I can't reply to the below, but you have to consider the difference in the signal to noise ratio for why it should be considered a different problem.
If I told a binary image classifier to classify a clear image of a cat as either a "cat" or a "dog", and it said "dog", then that would be an accuracy problem.
If I gave the same classifier an image of a black cat standing in a very dark room, where even a human would have trouble identifying it, and it says "dog" it's not an accuracy problem as much as a signal to noise ratio problem.
It seems like you're making the assumption that all of these have the issues you describe have the same root cause. I don't think that's a sound assumption...tehe.
[1] https://discuss.huggingface.co/t/open-to-the-community-whisp... [2] https://youtu.be/fZMiD8sDzzg?t=1226
[1] https://twitter.com/lunixbochs/status/1574848899897884672
If you have an always running server and only need access via a browser, forwarding port 80/443 to a reverse proxy works well. Dynamic DNS can be used if you don't have a static IP address. Alternatively, if you are more security conscious and don't mind the reliance on CloudFlare, you can use CloudFlare tunnel which won't expose ports from you local network at all. It can also be combined with a third party authentication mechanism.
Large-V1 (transcribed at 2022-10-02):
* SRT: https://pastes.io/yrggofqhof
* Transcript: https://pastes.io/rtp9buhsm0
Large-V2 (latest version at 2022-12-07):
* SRT - https://pastes.io/uiqblpw1qk
* Transcript - https://pastes.io/rheqgnftzl
There's still some timing issues after a period of silence, but using a VAD as a workaround (like I do in my WebUI) may no longer be strictly necessary:
Also, OpenAI is basically a subsidiary of Microsoft at this point.
See Imagen for example - https://imagen.research.google/
If Google is to be believed, this outperforms Dall-E - and I’ve heard from people that use it that in general, it does perform better than Dall-E.
I don't get Google's brand or leadership anymore. They're staring disruption to their core revenue stream in the face and acting like the "everything is fine" dog.
Meanwhile Microsoft is playing dimensional chess across multiple industries and key developing areas of research.
Are you saying they should have more bombast in their product announcements?
However although the models of today are interesting, they’re not consistent enough for many business applications. It’s not good enough to generate realistic cartoon characters most of the time but sometimes they randomly have 4 arms. Or you have a language model that can say interesting things but can’t be relied upon for basic facts. Maybe some day we will be there but not today.
Google actually surpassed its own model[0], first with Parti[1], which unlike DALL-E could even correctly insert text in the image, then with Imagen Video[2], which does as the name implies.
[0]: https://paperswithcode.com/sota/text-to-image-generation-on-...
In terms of granting direct access to machine learning models, OpenAI has Google beat. But that’s not Google’s business model. And it remains to be seen if it’s a viable business model at all.
I uploaded it to a hugging face space if you just want to fiddle with it without installing
For instance, I know Turkish is very consistent: it was refactored in 1928 with the birth of Turkey. Turkish is quite high in the rankings. I don't think because there's loads of data available, but because of its consistency. Contrary, English has loads of data, which should compensate for it inconsistency.
Radio dispatch is very hard to understand, low quality audio. AWS transcribe was essentially useless, incomprehensible transcriptions. Whisper was 95+% accurate, with maybe 1 incorrect word per few sentences, and it was often easy to tell what the correct word would be.
I could believe that other languages with more regular grammar might be easier to parse. Or maybe Romance languages lean on more easily distinguished phonemes.
English also tends to take delight in stealing words from other languages, and then using them in the wrong way. Seemingly in an effort to drive speakers of the other language up the wall.
WER for the exact same model can vary wildly between datasets, it’s really only useful for comparing model performance on a single test set.
(I’m under the impression they don’t charge)
Anyone ran tiny en on t4 gpu (aws g4dn iirc)? What’s the speed up
on tesla-t4-30gb-memory-8vcpu google cloud
on tiny and tiny.en
for 10 minute = 30 seconds
on medium
for 10 minute = 1m 30s
for 60 minute = 7m
on large
for 60 miutes = 13m
on NVIDIA GeForce RTX 4090
on tiny
for 10-minute = 5.5 seconds
for 60-minute = 35 seconds
on base
for 10-minute = 7 seconds
for 60-minute = 50 seconds
on small
for 10-minute = 14 seconds
for 60-minute = 1 min 35 sec
on medium
for 10-minute = 26 seconds
for 60-minute = 3 mins
on large
for 10-minute = 40 seconds
for 60-minute = 3 min 54 sec