Universal Speech Model
sites.research.google
sites.research.google
I messed with a combo of Whisper and ChatGPT. I took a whisper transcript and asked chatgpt to fix mistranscriptions using the context of the transcript and based on potential phonetic issues. I asked it to replace transcribed words that don't make sense with "[unintelligable]", which improved the output even more.
Transcription error rate was almost nonexistent, even on the smallest whisper model.
Try to pre-process it where just "voice" is detected, not the meaning, just some speaking, and cut the audio into snippets which only contain speech, so that Whisper doesn't have to guess if the segment will contain speech or not.
Also, if you cut it up into chunks and let it transcribe each chunk and expect JSON as the output, instead of the other output methods, then you'll get a bunch of extra parameters with it which will help you identifying problematic sections. For example hallucinated repetitions usually have a higher "no_speech_prob" parameter, or segments with lower "compression_ratio" will also not be that accurate.
Personally wouldn't recommend my approach however unless you are okay with doing some hardcore text manipulation and fuzzy math. It fails to produce text that matches up with the prompt 10% of the time and lots of other caveats with what to do with text that doesn't fit into the prompt.
In my case I'm recording notes while exercising on my bike, so there's wind and friction and breathing, then I transcribe them with Whisper, do try to filter out things that don't make much sense (where the temperature, compression_ratio, avg_logprob and/or no_speech_prob are not acceptabe).
Since I start recording at will and don't prepare the sentence to be recorded, the recording is a bit chaotic, and ChatGPT does a good job in correcting Whisper's mistakes and transforming the sentences.
But, as I said, sometimes this fails and the real message gets lost.
I'm using them to record multiple people in the same room but avoid cross-talk (for data collection in a research study). The audio isn't great, but I've used transcription services and they seem to be able to make the words out just fine.
Not ideal, maybe I should've asked to censor with asterisks instead of changing the whole world
[0]: https://github.com/ggerganov/whisper.cpp#real-time-audio-inp...
I find this part to be the most impressive thing in the OP. Most of those languages are spoken by fewer than 0.1% of the world and Nzema is spoken by less than half a million people. Where are they even getting enough training data to figure those languages out?
What this implies is that the core encoder model (trained on popular languages with a ton of data) does a really good job at learning the generalized basics of any language (period).
Then, they enter a comparatively minuscule amount of labeled data from the small languages at the end (supervised fine tuning - since parameters are tuned iteratively, the latest training data has the biggest impact on the final performance). The model has enough of a general “understanding” of grammar across languages that it can fill in the gaps it doesn’t get from that small labeled set.
Thats the true crazy part. The pre trained encoder is an ersatz Rosetta Stone.
Understanding those generic internal structures would be fascinating.
Also, for all we know, they failed on a bunch of languages (unless there’s a paper somewhere I missed, but sadly the days of good AI papers from for-profits seems to be over).
For example, if it had historically turned out (or will turn out in some future) that everyone on Earth speaks some variant of English or other Indo-European language; or some variant of Chinese or other Sino-Tibetan, then there would be many features of grammar that technically are universal across the population but that wouldn't be evidence that "universal grammar" (as used by Chomsky) is indeed a thing.
It’s like Chomsky, but with the Norvig approach.
The more languages from a language family a human speaks, the easier it is for them to learn each additional language from that family. It stands to reason that deep learning can benefit similarly.
what does this mean?
Norvig was one of the pioneers in proposing that NLP could be solved by statistical machine learning on large quantities of data.
So the joke here is that we're using large-scale machine learning to infer the underlying universal structure of language.
---
EDIT: After writing the comment above I asked ChatGPT-4 the same question, and it came up with the following answer, which I think is flawless.
---
The last sentence in the paragraph is referencing two well-known figures in the field of linguistics and artificial intelligence: Noam Chomsky and Peter Norvig.
Noam Chomsky is a linguist who proposed the idea of a "universal grammar," suggesting that all human languages share a common underlying structure, and that humans have an innate ability to acquire language. This idea implies that even with limited exposure to a specific language, humans can still learn it thanks to this innate structure.
Peter Norvig is a computer scientist and artificial intelligence researcher who is known for advocating a more data-driven, statistical approach to natural language processing. This approach emphasizes the importance of large datasets in teaching machines to understand and process language, rather than relying on pre-defined rules or structures.
The sentence "It's like Chomsky, but with the Norvig approach" suggests that the described language model incorporates elements of both Chomsky's universal grammar and Norvig's data-driven approach. The model benefits from an underlying structure (akin to Chomsky's universal grammar) while also leveraging a large corpus of various languages (as advocated by Norvig) to improve its performance in understanding and processing language.
Also, have you heard of Conformer-1 by Assembly-AI[1]? It released a few days ago and supposedly scored higher than Whisper on various benchmarks.
As for speed I have no idea how they make it so fast, but I'm sure they've written about it somewhere. My guess is at least that they are slicing the audio and parallelising it. Will look into Conformer-1 as well!
Whisper is mostly an academic toy.
Yesterday I wrote virtually all the prose in the manuscript while walking around with a friend and discussing it. We didn't even look at the phone.
Obviously there's an academic element here because I'm saying I'm using it for writing. But it's more of a human-centric computing thing. I'm replacing a lot of time that my thumbs are spent tapping on keys, my fingers are spent tapping on keyboard, and my eyes are spent staring at the words that are appearing, looking for typographical errors to correct, with time organizing my thoughts in a coherent way that can be spoken and read easily. I'm basically using whisper to create a new way to write that's more fluid, direct, and flows exactly as my speech does. I've tried this for years with all of the various ASR models on all the phones I've had and never been satisfied in the same way.
I also use whisper.el on emacs. It's amazing. Much more powerful but computer based of course.
That's a big deal for someone who wants to build something using speech recognition (voice assistant, transcription of audio or video) without resorting to APIs.
Look, it's nice that it's out there and free to use. For the sake of my wallet I hope it gets really good. But it isn't competitive with top of the line commercial offerings if you need to ship something today.
While true real-time would definitely be nice, I can approximate it well enough with various audio slicing techniques.
Vocode uses Whisper for real-time zero latency voicechat with chatGPT. Give their demo line a call to see how well it works: +1-650-729-9536
Another is Deepgram. Even this obscure vendor seems to be able to handle the samples I tried better than Whisper: https://picovoice.ai/platform/cat/
But yeah, go with Azure as your starting point. It is good and the price is likely acceptable unless you're transcribing all of youtube.
Lowest WER in the industry, cheaper than Whisper API, and we have an on-prem solution
It is comically bad, with nonsense words that don't make any sense in context about every other sentence in English, and an absolute inability to produce a coherent thought in Japanese.
But it has come a long way since then.
I think this says more about the generally-low quality of the captioning services used by YouTube content creators than anything, though.
My biggest complaint is that I often find that human captioners don’t recognize domain-specific words. Like, names of ethnic foods. The word might be literally showing on the screen while the presenter is talking — and the “professional” captioer will just put [indistinct] there, as if they aren’t even watching the video as they’re captioning it. YouTube’s auto-captions get these words right every time.
My impression is that USM is seemingly uniquely good at code-switching from word to word within a sentence — which makes sense, given its “universality.” I think, if they allowed it, it would even be able to embed clauses and quotations generated in one language and alphabet, into a sentence generated in an entirely different language and alphabet, keeping syntax and grammatical structure correct for the given language within its clause.
I have had some channels where I was laughing out loud at the autocaptioning, which was probably more so translation, but I did get a laugh out of it after all, and I generally knew what they were saying.
I've also noticed that, at least with the videos I watch, usually the autocaptioning errors seem "phonemically correct", so it's substituting a word that sounds the same, and I can easily figure out what was meant. Usually I've noticed these problems with more so with non-American English (British or Australian for example), especially where there are multiple people all speaking English, but with different accents. It does seem to me the English speech recognition is honed in on some west coast or midwest US English accent.
I am surprised contextual cues aren't being utilized more but I'm very happy with the YouTube speech recognition.
It's only 1% better than the current state of the art. But it's still noteworthy. From the end of the abstract:
> We demonstrate that utilizing a large unlabeled multilingual dataset to pre-train the encoder of our model and fine-tuning on a smaller set of labeled data enables us to recognize these under-represented languages. Moreover, our model training process is effective for adapting to new languages and data.
It's amazing to me that the chaotic process of "machine learning" can end up with an internal state for languages that is readily adapted to entirely new languages.
For now, they've got this handling audio transcription, but with some hints that this approach could work well for translation. Perhaps we'll be able to use these improved models to decipher Linear A[0] or other un-deciphered languages. It sounds like "magic", but it's the kind that could maybe exist.
Yeah and interestingly, that was roughly Chomsky's breakthrough finding with respect to how humans learn language as children. We are born with an innate language acquisition device.
Here's an alternative theory/approach. What if natural languages are just the way the device starts working when the number of neurons grows quickly? NL properties sort of emerge out of low-level details of brain work? Neurons are simple but the brain is not. Complex brain properties emerge from trivial parts the same way our full bodies emerge from a simple DNA/RNA system. Any details in these systems would be too statistical to expose a limited rules system.
Obviously, powerful enough ML system can infer the system's properties. In fact, it can infer any function. The thing is that this doesn't mean there's some kind of simpler model explaining details of emergent system's work.
What is surprising is the way LLMs imitate a stateful function (our brain, with memory, fluid biology, etc) using a stateless inferred function (the model). I suspect this statefulness might be the answer to the question of "poverty of stimulus" problem.
YouTube live transcriptions are terrible, because they get confused by homonyms and can't follow the context in a sentence.
In the same manner that Dall-E joined an LLM to an image generator, they ought to train a combined speech model + LLM so that the uncertainties in the speech model output is disambiguated by the LLM.
Surely the speech recognition model itself learns some basic language statistics in order to recognise homophones?
> YT-513-U: We create an additional dataset called YT-513-U to ensure coverage of lower resource languages in our pre-training dataset. We reached out to vendors and native speakers to identify YT videos containing speech in specific long tail languages, collecting a dataset of unlabeled speech in 513 languages. [1]
Totally free.
This needs to be much better to make sense and their own graphs show only marginal improvements in specific scenarios.
The results from Whisper are incredible, with very few mistakes. Though it did get Nelson Mandela's first name wrong (transcribed as Nesson). What's more, Whisper finished transcribing a 60-minute audio stream in 20 minutes on commodity hardware (T1000 G8 NVIDIA GPU). Broadly, here are the steps I used:
* Download and install podman.
* Download and install git.
* Download and install curl.
* Open a command prompt.
* Run the following commands to containerize Whisper:
git clone https://github.com/lablab-ai/whisper-api-flask whisper
cd whisper
mv Dockerfile Containerfile
podman build --network="host" -t whisper .
podman run --network="host" -p 5000:5000 whisper
* Download MP3 file (e.g., filename.mp3).* Run the following command to produce a transcription:
curl -F "file=@filename.mp3" http://localhost:5000/whisper 33,53s user 2,05s system 443% cpu 8,023 total
with the 'tiny.en' model whereas whisper.cpp gives me 22,71s user 0,12s system 745% cpu 3,062 total
with the 'base.en' model for a 15s audio clip on an i7-3770 (8 threads).In my workflows I've found rare but noticeable quality differences between the model sizes. So when practical I try to use the larger ones.
Now installing the dependencies of every git repo I want to try on my host system, that's how an environment becoming needlessly complicated
For instance, this was added to the transcription of a silent section of audio:
> Hello everyone welcome to my channel. Today Im going to show you how to make a very simple and easy recipe. I hope you enjoy the video. If you like my recipe dont forget to subscribe to my channel
It makes me wonder how much of Whisper is trained on audio from Youtube, which was transcribed by this model.
Now if only the timestamp timings were correct for noisy audio... I've tried stable whisper and another fork I forget the name, but I need to run the audio through RTX voice if I want consistent timestamps...
I'm imagining the near future will see a portable fast streaming model for real-time voice translation, piped into text to speech. Hook it up to an earpiece and you've got a real-life Babelfish
Source, me with my Eastern Europe accent. There are engines out there that do far better with my accent. FAR better.
I really don't care about closed API models of anything that has a good/usable open source version. Whisper works well enough, I'm never going to follow up on this USM research or use it. The only reason to pay for the API access would be for some super niche language. And if Google is paywalling this only for the few customers who need it for use in terribly under-represented communities ... that's a kind of douchebaggery all its own.
The only reason people are paying for OpenAI's GPT-4 is because there's literally no usable open-source LLM. The instant a "good enough" one exists, OpenAI's revenues will drop by >95%.
Hopefully Google will at least use this in Google Home because it's still bad enough to notice.
Regardless, I don't want to use the API, but I'm working with public information anyway; and so, while I have considered moving to Whisper now that that's an option, it hasn't been a priority and it isn't clear to me that Whisper is good at random non-English languages anyway.
Which probably works out to one less error per thousand words or something crazy like that.
The state of the art is pretty much better than humans at this point iirc.
I know that's not the point of this model (the point is that for a lot of languages, its the only model available). But paywalling it seems greedy, you'll only extract money from those under-represented communities. On the other hand, maybe this never would have been built without the profit motive. Idk. I wish we could fund these things as "basic science research" without a need for direct profit. Let positive externalities pay us back down the road.
That being said, Speaker Diarisation is still a problem that hasn't been fully solved. As of yet, AI hasn't been able to outperform humans in this area.
I feel like its watching an ad to buy a product but no 0800, QRCode or Web URL to go and use it. So frustrating.
Is there any good quality, open source multi-language TTS setup available? Flite works, but doesn't sound very good.
Three design documents, committee review LGTM, director-level LGTM, VP approval for headcount of person to make the form, HR forms to process head count to make the new form, legal review, privacy review, security review, and finally one of the original LGTM from committee was laid off, so you need to start over.
Maybe it works well with native speakers? But since it's supposed to be so multilingual I hoped that it would work well with my accented speech... maybe that's a wrong conclusion to draw.
(Of course that might just lead to picking more common translated phrases.)
Humans train to recognize speech on much smaller datasets. If a small human is awake 16 hrs/day, that amounts to maximum 5840 hrs/year or 58400 hrs per 10 years. Why do mathematical models use more data and produce lower quality results? Is it because they don't understand the meaning of words?
But no human can understand 300+ languages, especially not at 10 years old.
Still a very approximate comparison, but if you multiply your upper bound by 300, that gives 17.5 million hours, so more than what was used to train.
We evolved into having these brains that learn in stages, when the early years are responsible for core functions (walking, running, talking, etc). There's a lot of input (all the senses) that feeds learning, while in later years we mostly just take for granted whatever we learnt as children.
It's not from a department I'm a fan of, but they do know how to catalog gestures.
I did that with Whisper output and it helped improve a podcast transcription and logically inserted proper line spacing.
I pasted the auto-translation into ChatGPT and asked it to summarize:
> The speaker seems to be discussing the importance of Nyepi, a Balinese Day of Silence. They mention that happiness, peace, and prosperity can be achieved by being in harmony with space and time. The speaker also references ogoh-ogoh, which are statues symbolizing negative influences, and the Panca Maha Bhuta, or the five elements of life. They suggest that negative emotions and behaviors can lead to a "darkness of the mind."
> The speaker emphasizes the importance of self-control and using the Nyepi celebration as a milestone for personal growth. They mention the parading of the ogoh-ogoh, which represents negative behaviors being confronted and released. Following the Nyepi celebration, people should aim for a new and better life, leaving behind their past negative behaviors and emotions.
I'm not convinced of the accuracy, but I'd definitely say it's useful.
Intent translation generally requires you to say your sentence in your language, then the translator adds/removes words and can swap order of stuff around.
OpenAI – write that down, write that down!
I keep thinking that there's no way that Brain/DeepMind are just getting stomped, lapped, generation-gapped by e.g. `ChatGPT`: they must have had an internal demo of this sort of thing like 2 years ago, right? At some point the Empire strikes back?
But the rollout and product integration has been so well done, so coordinated and cohesive that it's now just obvious that it was way too soon to count MSFT out of the game. It's all through search and Office 365 and GitHub/CoPilot/etc. and the whole stack in such a legitimately compelling way that you can almost forget that the DNA is Win32.
It's a bad thing to let Microsoft get a stranglehold on developer and user mindshare network effects: the 90s were rough. But with how cool it all is it's very tempting...
Everything since then has been a combination of algorithm and compsci research (which they are world-class in; credit where credit's due), vague ideas about things people might like, and copies (or buyouts) of their competition. They remind me of my engineering friends who tried to come up with business plans in college...good at building things, but clueless about figuring out what people actually need, what they should therefore actually build, and how to make it user-friendly. You know, all the stuff that you need if you want to run a business. (As I've said before on hn, their initial competition against youtube is a great example of this)
It's a surprise that a technology came along which upended them so abruptly, but it's been clear for a long time that they were only alive because their search engine couldn't be beat, and they didn't have a clue how to replicate that success.
Google hasn't needed to generate another monster revenue stream outside of ad sales, so it's possible that they never really tried all that hard (they've certainly killed it on the things that protect it, notably Android and Chrome). An utterly dominant position in how people access information that lasts for decades is probably "a hell of a drug".
For example, GCP is technically a really, really good cloud offering, maybe even the best for a lot of use cases (if you haven't looked at it lately, Cloud BigTable looks friggin amazing, I wish I'd had that database for the last ten years). They've obviously failed to parley that technical achievement into dominant market share, but maybe with the pressure on around search they get serious about whatever combination of pricing and marketing and customer support that gets them some serious market share.
YouTube has been quietly building their TikTok competitor into something I'm actually starting to waste some time on, they people who work on that are clearly really good at their jobs even if they started a little slow.
And on the LLM space, honestly I'm rooting for them: MSFT/OpenAI/ChatGPT need at least some competition and they are probably the best positioned to do so. Facebook/Meta is also doing this stuff in a more "open" way and that's keeping the pressure on around some competition too.
In general this LLM stuff is going to be a great thing long-term, but letting one company dominate both mindshare and marketshare is going to make that a much rougher road for society than if it's avoided.
Its been going on for some time. Something that was once a joke in good jest AKA Google's graveyard, is now their actual reputation, and helping their strong big tech competitors when competing on new services.
Oh and of course their model is hidden behind an API.
If it can memorize all possible waveforms across all those languages and how they convert into words in only 2B parameters then I'm incredibly impressed.
(Of course it isn't doing this, and this criticism is just wrong.)
> Model implies they had some sort of insight, a simplification, an understanding of how natural language works.
Well it shows that using the same model for multiple languages increases performance which was something that was not at all clear 10 or even 5 years ago.
> This is just bragging about the number of parameters they can pull off.
A humble brag about how small the model is maybe?
Well then I don't think you've looked at the topic very deeply. Our voices don't make every possible wave form. The IPA alphabet has existed for a generation now, is universal for all spoken human languages, and has a basis in physiological origins. Of more than 160 IPA symbols, relatively few will be used to transcribe speech in any one language. It should not take 2B parameters to do this.
Look at the inner ear and you'll see a big hint. Before even getting to the brain the sound goes through a spiral shaped chamber. Quote wikipedia: "The hair cells in the Organ of Corti are tuned to certain sound frequencies by way of their location in the cochlea...". Clearly looking at time series wave form data in the first place is a mistake. Rolling window fourier transform all the speech data first, then train on it. I will bet 2-digit sums of money that simple preprocessing step would outperform what they've done.
Really, take any recording of people talking, open it up with audacity, and check out the spectrogram view. The sort of sounds we make with speech are a lot like an FM radio transmission. There's a baseband average pitch that we speak at, and then information encoded on top by deviating over/under the base band in smooth, simple ways. If that spectrogram was an OCR problem, it would already be solved.
>A humble brag about how small the model is maybe?
I see a B at the end of a number I assume its part of a pissing contest. But if they really are playing parameter golf bragging about how small they can get it, then I admit that's a step in the right direction.
If this is right, one could interpret it as phonemes being digital (a native speaker can consistently transcribe them using a small number of IPA symbols), but (at least vowel) phones being analog (the typical frequency spectrum appearance typically representing a specific vowel could best be represented by floating-point numbers that are culturally chosen somewhat arbitrary by language evolution). Like /ɪ/ might be 0.82 closeness, 0.23 backness, 0.05 roundedness in some language, but 0.79 closeness, 0.26 backness, and 0.1 roundedness in another.
In turn, if that's right, I don't know how much variation there is among native speakers (from a particular region) in production of a given vowel (or how consistently and precisely children can learn these values). Presumably for recognition the brain rounds off somehow to the nearest vowel phoneme, kind of like picking the nearest point on the constellation diagram in digital communications demodulation?
that's not 3 floats worth of data, those are all ints over 100.
We generally think of people, but not languages, as having deep or squeaky voices. Furthermore we usually have no problem following when the speakergoesfaster or enunciates reaaaallllyy slooowwwlllyy or even ish they slur theirr wurds abit. I don't think a language specific absolute value of vowel tone is a real thing. Rather, what matters is that the direction and magnitude of the shifts in tone are all in consistent proportions relative to each other. When you look at the vocalizations of humans or even other mammals on the spectrogram there's very clearly just simple modulations happening, like moving linearly from a starting frequency to the ending frequency, or at most taking a x^2 curve to get there. When you think about it this makes sense as those sorts of patterns are trivial to implement on the hardware. Just squeeze or release the muscle on the vocal cord harder. You're right in principle any set of frequencies can be the vowels but every language is going to discretize down to a fixed number. So instead of having or needing language specific absolute "vowel frequencies", you can just look for k distinct levels in the individual speakers voice.
If you've gotten as far as calculating the spectrogram, everything I described can be easily implemented by applying curve fitting to that spectrum and then thresholding the coefficients
But if a native English speaker speaks Japanese with five appropriately-spaced English vowels, it will be directionally correct and comprehensible, but native Japanese speakers would notice that those aren't the normal vowels used by native speakers in Japanese.
So I guess I'm wondering about how much leeway there is in what sounds like a correct vowel to a native speaker, how much range there is in a given person's production, how much range there is in production among members of a community who interact with each other constantly, and how small a difference between vowels someone can detect as a listener.
Okay. I take this bet. If your simple preprocessing step outperforms whisper on, say, me reading a random wikipedia article, I'll happily send you any 2-digit sum of money (denoted in dollars presumably), or donate it to a charity of your choice.