In order to achieve modest improvements in dictation we're throwing entire GPU arrays at the problem. What happened in the middle? Was there really no room for improvement until we went full AI?
In order to achieve modest improvements in dictation we're throwing entire GPU arrays at the problem. What happened in the middle? Was there really no room for improvement until we went full AI?
Also, the original MacinTalk sounded a lot better if you fed it phonemes instead of text. It didn’t know how to pronounce that many different words, and wasn’t really good at making the right choice when the pronunciation of a word depends on its meaning.
For example, if you gave it the text “Read me”, it always pronounced “Read” in the past tense. That always seemed the wrong bet to me, and I would think the developers had heard that, too, but apparently, fixing it was not that simple.
I also think it didn’t know to pronounce “Dr” as “Drive” or “Doctor”, depending on context, or “St” as “Saint” or “Street”, to mention a few examples, and probably was abysmal when you asked it to speak a phone book, with its zillions of rare names (back in the eighties, that’s an area where AT&T’s speech synthesizers excelled, I’ve been told)
And that’s just the text-to-phoneme conversion. The arts of picking the right intonation and speed of talking are in a whole different ball park; they require a kind of sentiment analysis.
chee-hoo-ah-hoo-ah
You can also combine the linear and nonlinear approaches. LPCNet uses a bit of each, for example: a linear lpc prediction, which is then corrected by a neutral net.
The different systems (klatt, MBROLA, etc.) build on these basic techniques. The *OLA systems are more space intensive, as they need more of the audio data to synthesize the voice, but tend to result in more natural voices.
The other complexity is how the individual phonemes (building blocks of speech) are stored. Different languages and accents have different sets of phonemes, so more complex languages require more data. Systems like MBROLA store the data as diphones (from the mid-point of the first phoneme to the mid-point of the second), to make it easier to join the phonemes. There are more diphones, although not all combinations are in a given language and accent.
This data is controlled by different prosodic parameters (e.g. pitch and duration). Neural nets can be used to control these parameters, or techniques like decision trees and probabilistic models.
IIUC, there are two approaches to neural nets when using audio generation: 1) generate the LPC/formant parameters; 2) generate the audio directly. These can either operate on phonemes, or on the text directly. Operating on the text directly limits the voice to a particular language and accent, but potentially allows the neural net to infer pronunciation rules itself.
Hmmmm.... In my experience, for (neural) TTS the dominant option is to have one model to generate a melspectrogram from the text (handling prosody) and a second model for synthesizing samples from the melspectrogram. (Tacotron, Lyrebird, and this Facebook group are all doing this.) There's certainly research projects on going directly from text to samples, but it's not the currently winning strategy... Maybe eventually, though. The Text->Mel portion specializes the prosody and pronunciation problem, and provides a nice place to add extra conditioning.
On the vocoder side: LPC, F0, etc can all be estimated from a reasonably sized melspectrogram; for the most part, these neural models are just letting the big vocoder model handle all of these things which are traditionally (fragile!) subtasks. The question is which "classical" parts are both cheap and reliable: you can compute these on the side and lighten the neural network's burden. LPC is great for this.
But the latest neural techniques have added quite a bit of naturalness, and their computational requirements, while high, are within the reach of consumer level devices.
I think that's a little rose tinted. Classical speech synthesis was awful in comparison to this.
"How come we never go out anymore?"