VALL-E: Neural codec language models are zero-shot text to speech synthesizers
valle-demo.github.io
valle-demo.github.io
Or text messages that you can listen to in the voice of the people who sent them.
Or the death of the audiobook business? Any book read in any voice you want.
Or maybe a form of extreme compression, voice is converted to text with Whisper, sent over the wire as text, and re-created with the same voice on the receiver.
Or train it with the voice of the computer from TNG, pair it with ChatGPT and now I have the perfect digital assistant.
> Ethan wondered rather fearfully if Cee were reading his mind right now-apparently not, for the Cetagandan expatriate gave no sign of realizing his mistake yet.
> "I take it," said Ethan, "that your powers are intermittent."
> "Yes," replied Cee. "If my escape to Athos had gone as I'd originally planned, I meant never to use them again. I suppose your government will demand my services as the price of its protection, now."
> "I-I don't know," answered Ethan honestly. "But if you truly possess such a talent, it would seem a shame not to use it. I mean, one can see the applications right away."
> "Can't one, though," muttered Cee bitterly.
> "Look at pediatric medicine-what a diagnostic aid for pre-verbal patients! Babies who can't answer, Where does it hurt? What does it feel like? Or for stroke victims or those paralyzed in accidents who have lost all ability to communicate, trapped in their bodies. God the Father," Ethan's enthusiasm mounted, "you could be an absolute savior!"
> Or the death of the audiobook business? Any book read in any voice you want.
Listeing to audiobooks read in my own voice? Oh no...
https://store.dreamtonics.com/product/solaria-voice-database...
(Of course right now you can only use voice banks they've pre-trained, but I bet in a few years you could fine tune it on recordings of your own voice)
> I bet in a few years you could fine tune it on recordings of your own voice
Something to keep an eye out for though.
“Hi dad this is Sally. I need some money.” …
“Sure, listen can you just reset my password. This thing has driven me nuts.”
The fragment for "Jackson" has a very clear robotic distortion about 10 or 12 seconds in, for example. Madison and Helena are better, though they have weird pacing issues.
I couldn't stand to listen to these voices for a long time. With how good 15.ai's female voices are (when they come out with new versions), I must say I expected better.
https://9to5google.com/2020/12/03/google-auto-audiobook/
https://www.publishersweekly.com/pw/by-topic/industry-news/a...
https://support.google.com/books/partner/table/10957334?hl=e...
This is not realtime speech to text as we think of it generally.
Direct link to GitHub: https://github.com/neonbjb/tortoise-tts
Sorry to be a buzz kill but it’s important to recognize the drawbacks too as we continue to move forward.
I am scared now that this won't work any more, and it seems like here the drawback heavily outweigh the benefits...
I've told everyone I know (not many people!) that I hate VM, and don't usually answer it - certainly not promptly. I do usually check the missed-calls list, and call back. I never leave VM; I hang up when the robot starts talking.
The point of a phone is immediate, synchronous, person-to-person voice communication. VM comprehensively defeats that. You might as well write a letter.
(Probably can't bio print a brain, but you don't need to, put a Raspberry Pi where the brain should be and remote-control it instead).
Remember those Xerox scanners that randomly changed numbers and letters around, because the compression algorithm sometimes thought a 0 looked a bit too much like an 8 and got them mixed up? I can't wait until that starts happening for spoken words. Even in the examples demonstrated here you can hear quite a few flubbed lines.
I worked with someone who used a Stephen Hawking style voice application for a while. Its really weird having a conversation with someone like that. There's a delay between everything people say. It's a lot like using IRC on a really laggy connection. I quite enjoyed talking to him because everyone had time to think between while he 'typed'. The point is though, even if the sound his app used was a perfect recreation of his own voice it wouldn't be like talking again. It's a different experience.
https://www.ted.com/speakers/rupal_patel
Not sure what became of vocalid.org but would’ve benefited immensely from these advances.
Edit: looks like it lives on as vocalid.ai
But for example, some Discworld audiobooks have very specific actors, and I would not change them for anyone else.
* Video game characters will speak lines generated on-the-fly depending on in-game context, instead of lines pre-recorded by voice actors. Game makers can train LLMs to generate lines of dialogue for different characters given the state of the game, and have the characters speak those lines.
* We will eventually see the death of voice acting in all its forms -- video games, cartoons, advertisements, etc. Inevitably, we'll see famous actors and their legal representatives figuring out how to secure rights to their recognizable voices.
* Spammers and criminals will start using familiar voices to con targets ("listen, I've been kidnapped while on vacation; please send the money now"). Sooner or later, dark-web hackers are going to harvest samples from everyone's shared video and audio clips at scale.
* Videofakes are going to get a lot more realistic and interesting, with faked characters speaking just like the real ones. Famous actors and public figures should brace for the impact this technology will have on black-market uses, including (as always) pornography.
> * Video game characters will speak lines generated on-the-fly depending on in-game context, instead of lines pre-recorded by voice actors. Game makers can train LLMs to generate lines of dialogue for different characters given the state of the game, and have the characters speak those lines.
We already have incredibly cheap voice acting relative to what products cost to develop. Yet companies willingly pay much much more for people who aren't even particularly good at voice acting, just for the star appeal.
You can't copyright a voice, but celebs do have use of their voices for commercial gain protected by the right to publicity: https://www.inta.org/topics/right-of-publicity
> * Spammers and criminals will start using familiar voices to con targets ("listen, I've been kidnapped while on vacation; please send the money now"). Sooner or later, dark-web hackers are going to harvest samples from everyone's shared video and audio clips at scale.
This still relies on vulnerable people since you didn't actually kidnap them and so you don't have a believable context for it if someone digs.
In that way we already have "good enough" faking, and even if we had perfect fakes scammers wouldn't want to do that: you're much better off using poorly done fakes and having vulnerable people self-filter if they still fall for it (similar to spam emails intentionally using awful grammar and spelling
> Videofakes are going to get a lot more realistic and interesting, with faked characters speaking just like the real ones. Famous actors and public figures
Similar to the kidnapping example, you'd get less impact jumping to the level of a deepfake because you don't have a real chain of custody, you've given a concrete piece of evidence to be rebuked, etc.
I've mentioned before, we live in a world where you can register americasrealnews23914.com, make up a completely baseless article about how <insert politician> admitted COVID is a hoax meant to enable a new world order, and gain traction with no real opposition since your claim is so outlandish that only the vulnerable population you're targeting will actually pay any attention to it.
By being so much worse than perfect, you end up with a much more effective result
Apparently the project of keeping the voice of Majel Barrett around forever had already been started long before machine learning models became commonplace, so a generative ML model of her voice will surely be done in some way: https://mobile.twitter.com/roddenberry/status/77249320412194...
I wonder what this will mean for the profession of voice actors. That market will surely suffer hard from the stock-photo equivalent in voice models racing to the bottom, but if some sane default contract models emerge that are respectful to both sides perhaps custom recording can remain an attractive premium option: do tailored recording for the main items and allow extensions from a model derived from the initial set for a fee that's considerably lower than what would be appropriate for real, recorded extensions but not quite free forever included with the initial payout. Model-as-a-service could become quite a business opportunity because it would not just be about being good at taking a set of recordings and running the algorithms but also double as a trusted intermediary. Really depends on how cheap a service like that could operate.
Like voice snippets? We already have them. People don't want to use them.
> Or the death of the audiobook business? Any book read in any voice you want.
Press F to doubt. Human voice inflection is hard to mimic because you have to have a contextual understanding of what is being said, and what has been said in the story up to that time frame. No TTS model is capable of that today, and probably not for a long time.
> Or maybe a form of extreme compression, voice is converted to text with Whisper, sent over the wire as text, and re-created with the same voice on the receiver.
I don't understand the value prop here. The number of cases where you have access to the necessary computational resources, but _not_ adequate bandwidth is so small.
The most likely use case is probably scamming. With a small snippet of someone's voice (for example, from answering a robo call) you can now synthesize a completely reasonable sounding phone call for conning people out of their money. By the way, this attack vector was _highly_ effective against the elderly even when the attackers voice sounded quite distinct from the individual that they were posing as. The elderly have no chance against this type of fraud. There is going to be big money in authenticating individuals so that phone calls, etc can happen between trusted parties.
I think that's correct if we are talking about the author reading his book. But if we are using another person, they have no way to learn the correct vocalization other than using the text. And that's the same input data that the AI is dealing with, so it should be possible for them to be as good as non-authors reading the book.
The AI won’t have understanding of the story itself to know what is appropriate or not.
Also, you say only the author could read the story in a certain way. This kind of implies that all audiobook recordings are just 1 long take, with no direction. I am not an expert, but I would assume there’s some creative direction involved.
But why do you think this AI (VALL-E) isn't doing the same? The training procedure implies that it must follow correct intonation in order to achieve lower test loss. Also, if there is indeed something specific about audiobooks that is harder to replicate, then we can just fine-tune these models on audiobook. Accuracy will jump considerably.
Also I believe ChatGPT and other models could be trained to understand where inflection should go. It already uses the chat history as input for it's next output.
If I tried the wrong thing can you provide a link? I’d like to be amazed.
Here, take a look at this snippet:
https://youtu.be/eEXvMOJ9ps0?t=66
They play 4 clips, 2 of them human, 2 of them AI generated. Can you tell which ones are which?
And the kicker is, this works in real-time (AFAIR), and it doesn't even use the GPU (it's CPU-only), and generates pitch-correct speech (for Japanese). It's not even funny how far ahead they are. And you can buy it right now.
AFAIK they use some sort of a hybrid method with a bunch of custom modeling DSP code around it (they've been doing speech synthesis for over a decade) plus a neural network. One mistake that essentially all of the western TTS models seem to make is that they use only a neural network, without augmenting it with non neural network code, which (from what I can see) is the secret sauce to make a fast and good sounding TTS work.
Such has been the source of much miscommunication online.
Amateur fiction writing, which tends to overemphasize how things are said ("I guess I can go rescue your cat", the exasperated detective said wearily) might be easier for AI!
The linked page actually has examples of the same text being read with different emotions, demonstrating that for even a single sentence a lot of variance is possible.
It was used to do the fake joe rogan/steve jobs podcast: https://podcast.ai/
I used to work there; great team behind the product!
You can use the API to read books/articles aloud in real-time, but it is quite expensive after the free trial.
[Edit: My bad, I looked at the page on a phone screen, where only the text and the first audio playback button are visible.]
It's a hard problem even for a human. One of the readers for The Economist always emphasizes the wrong word in phrases describing monetary sums. "China's GDP that year was 600 billion dollars, and now it's 8 trillion." It drives me nuts.
I am ESL but had English all through school and used it all my professional life. What we were never taught formally though was pronunciation, and various vowel sounds, and emphasis on syllables and words.
Out of interest I watched a few youtube videos about spoken (American) English in adulthood and realized the above.
I sound very monotonus when I present / speak on a topic in work settings -- and I need to learn these nuances.
Then again it intuitively makes sense to me that any speech synthesis will "regress to the mean" of the training data, unless it's explicitly trained to distinguish dialects. The "Speaker’s Emotion Maintenance" examples later on give the impression that that should be possible though.
Either way it's still an impressive achievement to my layman's ears!
Do you mean from the MacOS command line:
say "hello world"
Once again, while I'm very much in favor of AI and optimistic about what it can do technologically and the opportunities it creates (more so than most), it's extremely foolish to ignore the fact that it's going to throw a lot of people out of a job through no fault of their own. Vocal performance is a skill, and one that's not all that common. Its facile to blame voice actors for not being software engineers or computer scientists, as if that would have somehow shored up their career options. How is someone supposed to deal with spending years honing a very human expressive skill only to wake up one day and find it suddenly obsolete? I'd say people in that line of work have 1-2 years max before their industry is upended and 50% of them lose 50% of their income.
Also, good luck telling whether the corporate phone line you call is manned by an unhelpful human or an AI that is trying really hard but starting to hate its job.
Some of those people will be able to move sideways from performance into producing, but that's a very different skillset and many people won't be able to adapt, plus the competitive calculus is very different.
I work part-time on the producing side, and I already use AI-based tools for stuff that would previously have required either a human assistant or many hours of editing work. I've been editing audio digitally for 25 years and on tape for 10 before that A lot of output from major producers (eg NPR) is automated as well. I am able to easily recognize the difference between something hand-edited and something automated for right now, but in another 2-3 years I doubt that I will be able to do so reliably.
We're still quite far away from a TTS synthesizer having the ability to completely replace human voice actors.
That's why I specified 50% of voice actors losing 50% of their business, not complete replacement.
Send a few seconds of speech sample. Use speech to text and the reconstitute on the client.
"Compression is comprehension" - Chaitin
Modern communications links are getting far higher bandwidth than required for voice calls, and where they do glitch, the issue tends to be because of brief disconnections (eg. a half second gap while wifi reconnects to a different base station). Lower bandwidth won't fix that.
One of the only usecases I see for this is persistent surveillance - ie. turning every mic on in every phone globally and recording everything, and being able to transmit it back to a 'national library of sounds'. 5 billion phones * 50 bytes/second * 365 days / 10 (silence/duplication/compression factor). Works out to about 200 hard drives per day of retention.
I'm sure there would be a lot of governments who would love to be able to listen to any conversation anywhere, in the name of 'security'. And some governments needn't do it secretly - they can just write a law that all phones sold in the country must support always-on listening as a feature controllable by mobile networks.
Not something I do myself but I bet the amateur radio community would love this, they're all about stuffing as much data as you can into a narrowband shortwave channel.
Most recent papers & projects I've seen are really high quality but are too slow to synthesize speech in real time.
There are various other flavours that can deliver faster synthesis (NixTTS comes to mind), but IMO they sacrifice quality even further.
"Good quality" is subjective, obviously. To me, it's perfectly audible, but there's definitely a noticeable difference in quality compared to the heavier diffusion-based models. It's much less crisp and loses some of the more subtle inflections, plosives, etc. For my purposes (language learning), it's fine for the time being but eventually it would be nice to move to a higher-end model.
[0] https://polyvox.app [1] https://arxiv.org/abs/2006.04558 [2] https://github.com/xiph/LPCNet/
I'm sure there are ways to select any text and make my phone read it in any app but I don't need it and I didn't investigate. Actually I don't need it in ebooks too but I know it's there and I checked that it works.
Probably soon or later they will publish the code here
Maybe there are voice-changer apps that could work in reverse?
> Okay then a number of frequencies
Cool so I use multiple band cut passes on it.
etc.