Show HN: Neural text to speech with dozens of celebrity voices
vocodes.com
vocodes.com
It has celebrities like Sir David Attenborough and Arnold Schwarzenegger, a bunch of the presidents, and also some engineers: PG, Sam Altman, Peter Thiel, Mark Zuckerberg
I'm not far away from a working "real time" [1] voice conversion (VC) system. This turns a source voice into a target voice. The most difficult part is getting it to generalize to new, unheard speakers. I haven't recorded my progress recently, but here are some old rudimentary results that make my voice sound slightly like Trump [2]. If you know what my voice sounds like and you kind of squint at it a little, the results are pretty neat. I'll try to publish newer stuff soon, and that all sounds much better.
I was just about to submit all of this to HN (on "new").
Edit: well, my post [3] didn't make it (it fell to the second page of new). But I'll be happy to answer questions here.
[1] It has about ~1500ms of lag, but I think it can be improved.
[2] https://drive.google.com/file/d/1vgnq09YjX6pYwf4ubFYHukDafxP...
[3] I'm only linking this because it failed to reach popularity. https://news.ycombinator.com/item?id=23965787
We'll re-up that thread (see https://news.ycombinator.com/item?id=11662380 for how this works generally). I'm going to move this comment there as well because it includes more background info than you posted there.
Tangentially related, if you have your voice print as a security mechanism at a financial institution (Vanguard), you should ask them to turn that off.
As I said, they lost on copyright claims but these cases involved their natural voice. If the voice was a made up voice, such as a character an animated character, I wouldn't immediately dismiss the idea that it might be copyrightable.
e.g. I tried to do an Alyx version of https://www.youtube.com/watch?v=koU3L7WBz_s but it came out sounding nothing like her.
I've been wondering about the possibility of using this sort of tech (or the API offerings from Azure or GCP) to provide voice overs in video games.
By that I mean for smaller budget Indie development, it would be certainly interesting to either be able to generate voice audio from transcripts in order to add voices to background NPCs and so on (or even the possibility of doing it at run time to produce much more dynamic worlds).
I guess the biggest blocker is the difficulty in conveying emotion with what is currently available as well as the difficulty in getting pronunciation correct (especially with nouns).
These companies tend to focus on off-the-shelf turnkey solutions, so they'll have a suite of a few voice actors to choose from for different character archetypes.
E.g. training off Schwarzenegger and offering an Arnold transform
I believe (I'm not certain) that celebrity voice impersonation is legal as long as it is not used to sell or endorse a product.
Most models are trained on the original speaker's voice, but maybe only a little bit. Models might incorporate learning from many speakers. We might even be able to boil down a speaker representation to a small vector encoding in the future. It'll be interesting if we can capture the representation of a person with just a few numbers.
I don't think the legislature should be overly protective against machine learning. It seems obvious to me that neural networks will play a huge role in creating entirely virtual musicians and influencers. We're already seeing this start to happen. r9y9 on github has published some models that rival Vocaloid in lyrical ability.
At the same time, we don't want these techniques used to commit fraud, slander, or have them be used to falsely accuse someone of committing some act. These are things we might need new legal protections for.
But I don't know what I'm talking about. I'm not a lawyer.
It's essentially the performance of a composition vs the composition question again: at what point am I mimicking someone to the extent they have a valid claim on a portion of my work?
I expect it'll enter the courts a few milliseconds after someone clones a dead actor (without their estate's permission) for a new performance.
There's always been an inherent tension in the US distinction between a law of nature and a creative work though. It seems a bit silly for me to claim patent / trademark on a vector that encodes my likeness.
For example, lots of people sound like Arnold Schwarzenegger. So if you trained a model with tall, deep-voiced Austrian man, you could probably get something that people will immediately associate with Arnold without actually being his voice, or someone emulating him. Because much of what Americans associate with his voice is really a regional accent which is relatively uncommon in the US.
There may be a little bit more difficulty getting away with with someone like Gilbert Gottfried, whose voice is much more unique. But I do think you could get away with creating a voice that people think sounds just like him, but doesn't hold up in a side-by-side comparison.
What I think will happen is celebrities like Morgan Freeman will use their voice to train models like this, then gift these to their estates for use in the future.
I think "passing off", an unregistered element of trademark laws, may be pertinent here. If the public think that there's an association and you're knowingly trading on that, even if the public are wrong, then you can be 'passing off' your output as someone else's goods/services/[vocal renditions].
It's likely you'd have to be very careful about use of copyright material for training the voice (eg extracting metrics that describe the voice). Fair Use might apply in USA though (even commercially).
IANAL, this is not legal advice.
Really cool that you got this to work. I used to work on TTS (a few years ago, now), and we trained on celebrity voices, but used full audiobooks. https://github.com/Kyubyong/tacotron
Here are some of our Nick Offerman samples: https://soundcloud.com/kyubyong-park/sets/tacotron_nick_215k .
Thanks for making this so open and accessible.
I saw once a company that offered to be the sole purveyor of a celebrity's synthesized voice. I haven't been able to find them again, but that seems like a much safer way to monetize this.
I'm not sure that making it easier to profit off of the likeness of others is a positive side. If it's legal for indie studios to do, it's legal for 20th Century Fox, Universal, and so forth.
This is purposefully not counting in the effect of being able to fake people and the damage that does to society, but I think that was implied by the previous poster specifying looking for the more positive side of the technology.
> I'm going to steal your soul. One injection at a time. Slowly, over the course of the next decade, the entire essence of your being will be demolished until your body is nothing but a vessel for my command.
Great work though!
I am worried about the potential abuse of this service, are there any existing services that can help to identify audio deep fakes like this one is for making them?
Found Resemblyzer: https://github.com/resemble-ai/Resemblyzer
On a more positive note, when deepfakes become a problem, we will see the emergence of a culture where unsigned authoritative content is not paid any attention.
If current events are any indication, that culture will only emerge 30 years after the tech becomes widely usable, and in the interim will lead to absolute chaos in the form of weaponized disinformation.
Lots of bad things happen, and they are only surfaced because the person in question didn't notice the surreptitious recording. When deep fakes becomes a problem, it will give these people plausible deniability and they can just reject it as "fake news."
For example, https://en.wikipedia.org/wiki/Censorship_of_images_in_the_So...
Also photographer friend of mine said; a great photographer doesn't need to Photoshop anything to lie to you.
I couldn't agree more if they used the word lie in the more general sense as a synonym of deception. The availability of fakes may not become a problem because the most effective deception doesn't involve telling untruths.
To give an example, suppose Russia Today and Fox News report on the same event. There's a set of facts. RT picks a subset and reports it from their point of view. Fox picks another subset and presents their view. The resulting articles may give readers vastly different interpretation of the event and no untruths had to be involved.
Don’t let perfect be the enemy of good. This has potential to literally cause spilled blood, fraud, etc. Better to have it for some than for zero.
The cat is out of the bag. Digital media should not be trusted blindly.
Governments have a history of putting up solutions that work in their favor by selective enforcement and securing power to self. Encryption debate is one such example.
An interesting quirk: some words seem to get dropped entirely? for example the word "cleverer" and any word with a hyphen.
Really fun to start with a quote from one person and switch between voices to hear others recite the same line. Alan Rickman doing lines from Aladdin as Iago is pretty funny
Tip: It's a cool idea to put some ready made samples under the photos. A lot of people like myself only want to hear some demos and pre saved mp3 samples are more than sufficient for that sort of thing. It will also help reduce your server loads.
The problem is that currently your training data has to be annotated with these tokens, and that adds a lot to the difficulty of creating data sets.
I imagine that over time this will get much easier to do.
Do you have a GitHub or technical documentation about how you build this sort of thing to work at scale?
A rust TTS server hosts two models: a mel inference model and a mel inversion model. The ones I'm using are glow-tts and melgan. They fit together back to back in a pipeline.
I chose these models not for their fidelity, but for their performance. They're 10x faster at inference than Tacotron 2. If you want something that sounds amazing, you're better off with a denser set of networks, like Tacotron 2 + WaveGlow. You should use these for achieving superior offline results for multimedia purposes.
Instead of using graphemes, I'm using ARPABET phonemes, and I get these from a lookup table called "CMUdict" from Carnegie Mellon. In the future I'll supplement this with a model that predicts phonemes for missing entries.
Each TTS server only hosts one or two voices due to memory constraints. These models are huge. This fleet is scaled horizontally. A proxy server sits in front and decodes the request and directs it to the appropriate backend based on a ConfigMap that associates a service with the underlying model. Kubernetes is used to wire all of this up.
I scaled for today, but it's pretty cheap to run day to day.
I also have some architectural optimizations to make that will greatly reduce the costs. Right now, nodes are responsible for two speakers apiece. This is an under-utilization since most speakers don't get used.
I ask because I help maintain an open source ML infra project ( https://github.com/cortexlabs/cortex ) and we've recently done a lot of work around autoscaling multi-model endpoints. Always curious to see how others are approaching this.
total 4.2G
-rw-r--r-- 1 bt bt 110M glow-tts_alan-rickman_ljstx_2020.07.22_expr-1_chkpt-4765.torchjit
-rw-r--r-- 1 bt bt 110M glow-tts_anderson_cooper_ljstx_2020.07.21_expr-1_chkpt-6622.torchjit
-rw-r--r-- 1 bt bt 110M glow-tts_arnold_schwarzenegger_ljstx_2020.07.16_expr-2_chkpt-9045.torchjit
-rw-r--r-- 1 bt bt 110M glow-tts_barack_obama_ljstx_2020.06.28_expr-1_chkpt-1729.torchjit
-rw-r--r-- 1 bt bt 110M glow-tts_ben-stein_ljstx_2020.07.21_expr-1_chkpt-7516.torchjit
-rw-r--r-- 1 bt bt 110M glow-tts_betty_white_ljstx_2020.06.28_expr-1_chkpt-1666.torchjit
...
melgan: -rw-r--r-- 1 bt bt 17M melgan_manyvoice5.0_2020-07-23_12d5838_10760.torchjit
(All the voices use the same melgan, or derivations of it.)I'll edit my post later with my deployment and cluster architecture. In short, it's sharded and proxied from a thin microservice at the top of the stack. I'll probably introduce a job queue soon.
Is this why some examples I tried seemed to skip some of the words?
There aren't entries for
- asdhfjahdsff
- rawr
I added around 500 new words, but I missed a lot of stuff.
The ultimate fix is to have grapheme -> phoneme prediction so that all unseen words can be mapped to potential phonemes (polyphones).
Thanks for sharing, though. Very interesting project!
A detailed blog post about this would be amazing! I wish there was a hn bot like Reddit bots to ping me when you do post it so i don't miss it.
There are a lot of neat research threads ongoing in terms of generating vocals.
Nvidia published Mellotron (code + paper + models), and the results are promising:
https://github.com/NVIDIA/mellotron
https://nv-adlr.github.io/Mellotron
The best results I've seen are from researcher Ryuichi Yamamoto (r9y9 on Github). He continually publishes astonishing results and novel architectures:
https://soundcloud.com/r9y9/sets/dnn-based-singing-voice
These results lead me to believe he's going to have a replacement for Vocaloid soon.
There's lots more stuff out there, and I can come back and edit my post later.
Some folks are getting good results by simply combining Tacotron with autotune:
- https://www.youtube.com/watch?v=3qR8I5zlMHs Mister Rogers sings Beautiful World (amazing, super charming, and shows the promise of this tech)
- https://www.youtube.com/watch?v=K1jrDgbRs9Q (Tupac, possibly NSFW lyrics)
- https://www.youtube.com/watch?v=QW16_W0K3qU (Tupac with various results, possibly NSFW)
There's a lot that gets posted to /r/VocalSynthesis and occasionally /r/MediaSynthesis
"My name is Bill, the lord of computers. I love computers, and they love me too. I'll give you a computer, maybe one, maybe two. If you are lucky it might not even crash on you. Love your computer, like your daughter or your wife, treat it with kindness, and it will reward you for life! I am bill the god of computers. Bow to me now or I will be sod you."
I am a writer and found that the best editing comes when I am reviewing audio files of my books from voice talent. Of course, then it is way to late to change anything. With a tool like this I can revise as much as I want!
The hardest part of this is in dataset creation. It's hard to clean and annotate the data and can be quite manual. That's why companies with lots of data will win.
There are automated techniques to help with segmentation, bandpass filtering, transcriptions, etc., but they're far from perfect.
Not legal advice, of course.
Can you give some more info on how you generated the models? I'm also interested in the tech stack you're using to implement this webapp... Would love some details!
..What's next?
Text To Video webapp that renders text to video + voice synchronised of famous people.
Who wouldn't like to laugh 5X more when social scrolling?
The first platform that enables creators with the ability to produce deep fakes of celebs from text that they can broadcast as HQ video content to their audience will kill both Youtube & Instagram.
Ranking based on likes so the best jokes of the day are trending on top of the feed.
Recommendation engine with a multibandid ML algo from the start so you can leverage all that incoming data.
glow-tts and melgan, which are somewhat unpopular choices given the proliferation of Tacotron2/Waveglow. I chose these due to their sparsity and speed.
> I'm also interested in the tech stack you're using to implement this webapp... Would love some details!
It's a Rust microservice architecture. There's a proxy layer that decodes the request and sends it to the appropriate backend, and then there's the tts service that is horizontally scaled and is responsible for loading the model pipeline and turning requests into audio.
> ..What's next?
For me? Voice conversion in the near term. This takes microphone input and turns it into the target speaker's voice.
I'm also spending a lot of time on photogrammetry. I have a 3d volumetric webcam system right now that I have much bigger plans for.
The implications for security are huge. If your friend calls you up for a very quick chat from an unknown number and asks you to remind them of your address, are you going to ask for authentication to prove it's really them and not a convincing synthetic voice?
Subreddit simulator is pretty convincing conversations, putting that to high quality voices? mannnn, so many good applications.
Speaking of which, why don't people just talk about the good applications. You'll get ostracized for speculating more bad things about COVID, but talk about how doomed we potentially are with deep fakes? Give that blogger a pulitzer prize!
Maybe, maybe not. You'll see some of the model sizes I posted in comments above. These are quite large, and adding models for multiple speakers gets quite large. These have to live in memory and probably can't be paged in selectively.
Once we achieve high fidelity multi-speaker embedding models (where multiple speakers are encoded in a singular model), then we'll have something compelling. I imagine the models will become less dense over time as well.
Furthermore, if the models are deterministic, then the designers will know what each line will sound like exactly before it's produced.
> There was an error and I still haven't implemented retry to make it invisible. You can absolutely submit your request again a few times; this is a self-healing Kubernetes cluster. Some models (voices) get more load than others and/or are scaled to fewer or more pods. There's also a rate limiter, but there aren't error messages yet.
I know there's legal (and perhaps ethical?) issues to work out, but I really wish tech like this, if fine-tuned, could be used to resurrect stuff like Jim Henson's original Kermit voice; the Muppets' new voices all sound horrible. I'd love for fictional character voices to become immortal.
Ask it to say " The Aristoocrats!" with Gilbert Goddfried.
I don't have generic grapheme -> phoneme/polyphone prediction, but that's something I look to add soon. In my literature review I didn't see anything in this space, so I was thinking I might have to come up with something novel.
Current cloud based solutions from AWS/Google Cloud/Azure are pretty expensive.
Using Chrome stable
Thanks for the help and info!
I'm using version 84.0.4147.89 (Official Build) (64-bit) and getting back responses.
I got the following response headers:
access-control-allow-origin: https://vo.codes
content-length: 151689
content-type: application/json
date: Mon, 27 Jul 2020 15:55:37 GMT
vary: Origin
x-backend-hostname: tts-group-1-965d444f5-7kvkm
I'll try to dump the cache and reproduce.edit: I must have an old browser. It works everywhere I'm testing it. CORS is hard. :(
I switched to Safari and Disabled CORS, but a 500 error is coming back now. So maybe the 500 response is the root cause, and the error handler is not returning CORS headers, masking the issue on Chrome.
Edit: by putting in a shorter input (sentence rather than paragraph) I was able to get a response.
What might've happened is that the instance your request was farmed out to might have been OOM killed. I've provided lots of memory, but these models are pretty massive and each inference run has to spin up a lot of matrices in memory.
This is all CPU inference, not GPU.
When the pods get OOM killed, they spin up again. The clusters for each speaker are about 5-10 pods apiece (with some double tenancy).
Do we have more data on male speech than female speech?
Cate Blanchett, Sarah Silverman, Katey Sagal, Jennifer Tilly, Laura Prepon, Viola Davis, Judi Dench, Whoopi Goldberg, Julie Andrews, Lake Bell, Jane Lynch, Joan Rivers, Martha Stewart, Katharine Hepburn, Sarah Vowell, Shoreh Aghdashloo...
I then transfer learned for each of these speakers. Some speakers have as little as 40 minutes of data, others have up to five hours. The resulting quality isn't strictly a function of the amount of training data, though more typically helps. It's also important to have high fidelity text transcriptions free of errors.
The transfer learning runs vary between six hours and thirty six hours.
I'm using 8xV100 instances to train glow-tts and 2x1080Ti to train melgan. I'm continuously training melgan in the background and simply adding more training data. The same model works for all speakers.
My reasoning for this approach: IMO, if the model learns a "universal human voice", it shouldn't need too much additional information to get a target voice.
I think you're right in that if we can get such a model to work, training new embeddings won't require much data.
If he thinks this is egregious, I'll take it down.
As an outlier not running javascript, I'm reaping what I sow, but it would be nice to me and others in the same boat if projects make their landing page viewable without the need for javascript.
Here's a request:
curl 'https://mumble.stream/speak' --compressed -H 'Referer: https://vo.codes/' -H 'Content-Type: application/json' -H 'Origin: https://vo.codes' -H 'Connection: keep-alive' --data-raw '{"text":"testing 12345","speaker":"david-attenborough"}' --output output.wav
The other speaker values:
Surely you can enable or use a browser with JavaScript when you choose to?