Show HN: Clone your voice and speak a foreign language
coqui.ai
coqui.ai
The fact that it can do multi-lingual voice cloning at all in that case is already surprising. You can find more details in the project page [0] and paper [1]. And here's the corpus. [2]
[0] https://edresson.github.io/YourTTS/
I also did Pt -> En and sounded like... me speaking English, though with some artifacts. VERY cool.
The idea of having a model of my voice out there that can say whatever is written in a text box is scary.
Perhaps a solution is a sound fingerprint requirement for voice imitation software so that it's easily identifiable in court if it's an imitation voice.
It's somewhat of a new frontier, imagine during a divorce proceeding your ex-partner fabricates voice recordings of you threatening the kids so you don't get custody, how do you protect yourself against that, how to you prove that's what happened? Soon enough it'll just be an app on their phone that they use to record your voice during a discussion, then later spits out a sound file of you saying whatever they want you to say. That's clearly a socially dangerous tool.
Fake calls to your relatives in your voice or even fake video with your face and voice asking for money! or illegal activities.
Few years later a company will come and say we can detect if it's fake or not pay $10,000 for solution, or get ready to be in prison. Oops! legal system doesn't accept this as a proof, now what? Welcome to the prison.
Both companies are making money, and you are paying by money and your life.
I can see the government banning using voice as a password. I can't see it banning the tech. The criminals will use the tech regardless of if it's banned. Looks like we'll need person to person authentication for our relatives soon.
If technology like this is plausible, then the recording shouldn't be considered a statement by me in the first place.
People are just going to have to learn not to trust audio. People adapted to photoshop, they'll adapt to this.
It's a valid concern to not want to give a random website a workable voice model. Just because you've talked on the phone or used speech-to-text before, doesn't make that concern invalid.
Probably the folks who could best use this nefariously are the folks we already know, who have much greater availability to our voice. Those folks are in the best situation to capitalize on a working voice model to, say, call our manager, bank branch, or local emergency operator. A random website would have to go to some effort to accumulate the needed information to use our voice for much, whereas someone who already knows us could have us fired for the contents of a phone call to the manager, up on charges for prank-calling 911, or worse.
At least there, it was optional.
The English->Portuguese sample sounded nothing like me at all except for one syllable where it sounded like it was playing back a brief snip of what I had recorded.
The English->French version did a little bit better, it sounded like the voice had been influenced by mine in some small way.
English->English (saying a very different sentence to what I recorded) was pretty impressive though.
The way it works is, a model is trained of all possible voices. Then your specific voice is projected into latent space.
That's why it can mimic your voice with only a few seconds of audio. It's not making a model, but rather using an existing model.
It may seem like a pedantic distinction, but it's why the model isn't as worrisome as it seems. It can't target you specifically, just the average voice near yours.
It's closer to a really talented parrot than a model that can impersonate you on command. I suspect if you try it out, you'll be surprised it's so far off from your actual voice.
English -> French seemed to work best, with the AI output have a very similar timbre to my real voice. Not hyperrealistic for me, but decent enough given I gave it a ~20s sample.
French -> English was less good in terms of the timbre and pitch of the voice---way higher than my real voice. It did have a bit of a Canadian accent, though, which is funny because I speak French with a Quebec accent. Maybe that's what I would sound like if I had a Canadian accent in English?
In the demo we specifically disallowed bulk uploads to hinder such abuses.
It's also a new possibility to somewhat personalize the text to speech engines. The above example is not really close to my voice.
https://fakeyou.com/tts/result/TR:eyfam30e255zxy69vn6a7z7yn9...
Background music makes misuse/abuse less likely (both intentional and unintentional)
Read more here about in our open discussion: https://github.com/coqui-ai/TTS/discussions/1036
Maybe if we can get watermarked stuff out first and the average person gets up to speed with what tech can do, we can all adjust our expectations before the real wave of abuse hits.
It's very hard to curb intentional misuse.
There is no reason to blame the creators, this is going to go mainstream one way or another.
That will help ensure that this is only being used by the person visiting the web page.
(That will only help with the hosted version, of course, not if you make the model code/weights available. I didn't generate this idea myself but also can't remember where I saw it. I think it was from someone offering a similar service.)
In fact, just a year after this post was written, CoquiAI started their open source projects [1].
[0] https://news.ycombinator.com/item?id=22869365 (https://thegradient.pub/towards-an-imagenet-moment-for-speec...)
I suppose that if I ever take proper English pronunciation classes, I now know what to strive for.
Btw idea is really cool, its like how will you speak in same tone in other languages.
If someone wants to fake a statement there are already 100 ways to do it. Not making their servers the ones doing the deed puts a meaningful barrier in place for more casual misuse. And for serious cases like impersonation on a large scale, the resources are there to likely do better than this instant feedback model can.
How do y'all intend to profit (succeed as a startup) if you're releasing so much publicly? I'd love to see you guys succeed.
Really great to see where some of the Mozilla TTS folks wound up, too.
The input sentence generally should be in the language you selected from the dropdown. For example, if the dropdown has "French" selected you could enter the text "Allons enfants de la Patrie, Le jour de gloire est arrivé!"
Clicking "Submit" then generates a TTS reading of the sentence you input in the language selected from the dropdown.
For fun you can mix and match. In other words, select a language from the drop down and enter text in the text box not in the language selected from the dropdown. (For example, the dropdown could have "French" selected and the sentence could be "O say can you see, by the dawn's early light". This gives interesting results, it sounds as if a native French speaker is speaking English.)
How do I embed this?