OpenVoice: Versatile Instant Voice Cloning
arxiv.org
arxiv.org
Audio sample: https://storage.googleapis.com/dalle-party/sample.mp3
Cloned voice (converted to mp3): https://storage.googleapis.com/dalle-party/output_en_default...
All I did was install the packages with pip and then run "demo_part1.ipynb" with my audio sample plugged in. Ran almost instantly on my laptop 3070 Ti / 8GB. (Also, I admit to not reading the paper, I just ran the code)
From the README
Disclaimer
This is an open-source implementation that approximates the performance of the internal voice clone technology of myshell.ai. The online version in myshell.ai has better 1) audio quality, 2) voice cloning similarity, 3) speech naturalness and 4) computational efficiency.It was terrible. Absolutely terrible. Like, given how much hype I saw about this, I expected something half decent. It was not. It was bad, so bad bad bad.
I was thinking maybe I did something wrong, but then I watched some of the youtube reviews - these guys were SO excited at the start of the video and then at the end, they all literally said, "Uh, well, you be the judge"
I still can't help but feel there's some kind of trick to it - get the right input sample, done in the right intonation, and maybe you can generate anything
I have a friend with a paralysed larynx who is often using his phone or a small laptop to type in order to communicate. I know he would love it if it was possible to take old recordings of him speaking and use that to give him back "his" voice, at least in some small measure.
I recently saw a video from google about a custom voice they created for somebody with ALS but I can't seem to find it online (Does anybody have a link?). Creating custom voices is not yet available on Android though. The latest iOS release (iOS 17) does support creating personalized voices.
ModelTalker [3] is a long-term (research?) project to create custom voices for people with speech disabilities. Their TTS seem to support Android so that might be another option.
[0] https://www.acapela-group.com/ [1] https://www.speakunique.co.uk/ [2] https://vocalid.ai/ [3] https://www.modeltalker.org/
Unfortunately my friend's voice is too far gone for that to be possible. Hoping for something where they can use old recordings to generate a voice.
Right now Eleven Labs is your best bet.
xTTS is just not there quality wise. The version available in the studio is marginally better than the OSS version but it's still pretty far from being believable.
The non-nerfed version of Tortoise (the author decided to ruin their own project but forks exist) was decent at voice cloning but it takes a lot of tries.
I'm pretty sure we already have the technology to do what you want and help your friend, it's just a matter of time until it gets better and more software comes out.
iOS has this built in :/ which may bode well, there's no greater Google product manager than "whatever apple just shipped."
I'm doing some xplatform on device inference stuff (see FONNX on GitHub) and it'll be one of 100 items that'll stick on my mind for a while, I hope I find time and I'll try to ping you
Edit: is an Android app with a keyboard and "speak" button that does API calls to eleven labs sufficient for something worth trying?
Thanks. Use the email address in my profile if anything eventuates.
> is an Android app with a keyboard and "speak" button that does API calls to eleven labs sufficient for something worth trying?
Maybe. Obviously something with local processing would be preferred, but it might be an option when internet connectivity is good. Is there such an app?
(warning: detailed opinionated take, I suggest skimming)
Why? Local inference is hard. You need two things: the clips to voice model (which we have here, but bleeding edge), and text + voice -> speech model.
Text to voice to speech, locally, has excellent prior art for me, in the form of a Raspberry Pi-based ONNX inference library called [Piper](https://github.com/rhasspy/piper). I should just be able to copy that, about an afternoon of work! :P
Except...when these models are trained, they encode plaintext to model input using a library called eSpeak.
eSpeak is basically f(plaintext) => ints representing phonemes.
eSpeak is a C library and written in a style I haven't seen in a while and depends on other C libraries. So I end up needing to port like 20K lines of C to Dart...or I could use WASM, but over the last year, I lost the ability to be able to reason through how to get WASM running in Dart, both native and web.
Re: ElevenLabs
I had looked into the API months ago and vaguely remembered it was _very_ complete.
I spent the last hour or two playing with it, and reconfirmed that. They have enough API surface that you could build an app that took voice recordings, created a voice, and then did POSTs / socket connection to get audio data from that voice at will.
Only issue is pricing IMHO, $0.18 for 1000 characters. :/ But this is something I feel very comfortable saying wouldn't be _that_ much work to build and open source with a "bring your own API key" type thing.
I had forgotten about Eleven Labs till your post, which made me realize there was an actually meaningful and quite moving use case for it. All of Elevens advantages (cloning, peak quality by a mile) come into play, and the disadvantages are blunted: local voice cloning isn't there yet, and $0.18 / 1000 characters doesn't matter as much when it's interpersonal exchanges instead of long AI responses
though, now that I write that...
Native: FFI.
Web: Dart calling simple JS function, and the JS handles WASM.
...is an excellent sweet spot. Matches exactly what I do with FONNX. The trouble with WASM is Dart-bounded.
(n.b. re: local cloning for anyone this deep, this would allow local inference of the existing voices in the Raspberry Pi x ONNX voicer project above. It won't _necessarily_ help with doing voice cloning locally, you'll need to prove out that you can get a voice cloning model in ONNX to confirm.)
(n.b. re: translating to Dart, I think the only advantage of a pure Dart port would be memory safety stuff but I also don't think a pure Dart port is feasible without O(months) of time. The C is...very very very 2000s C. globals in one file representing current state that 3 other files need to access. array of structs formed by just reading bytes from a file at runtime that matches the struct layout)
(Checkpoint link defanged because I’m allergic to direct links to zip files hosted on Amazon. Nor have I reviewed what the file contains.)
As for the checkpoint, I'm not allergic and I don't do security theater:
https://github.com/myshell-ai/OpenVoice?tab=readme-ov-file#i... links to
https://myshell-public-repo-hosting.s3.amazonaws.com/checkpo...
Your comment comes off as passive aggressive.
I’d suggest that those who think it is theatre probably don’t understand the implications of that action.
Simply downloading a zip from Amazon has zero risk. Even opening an arbitrary zip has essentially zero risk. RCE from opening a zip is obviously a really critical and valuable vulnerability and would not be wasted with a public link.
Combine that with the fact that this comes from a voice cloning GitHub repo and the chance of this having some 0-day are infinitesimal.
Finally just making the link non-clickable does not add security. Nobody can take any action to increase their security because they have to slightly edit a link (not that they would because it's sensible a clickable link in the GitHub readme).
So yes, I fully understand the implications and it is definitely security theatre.
I suggest that those who think that it isn't probably haven't really thought about the threat model.
But in the spirit of hacker news, I'll continue the argument.
> There are no implications.
Untrue and absolutist.
> Simply downloading a zip from Amazon has zero risk.
Agreed.
> Even opening an arbitrary zip has essentially zero risk. RCE from opening a zip is obviously a really critical and valuable vulnerability and would not be wasted with a public link.
Broadly agreed. History is full of unzip vulns, but I agree that this doesn't seem likely. Much easier to persuade folks to deliberately run their malware by using the latest fad as a hook. I'm not claiming that happened here.
> Combine that with the fact that this comes from a voice cloning GitHub repo and the chance of this having some 0-day are infinitesimal.
Maybe you know these authors and this repo and trust them. I don't. I'm sure they are lovely, I have no idea, I've done no research, and I've never heard of them before. That being said, if I wanted to distribute a backdoor or cryptominer to a bunch of people with powerful computers, I'd definitely hop on the AI bandwagon.
> Finally just making the link non-clickable does not add security.
I disagree. Some of the commenters here are rather savvy and will properly evaluate what they are downloading. Some are not. Making a link unclickable will prevent a percentage of people from downloading. If shenanigans are discovered, someone will make a very loud comment warning folks to avoid the download. In that case some of those non-downloaders may have been saved from themselves.
Again, this wasn't a well thought out decision, but it was also a rather low impact decision, and I stand by it.
And write and entire novel research paper and open source the code and put it on GitHub? No you wouldn't. Don't be ridiculous.
You are moving the goalposts, why?
Regardless, generating a plausible sounding paper with source code is trivial with gpt4. Obviously it wouldn’t withstand scrutiny, but neither would my coinminer.
I wanted to expose it so people didn't have to comb through the github, but decided to make it unclickable out of an abundance of caution. This appears to have offended people.
I would not have hesitated to link to hugging face. That is a known quantity.
Seems impressive!
So it is not 'open' then and you cannot make money out of this?
No wait…
EDIT: I guess there’s a third option, “work another job and use OSS on your off hours”. Which feels… idk, disrespectful of the whole enterprise. OSS software development is important enough to deserve a wage IMO, to say the least
> You can view, use and modify the code to your hearts content.
The non-commercial clause of their license specifically prohibits commercial use, so we cannot use this source, and presumably the data that the source uses, to our hearts content.
The OSI has a definition of open source that clearly states commercial use is required [0].
Wikipedias entry on Open Source Licensing also stipulates that commercial re-use is required [1].
There is a term called "source available" which is more in line with your intent.
While this is very true, the context of "open source" can't be assumed.
It will definitely stop those bad actors from scamming people this time, right? Right?
Additionally, it only really hurts small businesses & startups as the big companies all have teams that can make their own version or pay for 3rd party apis for easily. So yeah, us startup folks won’t like this license much as it basically is aimed at us the most.
Either way, congrats with the tech. It does look very impressive!
unless you're proposing it's use in detecting itself is some how symmetrical, which I really don't think is anything but unproven conjecture.
Now I've got Just Another Item on my ToDo list, to get that undone. Gawd, does every company promote it's stupidest people to management?
They have so much money that competence no longer matters and bootlicking will get you much farther.
Yes: https://en.wikipedia.org/wiki/Dilbert_principle
Ironically, this is the place where they can do the least damage.
Most likely this is caused by SCA another European directive that ruined our lives with extra security hoops (for payment providers) for little extra security - or even worse in case of voice password or security questions
> Processing personal data is generally prohibited, unless it is expressly allowed by law, or the data subject has consented to the processing
Given the same voice get processed and recorded during a normal phone call to the bank so you would need to give consent just to talk on the phone (and they do have a disclaimer when you are calling in Europe).
Most likely this is buried deep in some massive EULA you accept when you open an account.
EULAs are a bit more of a mess, as all the advice I've been given says "don't hide stuff like that" while all the websites I visit are "we're going to do this anyway because we think we can get away with it".
Europe is deteriorating incredibly rapidly in terms of crime (due to a combination of economic poverty and uncontrolled immigration from third world countries) - but I think some of the low tax EU countries (Malta, Cyprus, Gibraltar, etc) are a good bet for a few more years.
My top choice if I had family (or friends I want to be close with) in the US would be Cayman.
I think long term, either South America drops the level of crime considerably and becomes the new place to be or China start building futuristic cities attracting wealthy western talent to offset their declining population rate.
That's a British Overseas Territory, it isn't in the EU.
Cyprus has a lot of English speakers (and indeed a lot of street furniture that looks just like the UK, plus two UK airforce bases[0]), but the national language is Greek… I don't know if I'd risk that, given the one time I tried to ask for «Ένα σάντουιτς και ένα τσάι παρακαλώ»[1] in Athens[2], the person behind the counter replied in English to correct my pronunciation.
[0] https://en.wikipedia.org/wiki/Akrotiri_and_Dhekelia
[1] https://translate.google.com/?sl=el&tl=en&text=Ένα%20σάντουι...
[2] I know that's not in Cyprus, but it is, as you may guess, another place where Greek is the national language.
But if it's trivial to use somebody's voice to say any arbitrary thing, then it'll be done. Combined with deepfake videos, the result will be the ability to show anyone saying anything, including lies and things they find incredibly objectionable, in a disturbingly realistic way, and more so as time wears on.
The fundamental issue is that we don't live in a rights-respecting world. Making it easy to utter anything in the voice of anyone will lead to many more abuses than legitimate instances.
1. licensing voice to other uses - people with recognizable trademarkable voices (actors, singers) have another potential revenue stream. yay!
2. use of past voices - voices that are not 'owned' from the past - let's say Humphrey Bogart's voice, can be used in projects without having to pay for imitator. This would be useful for both marketing and artistic projects. But probably less for marketing because they will want to go with step 1.
3. Teach yourself to talk like X. People who need to learn to talk like a particular person / have a particular accent could learn quicker. Just think - you will be able to supplement your comedy routine with kickass Christopher Walken impersonations any day now!
Variations of 3 and 2 together open up interesting modes of aesthetic impression, but I won't go into that here. But definitely I have some ideas that might benefit from being able to do this.
Or...! Indie game development. I can learn basic voice acting (to get rid of the cringe), and act out all of my characters using different voices.
And in case anyone is concerned, I intend to make the purpose of the vocal samples clear to the provider and then arrange appropriate credit and compensation to those whose voices I used. I also don’t intend to train with anything but public domain and purchased data.
I think the most common use case will be making art & content programmatically without voice actors (and most likely without actors at all once we nail video or a 3d model pipeline + frame by frame transformation to make it look realistic)
Some banks have voice authentication when you call in and you have to ask to opt out.
Creative Commons Attribution-NonCommercial 4.0 only prevents the use of the licensed code within other projects. It says nothing about output. People can (and will) use output generated by this for commercial projects, and they will be completely within their rights to do so.
So detect away I guess.
This is zero shot TTS. Samples create vector encodings that serve as input to inference. There's no retraining the model unless you want it to generalize or perform better.
Funny enough, a lot of RVC packages are using VITS to do RVC for TTS.
It's allegedly the basis of the tech used by Eleven Labs.
You're banning genuine uses like that or just creators who want to fix a fumbled or awkward line without completely re-recording if you ban it.
“MyShell reserves the ability to detect whether an audio is generated by OpenVoice, no matter whether the watermark is added or not.”
Call me skeptical…
https://docs.myshell.ai/tokenomics
Tokenomics
Disclaimer: MyShell is currently in the testing phase, and the content of the whitepaper may be subject to change in the future.
$SHELL is the token used for user incentive, governance and in-app utility.
The total supply of $SHELL is 1,000,000,000
https://github.com/myshell-ai/OpenVoice/blob/a33963c3d764bee...
$SHELL is the token used for user incentive, governance and in-app utility.
The total supply of $SHELL is 1,000,000,000
Team, Treasury, Advisors & Private Sale = 55% allocation
Community Incentive = 40% allocation
Liquidity = 5%
Auto non-toxic rephrasing of online chat in video games, let people hear their voice but paraphrase what they said in a manner that doesn't turn the platform into a cesspit.
Cloning your own voice so that you can turn a script into audio without 50 takes and then having to remove a million Ums and errs.
that feels very orwellian
I think this is closer to the direction of Huxley in Brave New World, where a deeper understanding of how to manipulate without brute force creates a very different dystopian society than 1984.
BNW had a similar effect by conditioning, rather than by applying the strong form of the Sapir–Whorf hypothesis.
Mr Beast talked about translating his videos to other languages to get more reach. This can be done for people with limited budget or just in general so people can watch videos without needing subtitles.
I wouldn't be surprised if we saw this incorporated into YT in the near future.
Certainly entertainment. Movies / TV. It opens a new opportunity for videogames with generative characters.
Of course, there's also the preservation of the voice of a loved one. I would probably pay to hear my father's voice again but there"s probably only one or two VHS tapes with his voice on it.
Outside public speakers, there’s probably other people whose lost their voice or have trouble vocalizing who might want to sound like their old selves. This could help them.
Disclaimer: I think these techs will more often do damage than good. I’m just brainstorming an answer to your question.
The real answer is yes, I could probably come up with some contrived examples, like I lost my voice in a freak LLM accident and now want to clone my old voice. But this doesn't (you don't?) really need a net benefit reason to figure it out and publish it. Because why? I assume, because "this shouldn't exist!" which is just a more palatable wa to phrase "won't someone think of the children".
Society doesn't benefit from ignorance, so given it can exist, what's the problem with it existing? Why does it need a practical reason? Because people will do bad things with it? Duh, but I'd rather everyone know then just the bad guys
To at least give us something as a consolation for all the havoc all sorts of deep fakes will wreak on societies. It's like asking what a knife can be used for other than murder. It's a valid question.
Scams aren't going away. Will this make it easier to scam some people? Absolutely, so did the internet. I'm not claiming this is anything like the internet. My argument closer to, the reason people get scammed isn't because [thing exists] it's because bad people lie, and kind people trust them. We can all wring our hands in fear over what the new technology might do, or we can solve the problems we care about. Authenticity was hard before this, and it'll be hard after.
> If you look at the enormous fruits of human genius that mankind has developed in the last 50 years, atomic energy and rocketry and flying to the moon and coherent light, and it goes on and on and on -- and then it turns out that every one of these triumphs is used primarily in military terms. So it is not reasonable for a scientist or technologist to insist that he or she does not know -- or cannot know -- how it is going to be used.
-- Joseph Weizenbaum
That is not fear. That is being serious and unflinching, if anything.
> We can all wring our hands in fear over what the new technology might do, or we can solve the problems we care about.
I'm doing neither. I said it's a valid question, with which you agree. The rest is a straw man apropos nothing anyone actually said, here, and wringing your hands about it. It's a way bigger waste of time than asking a simple question and let those who want to answer that, and let those who don't want answer it simply don't answer it, instead of making up this "issue" with the question itself.
> Authenticity was hard before this, and it'll be hard after.
So "nothing changes", but technology is super important? You could say the same about, say, curing cancer. People will live for a while and then die, with or without it. Why since it makes no difference, what'd be the problem with "fearing" it?
I was curious to see if anyone could name at the top of their head some practical use cases that they feel net out the potential harms of cloning and misusing someone else's voice.
There's some nice and certainly practical examples, but I don't feel any of them would net out the harms.
Perhaps there's a use case that we can't even comprehend yet that would though!
While you can’t make it go away, you can disincentivize propagation and use which can be the difference between thousands of cases of scams/extortions and millions. Until there’s a stronger argument for voice cloning models (talking to a dead loved one is creepy and not a positive argument) then we shouldn’t encourage tools with overwhelmingly nefarious utility.
Hurting people, lying, that's already illegal.
I think Maybe you misunderstood my argument. My argument isn't that good guy with a voice cloaner is the only thing that can stop a bad guy with a voice cloaner. That's, as you pointed out, stupid. My argument is that no one benefits if how easy it is to make one remains a secret to everyone but the bad guys.
Small, independent film makers can now use a skeleton crew to voice parts.
I can't imagine it would be anything other than a niche service, but hearing the voice and, potentially, interacting with a chatbot/LLM with the voice of a passed love one.
This is off the top of my head. I would also guess that this technology is a stepping stone for other weird, interesting and profoundly helpful uses.
[0] https://www.theverge.com/2022/9/24/23370097/darth-vader-jame...
Alexa, siri, and similar, are all common place.
Another huge usecase would be anything to do with voice acting. Either in video games, cartoons, or the like.
This would completely democratize voice acting material, and would empower anyone to be able to do this for cheap.
Democratization will always be the enemy of those who profit from preventing others from being empowered.
At least for now there's too much lag to do a real time conversation with a cloned voice.
Speech to Text > LLM Response > Generate Audio
If that time can shrink to subsecond, I think there'll be madness. (Specifically thinking of romance scammers)
It worked a bit too well, as it could parse the sound file and generate a complete response faster than real-time, leading people to ask if he'd actually listened to the messages they sent him.
Also they had trouble believing him when he told them how he'd done it.
This is a society-destroying idea.
Most of us, especially younger people, only know how to vote, where there are wars, or even what our parents are doing by using digital media.
If digital media becomes untrustworthy, everyone will live in a warped and fragile alternate reality that no one can agree on.
> This is a society-destroying idea.
Believe it or not, this is how much of the population saw The Internet when it first came close to being mainstream. Everyone and their mother said "Don't believe anything you read on the cybernet", which ended up ironic as everyone and their mother ended up being the ones to believe anything on the cybernet anyways.
> everyone will live in a warped and fragile alternate reality that no one can agree on.
How is this any different from today? The various corners of the internet (which is mostly divided by languages: English, Russian, Spanish, Chinese and Portuguese) already have these vastly different realities and ground-truths.
I'm sure we could survive another Internet-Winter where people trust everything a bit less than today.
If it now becomes impossible to trust a voice received through the internet without being connected to a verified telephone number I don't know how that can be classified as society-changing.
https://github.com/underlines/awesome-ml/blob/master/audio-a...
The thing that changes is the complexity to run it. I was training my wife's voice and my voice for fun and needed 15min of audio and trained on my 3080 for 40 minutes.
Now it's 2 Minutes.
If it is a bad thing, should we cheer it on?