Lyrebird – a voice changer for Linux, written in GTK 3
github.com
github.com
I wish that they would release it in some capacity. Heck, even with a licensing fee. There's so many use cases beyond what Descript uses the tech for.
Not much else to say. Here's a lyrebird imitating a chainsaw: https://m.youtube.com/watch?v=mSB71jNq-yQ
Even old versions of this kind of thing are far more varied, eg: https://www.youtube.com/watch?v=sz53Dk37m-0
It's easy to try, just clone and run ./lyrebird, it runs right from the tree.
You can see demos of our voice conversion on https://storyteller.io near the middle of the page (section "3"), where my voice is converted into Donald Trump's voice. (I know, I should have used SpongeBob. We're going to have better product demos soon.)
(We're hiring if you're interested in virtual production. VTubing, Hollywood deepfakes and production pipeline inversion, or even SasS marketing tools.)
I dunno why people keep using scripting languages for when performance matters.
If they accept you for a job and then discover only later that you're actually female or non-white it'd be a big red flag if they rejected you only after seeing your face.
Even better, companies that were truly dedicated to diversity in hiring could remove names from any materials shown to interviewers and tell interviewees to use the same voice changer.
How we do it is to listen for other features of speech, accent and vowels, speed, rhythm, prosody, intonation, anomalies like glottal noise, dropped H's, nasal formant, murmuring diphongs.. the things that make us unique.
Impressionists (mimics) learn those. If they're not present in the source signal it's not easy to change or add them unless you move the whole signal to an intermediate form (speech to text) and then resynthesis the whole show (TTS) via a full articulation model that has those anomalous features.
If you get any good at this the people who you will piss-off are banks and folks who use voice as a blind authenticator (hint: your bank already does if you call them).
It's also pretty easy to figure out within the first week of employment if the person knows their shit or not.
Would only work for male -> female (and vice versa) changes. There is no vocal difference between a white male and a black male (or white female and black female); the aural difference is in the accent and actual regional slang, not in the voice.
Does it have to do with Twitch?
Sure, I’m nowhere near a bank director on the list of “attractive targets for voice cloning,” but who knows how widespread this attack might become in the future, and by the time one’s voice is out there on the Net, there’s no way to take it back.
I would like to use a voice changer when making phone calls to businesses too. I can totally imagine future corporations creating a new revenue stream by selling models of a known person’s voice to advertisers so they can later correlate the voice to that person.
Unfortunate that Lyrebird’s transformations seem easy to reverse. I wonder if there are any FOSS tools that make it harder to recover the original voice.
Your post gave me an interesting business concept. What if you could buy an anonymous, but totally unique voice, with its own speech patterns and pitches and everything, the way you can currently rent an anonymous email address? Each voice you buy could be generated as a downloadable data set, a one-time sale, subscriptions paid for updates to the software that interprets and reads it.
If it can do real-time mutation into the Ghostface voice (from Scream), that would be awesome and creepy AF.