AI voice actors sound more human than ever–and are ready to hire
technologyreview.com
technologyreview.com
Once, I was riding a two-car local train through a rural part of Kyushu when the train was delayed by heavy rain up ahead. Soon there were audio announcements with artificial voices about the delay in Japanese, English, Chinese, and Korean. Both the Japanese and English sounded a bit strange—the English especially because it was clearly machine translated—but the meaning was clear and they served their purpose. (I couldn't assess the Chinese or Korean.)
The pitch of the artificial voices must have been chosen for good audibility in a noisy environment, as they were easier to understand than most live announcements by human conductors on trains.
Professional human voice actors certainly would have been better, but there is no way they could have been arranged at a reasonable cost for up-to-the-minute train delay announcements in even one language, let alone four.
In my opnioning this is a good example of how AI is slowly gaining foot. It's a bit like the old adage of a camera phone vs a DSLR. Sure, it's inferior, but it's always better than a superior version that I don't have access to. It's difficult to compete on quality with something that's free.
Yes. Machine translation is like that, too. It makes mistakes and can be hard to understand and easily misunderstood, but compared with the time and effort required to learn another language, or the expense and bother of finding a human translator or interpreter, it clearly has its advantages.
It's those unexpected use cases that often become the killer app of advancements like this.
The real skill of a voice actor is to collaborate with the writer on deciding the specific vocal choices to make during a performance to give it specific character and intentionality. There's been no work in this area and, indeed, it's hard to imagine how you would even design a system to annotate text with specific intonation cues that isn't many times slower than just saying the thing out loud.
eg: Imagine you had a WebMD page and you gave it to two different voice actors and you told one to try and make this drug sounds as promising as possible and another to make it sound as scary as possible. One piece of text can produce two vastly different audio products depending on intentionality. But with AI tools, there isn't even a setting for getting different audio from the same text and it's not even clear how you would build the setting that would allow you to convey what intentionality you want from the AI.
In short, there's a market for just handing a raw piece of text over to a human to convert into audio with no other interaction and that market is amenable to be taken over by AI but is already quite cheap and quite a small market. There's an entirely separate market where the voice actor sits down with the author and talks over what the intentions of the text are and how best to convert them to audio and nobody is thinking about building AIs that can understand intentionality any better than a human would.
This isn't about just text-to-voice AI but about how we build AI systems in general.
The AI we have now is good for
a) problems with an unambiguously correct answer that AIs can perform better than humans (like camera tracking a scene).
b) problems where the deficiencies of AI performance are made up for by other characteristics (like machine translation which is often bad but at least its instant)
Where we have no real advances in AI are problems in which there are multiple valid answers and how to coach and give feedback to an AI to improve their answer more in accordance to your preferences. People who make pollyannaish predictions about AI takeover fail to understand this distinction.
eg: It's relatively easy to make an AI that can spit out a bunch of brand new recipes. But we don't know how to right now take an 85% good recipe that an AI generates, make it, taste it and then give AI feedback that would allow it to improve that recipe. Like, we don't even know the UI for that. About the best UI we have for it right now is to have the AI show you two versions side by side and you pick the better one until it gradient descents into what you want but that's unbelievably clunky for most tasks.
Where humans excel is at the interface and the interface makes up a surprisingly large part of many tasks, especially ones that white collar workers mistakenly regard as "menial".
In the bay area there's "Junípero Serra" which is usually read by text-to-speech systems as "June i pair oh sara" instead of "Wun ip airoh sara" (and I don't even know that's correct and I'm not good at phonetics to show where the accents go. I only know "June i pair oh sara" is not correct. It would be like if they said San Jose" as "San Joe's" as in like "This ball belongs to Joe. It's Joe's ball" -> "San Joe's"
In Hawaii there's Kalakaua Parkway, the main street through Waikiki. It's pronounced "Kah lah kah oo ah", not "kah lah cow ah" which is what all the navigation systems say.
It really feels like these languages are being destroyed by attrition. Each year a few percent less people know the correct way to pronounce something. I'm not into the whole French level of language protection but can't we at least put some effort into this?
Sorry, I might be a bit biased as I know several voice actors, and have routinely work with vocal tracks in the edit bay. We would be having a field day with the jokes if asked to use any of this.
It's a better robot voice sure, why not sell it as such? Selling this as human is harmful to the industry reputation and also plainly ridiculous.
They had trained on just about every character in "My Little Pony: Friendship is magic" and various characters from other games (Portal 2, etc), and it was remarkably good.
It's been down for months though, claiming that they're having a hard time training the latest iteration of the model :-/
Here is one of the characters that is available in both Vocaloid and Voiceroid.
https://www.ah-soft.com/voiceroid/yukari/index.html
You can see there are 6 tunables. The first three, "Speed", "Pitch", "Intonation intensity", seems pretty standard to TTS engines.
For "Anger", "Sadness" and "Joyfulness" though, I think very few other engines provides settings to those. (Can anyone else name one?)
The point is there is no one right way to speak a sentence. There is way more context than the text itself. Like the use of strange intonation in synthesized train announcements used in Japan to increate legibility in noisy environments, as another commenter mentioned.
Alternatively, voice actors are generally pretty cheap and they're usually not the limiting factor of much at all, though it does depend.
The cost of no-name actors is almost an insignificant part of production cost.
As someone note above - things like the voice enabled stop announcements for bus/rail/subway stops etc. - this would be ideal.
If I can dial in the pace of speech then I would almost prefer this to real actors, right now.
I am working on a product right now that's leveraging text to speech. I am presently using Microsoft and Google TTS because they seem to have models trained in a number of languages.
I'd be curious to hear what others are using, especially in the case where support for a number of languages is required.
There are so many hardware lead/bass/drum synthesizers being released these days, and hardly any (hardware) vocal synthesizers that I know of.
10 years? 50 years? 100 years?
Only a matter of time now.
You are talking English media to billions of new people, and massive libraries coming back to English speakers.
And dubbed way better than anything at the moment. You can mess with the speech to lips. Not just movies, also things like Khan Academy.
This is something that should be Moonshot.
I'd just be happy with some Russian Scifi dubbed for the first time to begin.
On Wish Dragon I was disappointed they referenced the game as chess, when clearly that was pandering to the English audience. It would be nice to fix that, it doesn't feel like that was the directors decision.
But sure, George Lucas could stuff a lot of dialogue up.
To me the concerns that original dialog will be lost are simply not real. If anything they can be restored if you have just the subtitles/original script
I think fan fiction re-dubbing could be really interesting. I think some movies could be turned around on the dialogue alone.
But to fear the authenticity of movies is at stake while spreading the worlds media across cultures, no, to many data hoarders to be losing original audio these days, I think we are safe.
And obviously on Khan Academy etc you do want to alter it. That's the point. Fix mistakes. Change things for different cultures. Speak a little clearer.
Especially gross is it when you hear the passionate voice of someone crying out - I particularly remember this when Myanmar was on the agenda - and then that perfectly understandable English voice fades out and is replaced with this all-artifacts-no-emotion-at-all voice.
Sounds pretty bad. This means we are going to make it artificially expensive for no other good reason than rent-seeking. It's not like the "actor" has to do extra work when someone licenses their voice.
I hope the answer will be to make true AI-based (as in, deep learning voices) that do not exist and screw all the rent-seeking opportunities.
Virtually anyone can use their own voice (or an entirely artificial one) in some work. If folks are willing to pay money to use one particular person's voice, what's the harm in that? If they don't want to pay to use that voice, they have literally billions of other options.