Emotionally Expressive Text to Speech
sonantic.io
sonantic.io
But I'm very curious what the emotional "parameters" are? There are literally at least a thousand different ways of saying "I love you" (serious-romantic, throwaway to a buddy, reassuring a scared child, sarcastic, choking up, full of gratitude, irritated, self-questioning, dismissive, etc. ad finitum). Anyone who's worked as an actor and done script analysis knows there are 100's of variables that go into a line reading. Just three words, by themselves, can communicate roughly an entire paragraph's worth of meaning solely by the exact way they're said -- which is one of the things that makes acting, and directing actors, such a rewarding challenge.
Obviously it's far too complex to infer from text alone. So curious how the team has simplified it? What are emotional dimensions that you can specify? And how did they choose those dimensions over others? Are they geared towards the kind of "everyday" expression in a normal conversation between friends, or towards the more "dramatic" or "high comedy" of intense situations that much of film and TV lean towards?
We can express emotions without words:
xxx: Distress
yyy: Support
xxx: Hope
It maps on music and we have dictionary to describe it. The one I'm listening to is Sorrow and Hopeful - entire track. May be a good start. Write first (classification).
Examples you gave I feel live on same scale but extreme values. So even harder.
I'd imagine it work like autotune - enhance human input
So that's not markup along "emotional" lines, but rather along "technical" attributes such as speed, pitch, volume, pause between words, and so on.
Obviously coding those things in XML manually would be a nightmare. Now I find myself wondering if 1) these technical parameters can be used to synthesize speech that does sound like a reasonable approximation of emotion (or if they're insufficient because changes in resonance and timbre are crucial too), and 2) if there are tools that can translate, say, 100 different basic emotional descriptions ("excitedly curious", "depressed but making effort to show interest", etc.) into the appropriate technical parameters so it would be usable.
Anyways, just a fascinating area of study.
Thanks for your thoughts and feedback thus far! I'd be happy to answer questions (within reason) about our latest cry demo / emotional TTS! Feel free to fire away on this thread.
Would you do demos for well known speeches/texts? It'd be easier to put this into context that way.
Clearly you can't give away too much on your "secret sauce" but is there any insight you could share on two questions:
1. Do the individual voice talents need to express the emotion types you use or can you layer it on after? (ie do they have to have recorded say "happy" to get happy outputs or can that be added to neutral recordings retrospectively)
2. What are the ball park audio amounts you need per voice? 10 hrs, 20 hrs or more?
This sentiment definitely gives you lots of credibility, only those who have seriously endeavored in this space are able to acknowledge just how true this is.
It's quite antithetical to how some ML folks like to think.
So the question is - what's there? Is it formants? Is it universal? Can we map them like syllables?
And music, it touches same emotions. Does it use same mechanism?
Edit: found "Emotional speech synthesis: Applications, history and possible future" [1], looks like melody is part of emotion processing.
If mapping is possible I'd love to see application in dubbing. Both as translate and TTS with mapped emotions and dubbing actors evaluation/autotune.
[1] https://www.researchgate.net/publication/268260426_EMOTIONAL...
I don't know either of the co-founders, but it seems like a logical, good idea to have a pair of co-founders where one is technical and the other is non-technical (maybe marketing, or sales, or very strong soft skills, etc).
Hence, I don't see the issue you (obviously) have with only one person having done the technical work. Is there any context you're not telling?
As an engineer, I've been approached by those experienced in sales and offer me to build a mutually agreed upon product that we agreed would make money for equity. I would like to know if this model works. If so, how does it work? Is this common?
You can take offense if you like, but if you meet a guy at an entrepreneur conference who already built a prototype, you're not a co-founder.
Are you looking to make this accessible (read: affordable) for small time content creators / hobbyists? What sort of pricing model can we expect (one time license fee / subscription)?
It looks and sounds awesome!
The same goes for sub titles, she'd be perfectly fine with a robot voice for the actors if they sounded real enough like this.
Game changer.
I generally use Kurzweil 3000 (http://KurzweilEdu.com) which is made by Kurzweil Educational Systems, as a screen reader. You should definitely considering partnering with them in particular, as it would be very strategic.
[0] www.blockstud.io
The fault is with the owners of the sites.
It's like complaining that your carbon monoxide alarm is way too loud and beeps too often.
I was just mucking around with Nvidia's latest, called flowtron, and I know from that experience there's a significant amount of work between getting a tech demo out and launching a usable product, whether API-based, or with some visual workflow like your video shows.
One thing I think worth considering on the commercialization front is whether or not the core offering is the workflow niceties around your engine, the engine-as-API, or both. I'm just a random person on the internet, so take these thoughts with a large grain of salt, but thinking about it, it seems to me that prioritizing integration with say unity, unreal engine, video compositing tools, blog posting tools are all interesting and viable market paths. The underlying networks are going to keep improving for some time, so you're really trying to buy some long term customers.
Some stuff that's obvious, but I can't resist:
I could off the top of my head imagine using this for massively reducing the cost to develop games, for script writers pulling comps together, for myself to create audio versions of my own writing, for better IOT applications inside the home... I'd really love to be able to play with this.
There still isn't a truly non-annoying virtual assistant voice; when the first tacotron paper came out, I was hopeful I would see more prosody embedded in assistants by now, but the longer we live with siri and google, the more sensitive I think we are to their shortcomings. I have a preference for passive / ambient communication and updates, so I would place a really high value on something that could politely interrupt or say hello with information.
At any rate, congratulations, this is cool. :)
We soon can create emotionally expressive youtube videos with synthetic actors..
The comment: I noticed that your demo video also had "emotional" video layered on top of the dialogue. This could be considered manipulative; perhaps consider sharing a naked version so we could attempt to interpret the emotion based solely on the text to speech engine.
The question: You mention you met at EF. I was wondering if, beyond bringing you together, you found EF to be worth the cost of admission?
Close your eyes
I thought the demo was impressive, but these things do seem like an effort to distract from (or more accurately bolster the effect of) the core technology.
Though maybe the right call since this is less a strict technical demo and more a way to drive interest/marketing.
The 'high levels of expressivity' comment was more of a flag to me, it's a meaningless phrase alone but it's suggested as an obvious answer. It feels like a mysterious answer [0].
I recognize though this is a marketing video, the core tech demo is cool, and I'm probably being unfairly critical. Flags like that make me more skeptical than I would otherwise be by default.
[0]: https://www.lesswrong.com/s/5uZQHpecjn7955faL/p/6i3zToomS86o...
As others have said, this is first and foremost a marketing video aimed at attracting target customers. We've got additional clean samples (without background music) further down on our homepage and we plan on adding even more on their own subpage of the site in the future. We've also done a few technical demos at conferences over the past year and will continue to do so.
We did meet at EF and it was totally worth it! There is no cost of admission for EF, they actually pay you to complete the program! Granted the monetary funds they provide could be a heck of a lot less than what you are earning at a full time job, so everyone's opportunity cost is different. EF's biggest selling point is their world-class network of highly ambitious individuals, so if you're interested in founding a company (pre-team / pre-idea) I would absolutely recommend looking into it.
We don’t see FOSS pharmaceutical research for instance, I believe for the same reason. The amount of coordination needed and the impossibility to separate TTS projects into sub-parts could also factors.
“Common Voice is Mozilla's initiative to help teach machines how real people speak.”
For something like Sonantic, you need clean recordings from professional actors in proper recording environments (not to mention the in-house expertise to then filter these down to curate the training/test datasets). That costs money. A million people with laptop microphones will just never get there.
https://github.com/NVIDIA/tacotron2
https://github.com/CorentinJ/Real-Time-Voice-Cloning
https://github.com/mozilla/TTS
I think a lot of the remaining gap is due to a lack of high-quality training data -- most of the open-source models are trained on public-domain audiobooks (e.g. LJ Speech).
However, good training data (large amounts of annotated recordings by professional voice actors) is expensive to create, and unlike code, there's not a tradition of people sharing it.
I’m not really an expert. From what I understand, the “cutting-edge” stuff requires pushing past the point where we are splicing segments of speech together. Splicing segments together is hard enough.
There are a couple open-source efforts like Mozilla’s, but if you want something like Lyrebird, well, that technology isn’t even really productized commercially yet.
It would be nice to try with actual text inputs right on the page, that this doesn't exist is tiny flag.
A great choice to work with voice actors, because there isn't any 'pure' TTY that's good enough in the most general sense, having the actual voice actor as a working basis will help.
Perhaps for small game houses, they can just use something off the shelf, big houses can use a customized voice, and then not worry if they have to make tweaks or changes, they don't have to do a whole production.
As you've mentioned, we do work with real actors to create our TTS and take misuse of their (artificial) voices very seriously. Because they sound so lifelike, we've made the decision not to allow public access/personal use at this time.
Lastly, your assessment is spot on regarding standard vs custom voices. Lots of interest for both!
Not diminishing the quality of your product, just pointing out an obvious expectation of the audience that it's presented to. Perhaps, there could be a way to test-drive it directly, with limited choices or combinations of the input text.
Next time be honest about what you have when presenting it; every human with functioning ears is attuned to the sound of speech. This sort of technology would be amazing for narrative video games even with the less than perfect vocoding.
Amazon's Polly English voice, Matthew is pretty nice. But they don't have Hebrew. Also Google doesn't have Hebrew. Bing has some attribution requirement that I haven't fully investigated.
I wonder if attaching this to a modern-day Elisa will improve the Turing test scores? Emotional load can reduce the requirement for semantic coherence.
If so, have plans for a Web Speech API plugin? I'm about to release a reader demo based around it. https://developer.mozilla.org/en-US/docs/Web/API/Web_Speech_...
Obvious application: H-anime. Reduced parameters for the "emotion" as well.
Ideally, it would be able to infer the emotion from the text itself, but I think that level of sophistication is a long way off.
Edit: Actually, this might be a perfect candidate for some sort of crowdsourcing. Imagine Wikipedia pages containing hidden annotations for the proper text-to-speech "tone/cadence/whatever" of each sentence or paragraph.
<dismissive>how would it know?</dismissive>
<sorrow>how would it know?</sorrow>
<angry>how would it know?</angry>
This is obviously an early demo, but this isn't yet to the level you could narrate an audio book - those little problems will quickly become noticeable.
One of the other challenges with using outside voice talent is that it can be inconvenient/expensive when you need to add/change something. I've been involved with podcasts using an external host and one of the negatives with that process is that if you discover a minor mistake/glitch in the narration late in the process you can't easily fix it.
But then I remembered voice acting fluctuates wildly in quality outside AAA games.
I'd take well-acted (with rough audio quality) over poorly-acted (but high-audio-quality) voices for most games.
I want the know the price and when can we use it in production.
haha thanks, we'll take that as a compliment!