Deep Learning for Siri’s Voice
machinelearning.apple.com
machinelearning.apple.com
I would love to have my home speakers announce things in this voice.
Will be interesting to see how Siri on Home Pod works out.
Hopefully it gets ported to MacOS's say CLI utility. I typically use that with `pbpaste | say` to read my articles.
Is the one in Sierra different from Siri in iOS 10?
Nope, macOS uses a different set of voices for TTS, which are named differently as well. They work offline and are nowhere as good as Siri.
I'm guessing Apple simply had to make this change in order to stay in the game.
Did you ever see the list of authors on publications of CERN? Or take this example: [1].
[1] https://www.nature.com/news/physics-paper-sets-record-with-m...
Not that there's anything wrong with that and it certainly seems like Apple has been investing in-house pretty heavily in recent years for Siri improvement.
That said, in my limited experience, organizations doing NLP also have something to offer by way of TTS too.
One of the ideas is that most ML researchers want to publish their work and Apple wasn't allowing it. Allowing ML researchers working at Apple to publish in this journal was they only way they were get more ML researchers to work for them.
Google also cares about privacy, but only in reverse. They don't want you to have any ;)
Could offline models be better? Definitely. But they only way to make them as good as cloud models it to make the cloud models worse.
Anyway, it's the only way I can still identify this as a fake voice. The intonation always follows the same cadence (not sure if that's the word?). We really shouldn't have overused the word awesome before this kind of thing came along.
There's also a kind of dread too, tbh, this kind of seamless TTS has the potential to change a lot of things. First of all criminals are going to love this, youtube pranksters too. Eventually this will shake up the voice acting industry in a possibly not healthy way for the voice actors, while at the same time allowing projects with a shorter budget to have incredible voice work (also dubbing).
What I think is really important, tho, is that as we move away from the uncanny valley we change our relationships with those voices, our brains don't have the capacity to listen to a voice this real and not imagine it as a person, even for adults.
Ironically at this moment I'm using an old threadless sweatshirt that says "this was supposed to be the future" but nowadays I can honestly say we're getting there.
I do think there are applications that we don't just have today because TTS just isn't good enough. I've had some ideas around Alexa apps related to content that would be TTSd. But the current Polly just isn't human enough. I don't think this is there yet either but it's getting close.
Similarly, we don’t see CGI motion capture replacing Andy Serkis any time soon.
I'm pretty excited about the video game side.
Listening to those samples I remember how big an advancement iOS 10 felt, but it's nothing compared to 11.
But I suppose an AI might choose to use a computer-sounding voice to remind us that it is a computer. Kind of like those inaccurate sound effects in movies - they have become so common that it seems more wrong to omit them. (TV Tropes calls this "The Coconut Effect".)
There is always the chance that as we get better at this stuff we'll start to find it creepy that it's so realistic (either due to the uncanny valley or because we crossed the valley) and we'll start to prefer devices that act robotic even though we know we could make the indistinguishable.
I'm trying to think of another example. I know I've heard a good one with Roombas but I can't remember it.
Basically we may try to avoid a Bladerunner situation where we're not sure when we are or aren't talking to a real person and prefer the 'computery' voices.
Personally, I'm less pleased with the actual new voice itself, although that is more a subjective judgment. After listening to many hundreds of voice talent auditions for Alexa, it's hard to step back from that level of pickiness.
I actually tend to generally prefer some of the female British accents in several current TTS systems. (Amy is probably my favorite Polly voice.) Perhaps as an American, the robotic-ness doesn't seem quite as obvious or grating.
> For more details on the new Siri text-to-speech system, see our published paper “Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System”
[9] T. Capes, P. Coles, A. Conkie, L. Golipour, A. Hadjitarkhani, Q. Hu, N. Huddleston, M. Hunt, J. Li, M. Neeracher, K. Prahallad, T. Raitio, R. Rasipuram, G. Townsend, B. Williamson, D. Winarsky, Z. Wu, H. Zhang. Siri On-Device Deep Learning-Guided Unit Selection Text-to-Speech System, Interspeech, 2017.
Why not just add the names by default?
That's why the Easter egg in Adventure with the programmer's name exists. It was the only way to get his name out there.
What happened was the developers didn't like this and left to start their own company, Activision, which made some of the best remembered games on the 2600.
Apple already 'compromised' by letting their researchers publish at all. Maybe names will be allowed in the future but it's kind of surprising we're even getting this.
I totally agree that this was an unprecedented move by Apple, considering their past stance on such things. I'm hopeful for the future, though! They seem to have realized (at least a little bit) that community cooperation is valuable.
https://www.amazon.com/Asimov-Eliza/dp/B0184NR4P8
Not sure what their stance on privacy is though.
The VC world is interested in “traction” and not novel tech which means we have to divert effort into growing customers for our mental health practice management system to get “traction” before we can spend any notable time building AI therapists. As much as VCs talk about “looking for innovation” they really aren’t. They are just looking at current growth/revenue. The days of building something amazing and monetizing later seem to be over for all except for founders with marquee names.
We could launch AI therapists within a year, but in the meantime, I have to pay my team. So we are forced to subsidize moonshot R&D with our existing sales — but that is hard to do since existing sales have to finance customer acquisition. Finding an additional $500k per year to make AI therapy viable is impossible for us.
We are in a catch 22. The first question from nearly every investor’s mouth: “how many paid users do you have?” Not, what technology do you have or can develop that is truly disruptive. We could start preparing AI therapy tomorrow for a Summer 2018 launch if we could afford it. But if we diverted resources to that, we’d be out of business long before launch. Clinically effective AI therapy isn’t a weekend side project.
The biggest challenge I saw was gaining the trust of sufficient therapists and patients and then an IRB. If you've solved those problems and can actually demonstrate data collection and a computational workflow with any evidence of results whatsoever, I suspect money would not be a problem.
Not crazy at all. At least some therapies provide benefits even with simple non-AI processes: "A meta-analyses of 15 studies, published in this month’s volume of Administration and Policy in Mental Health and Mental Health Services Research, found no significant difference in the treatment outcomes for patients who saw a therapist and those who followed a self-help book or online program."[0]
[0] https://qz.com/1057345/researchers-say-you-might-as-well-be-...
The WaveNet method of predicting the output sample by sample yields great results but at a very high computational cost