say 'this is a test'
I for one think it's very good quality (at least for the Alex voice)I'm not even a stoner and folks seem to get a kick out of it whenever they're pair programming with me.
We use TTS extensively within the Pantelligent iOS and Android apps, and it's something our users requested and get a lot of value out of. It seems like the existing solutions are already good enough / dramatically above the threshold of usefulness for an interactive real-time-guidance application like ours, and just keep getting better from the OS side.
- The "bank" one referenced above: this is made of short recordings of a real person saying the words or phrases, cut up and then concatenated. For some messages they sound exactly like a real person (because all it does is play a single recording), but when numbers are inserted, the above characterization is quite accurate. There is no effort to make the inflection fit properly in a sentence or have it sound natural.
- ivona.com and OS X `say`. These generate audio in real time, and may have a few samples but are generally created on-the-fly according to what is around the text. This is where the research is at right now, but the main problem is the CPU required to generate these. Your car, or Madden 2015, or the bank might not want to use up too much CPU time to make their audio sound like that.
Either way, the incentive for the bank to change things must be minute.
It's not true that TTS hasn't improved (see below). Many people are working on this, both in academia and in private enterprise. It's an obvious and potentially valuable part of the human-computer interface.
This is not to suggest that it's easy -- the mathematics and vocal tract modeling problems are formidable. The only reason there are reasonable TTS resources now is because of the rapid increase in computer power -- power that's needed to support this feature.
Here's a site chosen at random that offers a high-quality TTS example:
It's pretty good based on prevailing standards, and it's the outcome of a lot of work.
To find the companies working on this, just Google for "high-quality tts".
Edit: Perhaps a Kickstarter or related would be a good idea since this type of feature would be useful by so many people. Nearly everyone has functioning ears. (no offence to those who don't)
A lot of really great work is happening in academia; I'm not going to name names because I'd forget someone deserving.
(Shameless plug: we [0] do speech and language consulting including custom TTS.)
[0] cobaltspeech.com
Also, I'm really linking to research groups, so take these names as starting points and look at their students and other professors working with them.
First off, the Blizzard challenge is a major hub of activity. [4] Festival is an important piece of software [5]. Interspeech is an important conference [6] (take a look at the speech synthesis track and the organizers for that).
Alan Black [0] @ CMU is kinda a giant.
Keiichi Tokuda [3] @ Nagoya also a giant.
Simon King [1] @ Edinburgh, just had a paper linked to on HN a few days ago and does important work.
Mark Gales [2] @ Cambridge does work here, too.
[0] https://www.cs.cmu.edu/~awb/ [1] http://www.cstr.ed.ac.uk/ssi/people/simonk.html [2] http://mi.eng.cam.ac.uk/~mjfg/ [3] http://www.sp.nitech.ac.jp/~tokuda/ [4] http://festvox.org/blizzard/ [5] http://www.cstr.ed.ac.uk/projects/festival/ [6] http://interspeech2015.org/wp-content/uploads/direct/INTERSP...
It switches back to normal settings when you go below 10mph for over 1 minute. Helpful with switching everything to TTS when you're driving automatically.
Lots of other functionality as well. Check it out: https://play.google.com/store/apps/details?id=com.org.imsono...
May be of interest to you.
In general, TTS is a better interface than STT, both are (in my very, very humble opinion) bad interfaces.
Read your comment or my comment aloud to yourself. Now feed it to your choice of TTS engine. The problem isn't the words, the problem is the lack of comprehension.
1. Nuance
2. AT&T
3. IBM Watson
However, I can definitely understand how TTS technology looks stagnant. Part of this is that going from nothing to something reasonable happened exceedingly quickly. Early TTS research was supported by the US government which saw that early systems were comprehensible (if not wonderful sounding) and declared victory. Funding went to other problems in computational linguistics (like speech recognition, information extraction, etc.) and so did a lot of the workforce.
Modern systems usually involve many hours of speech from a single person and use variable length units to form more natural speech. Many systems still sound pasted together because that's how enterprise technology goes. How many banks have online banking that seems like it's from the 90s? You can't compare what systems can do to what some call center has installed as its technology. Someone has linked to Inova. Google and Nuance have good systems as well, but there's a balance between resources and perfect speech.
In terms of some of the issues. . . When you're going through a finite amount of recorded speech, you have to choose something that fits. It isn't going to be perfect in many cases. You have to deal with things like F0 declination. You have to deal with how long phonemes are going to be. You have to deal with breaks in utterances.
And the fact is that we can understand Hawking's 1980s TTS.
If you want to start thinking about the problem more, try inputting these two statements into Inova "Do you really want to see all of it? Do you want to see all of it? I want to see all of it." Notice how it tries to rise around "really" in the first sentence. It's trying to match how we would speak - rising for "really" in the first sentence and rising for the question-ending in both questions. But it kinda misses in both cases. Still, in some ways it's amazing that it recognizes "really" as something that should go up. It recognizes that questions go up at the end. It recognizes how the non-question goes down as the sentence progresses. And it finds things within its data set to fit to how it thinks the sentence is going to go. But it doesn't have perfect language understanding so it doesn't know exactly how things would be said - a lot of sounding natural isn't making the phonemes more accurately, but the intonation and attitude of the speech. It also has to find something that fits. Lots of smart things are done, but it's pulling from a limited amount of recorded speech - speech that has been sliced in many useful ways, but still limited.
TTS has definitely evolved and I think that Google, Nuance, and others are definitely pushing it forward. You're going to interact with a lot of legacy systems that feel like they're still in the Hawking era. But most ATMs I use don't even have touch screens (opting for buttons on the side of the screen) and even fancy ATMs like Wells Fargo don't feel like an iPad. You don't want to compare to systems that are so far away from modern, commercially available systems.
There is definitely work being done on it and it's definitely become much better. But to an extent, it isn't something that a lot of companies are going to work on. How big is the market for TTS? Before you say, "it's useful in loads of things," think about the market for maps. A lot of it is Google or Apple Maps. Loads of apps integrate mapping, but don't want to map the world or run their own infrastructure for serving it. Some use OpenStreetMaps, but they're really just serving generated tiles rather than re-mapping the world. If you were to create a TTS startup, what would your business model be? Pay us money to TTS your text rather than getting it for free from Google's Android TTS? The issue is that TTS is more a feature than a product. Companies like Amazon might bring a company like Inova in-house. Nuance sells their stuff to companies like Apple. Google is large enough to build it. But you'd be pitching something that doesn't directly solve something for customers and trying to hire very smart (expensive) people who you hope will be able to come up with something new that isn't an obvious "with more resources, better" solution. Remember, TTS needs to be done in real-time, possibly on low-powered mobile devices (don't eat their battery or storage) or over the network (don't make our AWS spend go through the roof). If you're going to sell to an app maker as a no-network TTS, how much bloat are you adding to the app?
It's just a hard market to be in given that the 1980s solution works, even if it doesn't sound realistic and modern already-available systems are quite reasonable.
Has anyone else wondered why Stephen hasn't upgraded his voice? Maybe it is his signature of sorts.
That's exactly what he's stated publicly - it's so widely recognised it's part of his identity.