Text-To-Speech engine in JS using eSpeak
syntensity.com
syntensity.com
I was a little disappointed after hacking Google's Translate API to get TTS working. Downloading an MP3 from a remote server just felt like an awful cludge solution. And there were http delays, cut off audio clips, and trying to queue up and play multiple <audio> clips never worked right because of delays between each clip.
This kind of stuff should be part of HTML5 and not just input[type=range]. Most OSs from Windows to iOS includes text-to-speech for accessibility. Why can't we hook into it from the browser? Give me a list of available voices/types (gender, country, accent), make all the other parameters 0.0-1.0, and I'll be a happy dev.
You should be thanking Mozilla for building an Audio Data API capable of features like this, and not wait for Apple and Google to get their heads out of their asses.
Whereas if you use a JS library in your project, you will know exactly how things will sound, you can include exactly the voices and languages you want, and only have to test once.
Would be no different than making web pages. Even JS isn't same across browsers. The big question is do I want to use a JS-TTS library or native and I think native will always be better performance/resource wise.
> And each will have its own list of available voices and languages, etc.
That can easily be solved by window.speech.listvoices like I mentioned in my post above.
> exactly how things will sound, you can include exactly the voices and languages you want, and only have to test once.
I've been programming for 20 years and that hasn't happened for any system, ever, not even once. You always have to test on different devices, OSs, platforms, browsers etc.
I do agree that a native library will have better performance, I am estimating something like 3-5X faster in the near future. So that is a benefit to the system TTS approach. But even the fairly unoptimized version in this online demo isn't too slow to be useful, I don't think, and it can be made much faster if necessary. I haven't focused on speed yet.
My concern with each OS having its own voices is that the names of the voices aren't enough to know what your users will hear. Unless we have a standard for TTS that includes the actual voice data, otherwise say "male UK English" may sound very different on different platforms.
It's clear there are tradeoffs here, both ways, and you make sound points. But I prefer the JS TTS approach, unless you are writing something like an iPad-specific app. If you don't care about other platforms, then I agree, system TTS is better.
eSpeak has a lot of settings, in the demo I think I use the defaults, but maybe things can be tweaked for better results.
(I don't really know much about speech synthesis or eSpeak, I just compiled it.)
Please consider how much the javascript code can be minified.
Can you direct me to some docs for object URLs? I tried searching for it but can't find anything.
2 useless notes
"shit" doesn't sound right (but sounds fine from espeak on the command line)
japanese characters come out as random english characters eg す(su) comes out as "y"
you're fucking awesome :P