Reviving the 1973 Unix text to voice translator
spinellis.gr
spinellis.gr
Radio Shack used to sell the SPO256 text to speech chip. I had a bunch of them and the speech was barely understandable. It was worse than Talking Sam on the C-64.
There’s a recording of what the SPO256 sounded like on its Wikipedia page:
Yes, it is certainly difficult to understand the sound. However, all these years later, the same could be said of remote video connections in which speech is sometimes replaced with undecipherable squeaks and echos.
I'm surprised that zoom et al. have not developed a scheme to show text during these squeaking times. After all, any meeting has people and computers that could detect a temporary problem, mandating a temporary switch to a dictation mode.
In the same video, there's a fun moment where they demonstrate constructing a simple spell checker out of unix primitives, and the demonstrator comments that "unix" is a word unlikely to ever appear in the dictionary: https://youtu.be/XvDZLjaCJuw?t=551
Of course, now it's in every dictionary!
Is there any better TTS Tool people could recommend?
If/when that happens, and you go to look at the rest of the tools... They're pretty crap. Still.
I did experiment with Google's API at the beginning, the audio would contain recurring artefacts that made it painfully obvious it wasn't real. If I had to describe it, I'd say "50ms of the sound your computer makes when it hangs while playing an audio file".
No such problems with Azure. Zero artefacts and nice, crisp, natural enunciation, though unfortunately it is still quite unreliable when synthesizing single words (as compared with full sentences). You can actually test it out by using the "Read Aloud" feature inside Edge.
I don't know how well this translates to English, though. I definitely enjoy the benefit of working with a gender/language combination (female/Chinese) that's received the largest total investment of effort and resources. Many well-resourced Chinese companies (banks, tech services, gaming, etc) all use similar synthesized female voices in their phone/online/B&M services.
I took a shot at making a modernized port of speak.c myself not long after the code was found, but it didn't get very far, sadly. I couldn't figure out an easy way of dealing with the multiple-character character constants.