The best implementations make an advantage in markets so they are well guarded. We do not know about the best implementations, because we did not notice a thing. For example some phone operators have replaced their customer services with TTS / STT solutions. Because people tend to lock up when they realize they are talking with a computer, they have had to make those systems sound very natural.
I know a few are pretty crappy ones, but a few are plain spooky. Customers that tend to joke and flirt with the computer, hoping for an emotional response, probably are most likely to notice them.
Then there's the case where the US intelligence services demonstrated their capabilities to a politician (senator/congressman) by recording him, and producing a voice clip of him saying something like "death to america" so well that no one could distinguish the speaker. It also seemingly passed further voice analysis. Google it up, pretty interesting read.
I got a call two years ago from a telemarketer that kept asking me yes/no questions with a robot voice and didn't leave me any space in the conversation to say anything. I really felt like I was talking to a computer so I tried to Turing-test him by forcing him to answer an open-ended question. It took two minutes, and the caller turned out to be human after all. That wasn't very relieving though. Humans should not speak with such zombified voice.
Why. Not. Let. The. Telemarketer. Have. Some. Fun. At. Their. Crappy. Job.
Festival can produce very good results, but you have to get into the weeds with it and do a lot of planning. If you don't like Scheme, you're not going to like dealing with festival, and if you do like scheme, you're going to be frustrated by their particular scheme.
I've also noticed a lot of folks looking at post filtering, but that is a sea of dragons at the moment.
It will work better on the ZipitZ2 and pentium2 machines.
Enough to listen short stories, but it's not suited to listen a full book from you bed.
With a small bit of work, you can have festival use an HTS voice directly. I usually use festival to generate my label files, post filter the phoneme timing, synthesize with HTS, then apply a post-filter.
http://espeak.sourceforge.net/
http://simulationcorner.net/index.php?page=sam
I have to admit, that the quality is not good. But it also works on embedded systems.
Can't confirm how well espeak works, since I didn't try it, but I have tried pyttsx on Windows. (pyttsx is a Python library which uses SAPI5 on Windows and espeak on Linux.)
Worked okay for me. Not tested a lot though. Here's my simple Python snippet for trying pyttsx on Windows:
http://code.activestate.com/recipes/578839-python-text-to-sp...
(A commenter on the above ActiveState recipe said my pyttsx recipe worked fine for him with espeak on Crux Linux.)
and more details on the same here:
http://jugad2.blogspot.in/2014/03/speech-synthesis-in-python...
"The pyttsx library is a cross-platform wrapper that supports the native text-to-speech libraries of Windows and Linux at least, using SAPI5 on Windows and eSpeak on Linux."
Also, check out the synthesized train announcer's voice in my blog post above - somewhat eerie and cool :)
http://jugad2.blogspot.in/2014/03/speech-recognition-with-py...
Some better free-software offerings would really open this area up for experimentation, since the commercial offerings tend to be pretty black-box and targeted at very specific use-cases.