I've been a long-time user for OS X's built in text-to-speech [0].
IMO, it actually performs pretty robustly on these examples. Is Apple still using just diphones, or are they post-processing in some way?
[0]: https://en.wikipedia.org/wiki/PlainTalk#Text-to-speech_in_Ma...