Work like, say https://arxiv.org/abs/1806.04558 [paper]
Work like, say https://arxiv.org/abs/1806.04558 [paper]
Edit: Sample of ETI-Eloquence at my preferred speed: https://mwcampbell.us/audio/eloquence-sample-2021-09-25.mp3 (yes, it mispronounces "espeak")
Edit 2: To elaborate on what I mean by "mostly dead": In 2009 I was tasked with adding support for ETI-Eloquence to a Windows screen reader I developed. At that time, Nuance was still selling Eloquence to companies like the one I worked for back then. When I got the SDK, the timestamps on the files, particularly the main DLLs, were from 2002. As far as I know, an updated SDK for Windows was never released. I'm thankful for Windows's legendary emphasis on backward compatibility, particularly compared to Apple platforms and even Android.
Finally, a sample of espeak-ng (in the NVDA screen reader) at my preferred speed: https://mwcampbell.us/audio/espeak-ng-sample-2021-09-25.mp3 I use the default British pronunciation even though I'm American, because the American pronunciation is noticeably off.
This is exactly the speech synthesizer I use daily. I've gotten so used to it over the years that switching away from it is painful. On Apple platforms, though, using it is not an option. So I use Karen. Used to use Alex, but Karen appears to be slightly more responsive and tries to do less human stuff when reading. Responsiveness is a very important factor, actually. Probably more so than people might realize. Eloquence and ESpeak react pretty much instantly whereas other voices might take 100 MS or so. This is a very big deal for me. Just like how one would like instant visual feedback on their screen, it's the same for me with speech. The less latency, the better. My problem with ESpeak is that it sounds very rough and metallic whereas Eloquence has a much warmer sound to it. I pitch mine down slightly to get an even warmer sound. Being pleasant on the ears is super important if you listen to the thing many, many hours a day.
[0] https://github.com/RHVoice/RHVoice
[1] https://rhvoice.org/en-voices/
[2] https://f-droid.org/en/packages/com.github.olga_yakovleva.rh...
https://youtu.be/MmcLFJQpv2o?t=85
Edit: or on the online demo; select "HMM-based method (HTS 2011) - Combilex" > "SLT (English American female)".
There are of course great benefits to something simple to use. I remember cross-compiling flite to run on a custom android/windows/linux project to generate voice lines intended for a in-game robot companion (nothing came of it though) based on SDL. It probably would not be nearly as feasible to do the same for some dependency-heavy machine learning library.
Now, I haven't done any research to find better examples of projects. I was just surprised how identical the article describes the options, to what was available 12 years ago.
https://aws.amazon.com/polly/ https://www.youtube.com/watch?v=00D0YZ9GQX4
Either we're just not there yet technologically (hard to believe), or there isn't a will to make good speech synthesis available to commoners.
[1]: https://www.youtube.com/watch?v=hDVuh4A-q3Q&ab_channel=Vocal...