NATSpeech: High Quality Text-to-Speech Implementation with HuggingFace Demo
github.com
github.com
As I switched from macOS to arch linux, I was looking for a good text-to-speech option, cross platform and open-source. (I use it to get an first overview over larger, boring papers I'm reviewing ).
pretty impressed by larynx so far: https://github.com/rhasspy/larynx
It seems also quite hackable.
Will also play with NATSpeech.
You will find a lot of code but that is research code, i.e. you need to understand the science behind it and it needs some effort to get it to run, and it will have lots of hard edges.
When you ask for a usable cross platform and open source solution, this is different. Although, Larynx looks actually nice and usable, i.e. intended for the end user.
The hard part of TTS systems is also to get all the edge cases right, like special symbols, abbreviations, special names, etc. And most research code does not care too much about this but is more about doing research such as developing new models.
A manual lexicon with lots of custom rules is often used to cover all these cases but to maintain this is a lot of work and depends also on the use case. Or you could train another model for this part which will somehow work better than a simplistic lexicon but will still fail at many edge cases.
I plan to release a version of Larynx that uses eSpeak for phonemization, since it covers many of those corner cases.
I also have a Twitter account (@rhasspy) used almost exclusively for announcements.
Also, from The FAQ:
> Who/what is behind this project? > > I've had a couple people and institutions provide a bit of help along the way, but I am — and always have been — one person.
Never underestimate the intellect of a single person. Every corporate team, every open source maintainer, every researcher more or less* has access to the same prior knowledge to build something new.
* some have more money to burn, some are surrounded by mentors/helpers
I actually think Microsoft's TTS is the best and has many voices. IBM's Watson branded voice was a close second. But as near as I can tell, there's no way to run these locally, only through their cloud interface.
I'd say this portaspeech one was the best local I've heard but it annoyingly trips on words like "shouldn't".
https://en.wikipedia.org/wiki/Currah#Currah_Microspeech_for_...
That answer is basically a universal "nope" imho. The vocal inflections are contextual and subtle. The problem is hard. That's why it's the test