A CC-By Open-Source TTS Model with Voice Cloning
huggingface.co
huggingface.co
{Angry}What have you done?
{Suprised}Me, I did nothing?
{Shouting}Who else do you think I'm talking to?
{Sad}Why are you always shouting at me?
I wonder if this would also work for "Character" tags, like {Susan}How was your day?
{Peter}I had a great day.
That would open great new ways of having audio books read by cloned voices - switching between characters with the same voice like often done by the real narratorsEspecially with TTS in a language other than English (but also with English), the pronunciation of certain words is sometimes jarringly wrong. Until TTS systems can compensate for this themselves, it would be great if it were possible for humans to use such tags to hint the system to pronounce better. Even if you can't specify the exact correction, but the TTS would just generate a 'different' sound, that could help.
Here's how it's done: https://youtu.be/ASFoTNpkM8o?t=992
Can't wait F5-tts to support the german language. Do you know wether this is planned in the near future?
I've also found a couple of the ESPNet TTS models are decent. I've exported those models to ONNX to make them easier to use.
For what it's worth, here is a list of models that cover what I've worked on in the "Open models" TTS space.
https://huggingface.co/collections/NeuML/text-to-speech-tts-...
Why is good TTS so expensive and why are there no good open source options? Is it just from the need for high quality training data? I don't imagine these models are more expensive to run compared to SOTA LLMs, yet they cost so much more.
in other words, while FOSS TTS lags behind commercial options, it does get better and i expect within a few years it will produce results that are at least as good as the commercial options today if not fully caught up.
https://rhasspy.github.io/piper-samples/samples/en/en_GB/ala...
Of all the TTS APIs I have tried, I like OpenAI voices the best. Haven't considered things like elevenlabs because I find them ridiculously expensive.
I love voice to voice interfaces, but only when they sound natural to my ears, and the current pricing for good ones is prohibitive for a huge number of use cases.
Eleven Labs is most likely trained on stolen audiobooks, they've published a few Youtube videos in Polish, now taken down, of AI renditions of famous Polish audiobook narrators. This was all before they became popular, and before their voice cloning models were publicly available I think.
That probably explains a lot. I've tried listening to some of those audiobooks - very hit and miss, mostly miss. Definitely amateur hour and mostly bad quality.
The dataset and video tutorials are all available and linked on (also english):
https://github.com/neonbjb/tortoise-tts
It supports voice cloning, but I am indeed having trouble getting docker container working and the command line docs are not perfect:
https://github.com/neonbjb/tortoise-tts/blob/1e061bc6752f05b...