TTS: Text-to-Speech for All
github.com
github.com
https://github.com/coqui-ai/TTS
Mozilla TTS is not maintained anymore (at least ATM).
Disclaimer: I've created both of the projects.
Do you support multiple speakers?
Also, do you mind if I email you and ask you a few questions about Coqui TTS?
Come and join our gitter room.
Any idea if these issues still exist? Thank you.
Check our latest work https://edresson.github.io/SC-GlowTTS/
You can also check other released models here
We've come a long from from SAM and my Amiga, I tell ya.
Maybe this new organization can accomplish the goal of easy and open trainable TTS. I'd really like to see it.
Good job with the stuff you've built, and good luck!
https://www.charlieharrington.com/flow-and-creative-computin...
And the podcast RSS feed:
https://whatrocks.github.io/castellan/podcastjr.xml
It's great when these ML models link to a Google Colab notebook. It makes it super easy and dare-I-say fun to try them out.
- https://github.com/snakers4/silero-models#text-to-speech - https://habr.com/ru/post/549482/
Disclaimer, this is my independent project
I feel the biggest causality of online advertisement is accessibility, Those with eyes (eye-sight) are more valuable to the mega corps than those without and so Internet is full of rich graphics; making the lives of those without proper vision miserable.
e.g. News apps should have a play button on the top both for accessibility and convenience. Here's a discussion[1] regarding that on my problem validation platform with participation of someone with accessibility needs.
[1] https://needgap.com/problems/200-human-voice-summary-of-news...
It was originally based on Mozilla TTS, but I've since moved to exporting models to Onnx for speed.
[1]: https://soundcloud.com/user-565970875/pocket-article-wavernn...
There are two ways to avoid this: 1. predict 8 high (coarse) bits, 8 low (fine) bits separately as in the original waveRNN paper. 2. use a mixture of logistic distributions as the predictive output as in the recent Lyra vocoder from Google.
Specifically, how much slower this would be if the audio was, say, 10 bits?
I recall a lab exercise in college where we were supposed to increase the resolution of a quantizer until we reached a decent tone and 10 bits were the point at which we reached satisfying quality.
I suggest you the check the latest uploads on soundcloud.
Disclaimer, this is my independent project
You can check out the released models page for the other models and languages.
We are working on automatically extracting some insights for the user and using NLP to present them like news articles.
It wouldn't take a huge lift from that to use TTS to provide another way for user to digest the data.
Would make for a cool demo but wonder how sticky it would be.
Also, Mean Opinion Score (MOS)[1] tests[2] of the TTS engine have it coming out ahead of other commercial engines[3].
[0] https://erogol.github.io/ddc-samples/
[1] https://en.wikipedia.org/wiki/Mean_opinion_score
Edit: Oh, I see this project uses Espeak. Interesting.