Not saying this one is poor, but I recently surveyed all of the text to speech models on hugging face, and they were not usable. There is a massive gulf between accuracy from the main cloud providers text to speech API and hugging face
For now we only host models backed by external libraries on the HF Hub, but we've started working on native TTS for the HF Transformers library, which will hopefully be on par with the paid alternatives! :)
By any chance, are you aware of any good surveys comparing the open models to the cloud providers for other NLP tasks like speech recognition or POS tagging?
Reminds me of Apple Live Text/Google Lens vs tesseract… Sad.