Coqui.ai TTS: A Deep Learning Toolkit for Text-to-Speech
github.com
github.com
Personally I prefer StyleTTS 2, and it has a better license. But XTTSv2 has a streaming mode with pretty low latency which is nice. I did run into hallucination issues though. It will hallucinate nonsense words or insert extra syllables in words, pretty frequently.
As others mentioned they shut down so there won't be any updates to XTTS.
What is the situation exactly? Seems to be licensed MPL at a glance, so you're able to use it in commercial projects.
The main problem is quality, Eleven labs is so far ahead, even though their API is not very flexible.
Meta's Voicebox is the only other option that feels realistic - but it's for research only for now.
TLDR: Making money from open-source is hard.
Audio samples are easily obtained from their podcast, but manual data labeling is painful for a hobby activity. Further, from what I understand, the real difficulty in performant diarizer models is not speaker recognition generally, but specifically speaker recognition while there is overlapping speech between multiple speakers. I am not even sure how to best implement a labeling procedure for segments with overlapping speech.
I started to wonder whether I might bootstrap a decent sample by leveraging TTS vocal cloning models to simulate the five speakers in dialogues with overlapping speech segments. So I ask HN, is this hopelessly naive, or potentially useful technique? Also, any other advice?
[1] https://www.3d6downtheline.com/ [2] https://github.com/MahmoudAshraf97/whisper-diarization/
We tend to agree, the time for just one company to be seriously doing speech is over. It needs to be more diverse, and needs to be opensource https://github.com/Camb-ai/MARS5-TTS
CUDA_VISIBLE_DEVICES="0" python TTS/server/server.py --model_name tts_models/en/vctk/vits --use_cuda True
def play_sound(response):
#learning : you have to use a semaphore to serialize calls to winsound.PlaySound(), which freaks out with "Failed to play sound" if you try to play 2 clips at once
semaphore.acquire()
try:
winsound.PlaySound(response.content, winsound.SND_MEMORY | winsound.SND_NOSTOP)
finally:
# Always release the permit, even if PlaySound raises an exception
semaphore.release()I've been using Dimio's Speech for a decade now, but it seems silly now that much better voices exist.