https://github.com/NVIDIA/tacotron2
https://github.com/CorentinJ/Real-Time-Voice-Cloning
https://github.com/mozilla/TTS
I think a lot of the remaining gap is due to a lack of high-quality training data -- most of the open-source models are trained on public-domain audiobooks (e.g. LJ Speech).
However, good training data (large amounts of annotated recordings by professional voice actors) is expensive to create, and unlike code, there's not a tradition of people sharing it.
I’m not really an expert. From what I understand, the “cutting-edge” stuff requires pushing past the point where we are splicing segments of speech together. Splicing segments together is hard enough.
There are a couple open-source efforts like Mozilla’s, but if you want something like Lyrebird, well, that technology isn’t even really productized commercially yet.
We don’t see FOSS pharmaceutical research for instance, I believe for the same reason. The amount of coordination needed and the impossibility to separate TTS projects into sub-parts could also factors.
“Common Voice is Mozilla's initiative to help teach machines how real people speak.”
For something like Sonantic, you need clean recordings from professional actors in proper recording environments (not to mention the in-house expertise to then filter these down to curate the training/test datasets). That costs money. A million people with laptop microphones will just never get there.