Impressive! I guess the speech synthesis quality is the best available open source at the moment?
The endgame of this is surely a continuously running wave to wave model with no text tokens at all? Or at least none in the main path.
The endgame of this is surely a continuously running wave to wave model with no text tokens at all? Or at least none in the main path.