I'd love to read that if you ever put it together!
How many hours of audio did it need to train one of the voices?
How many hours of audio did it need to train one of the voices?
The key to training is that all of the models were transfer learned from the Linda Johnson speech dataset (LJS).