This looks really impressive, do you have any writeup on how you coded all of it?
There are a lot of short cuts I wish I'd known at the outset, mostly in terms of curating training data and monitoring the learning, but I also learned a few new systems-y things from the overall engineering project.
How many hours of audio did it need to train one of the voices?
The key to training is that all of the models were transfer learned from the Linda Johnson speech dataset (LJS).