Definitely an ambitious project. Is there sample videos to evaluate?
I only skimmed this but I'm curious if this also kills the need for a voice actor or if it requires sample voice data from a native speaker to sound good (i.e. is it training from a corpus of attempts of the original actor trying to speak that language or does it just need some samples of different phonemes to train from?).