Large-scale multilingual audio visual dubbing
arxiv.org
arxiv.org
I only skimmed this but I'm curious if this also kills the need for a voice actor or if it requires sample voice data from a native speaker to sound good (i.e. is it training from a corpus of attempts of the original actor trying to speak that language or does it just need some samples of different phonemes to train from?).