The innovation is spectacular, BUT, there needs to be a signal/low pitch sound that denotes this is generated audio in every single generated sample (likely legally enforced), or your grandma and kids will soon be getting legitimate sounding voice calls from you after someone calls/visits/interviews you first to record and train a model on your voice (as a simplest potential abuse vector, celebrities and anyone with public voice samples would be even easier).