Here, take a look at this snippet:
https://youtu.be/eEXvMOJ9ps0?t=66
They play 4 clips, 2 of them human, 2 of them AI generated. Can you tell which ones are which?
And the kicker is, this works in real-time (AFAIR), and it doesn't even use the GPU (it's CPU-only), and generates pitch-correct speech (for Japanese). It's not even funny how far ahead they are. And you can buy it right now.
AFAIK they use some sort of a hybrid method with a bunch of custom modeling DSP code around it (they've been doing speech synthesis for over a decade) plus a neural network. One mistake that essentially all of the western TTS models seem to make is that they use only a neural network, without augmenting it with non neural network code, which (from what I can see) is the secret sauce to make a fast and good sounding TTS work.
Such has been the source of much miscommunication online.
Amateur fiction writing, which tends to overemphasize how things are said ("I guess I can go rescue your cat", the exasperated detective said wearily) might be easier for AI!
The linked page actually has examples of the same text being read with different emotions, demonstrating that for even a single sentence a lot of variance is possible.
It was used to do the fake joe rogan/steve jobs podcast: https://podcast.ai/
You can use the API to read books/articles aloud in real-time, but it is quite expensive after the free trial.
If I tried the wrong thing can you provide a link? I’d like to be amazed.
I used to work there; great team behind the product!