Robocalls -> I want to know I’m speaking to a robot
Audio books -> reasonable. An accurate tone is pleasant
Singing -> ever heard of vocaloids? They’ve existed for at least a decade or two
That it was technically already possible does not mean there isn't benefit from improved quality. In fact, Vocaloid itself has been improving and now uses AI.
Would also add making movies, podcasts, news broadcasts, etc. available automatically in a huge range of languages. You wouldn't want movies dubbed by Microsoft Sam (beyond initial comedic effect).
You'd be surprised how common something like this used to be in Poland, though admittedly we used an Ivona voice for this, which was a lot more pleasant.
Having a single narrator narrate the entire movie, overlaying the original audio track, is already common here, much more so than dubbing or subtitles. This is for historic reasons, in the communist era, obtaining the raw audio tracks for dubbing was often impossible, all the translators often had was a normal copy of the movie in its original language.
In the early 2000's, we had a lot of early / unofficial pirate releases, and they had to be translated into Polish somehow. Subtitles were certainly one method, but as we're all used to the single-narrator style, many people didn't mind listening to a somewhat decent synthetic voice instead.
Commercials/tutorials/corporate training videos -> Voiceover work
TV shows -> Dubbing in various languages
Fast food drive-throughs -> Taking customer orders
E.g. scamming. For anything that is just about conveying information through audio, like voice assistants, traditional TTS already works fine.
Most of us just need a machine to have the ability to speak to us in a way that sounds halfway human, but not as horrible as old open source TTS systems.
My current go to is Piper TTS: https://github.com/rhasspy/piper
It's MIT-licensed, supports ~30 languages and multiple voices[0]/quality/licenses.
Voice output samples: https://rhasspy.github.io/piper-samples/
Discovered just now that recently there's also been some efforts to train Public Domain/CC-BY licensed voices specifically for Piper TTS too: https://brycebeattie.com/files/tts/
---- footnotes ----
[0] e.g. One of the English voice sets has ~900 voices(!).
however if i am going to listen to 3 hours of audiobook, i would still pay for a better TTS versus Piper. (now i buy a lot of audiobooks read by voice actors but not every book i want to audio-read has a market to support a voice actor being paid to read it)
however Piper is definitely going in the correct direction though, huge progress.
A couple of years ago I started development of a tool to help with the generation of game audio such as NPC dialogue, "barks" or narration for those without access to/budget for human voice actors: https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to...
One thing I found interesting is that writing a small "scene" and then hearing dialogue being spoken by a variety of voices often prompted the writing of further lines of dialogue in response to perceived emotion contained in voices in the generated output. Plus it was just fun. :)
The version of the tool on that page is based on Larynx TTS which has continued development more recently as Piper TTS: https://github.com/rhasspy/piper
I'm yet to publish my port which uses Piper TTS though: https://gitlab.com/RancidBacon/larynx-dialogue/-/tree/featur...
Though I did upload some sample output (including some "radio announcer" samples in response to a HN comment :) ): https://rancidbacon.gitlab.io/piper-tts-demos/
Obviously there's variations in voice quality, and ability to control expression is currently limited but beats hearing my own voice. :D
Overall, this tech is a boon for the disability community.