Like voice snippets? We already have them. People don't want to use them.
> Or the death of the audiobook business? Any book read in any voice you want.
Press F to doubt. Human voice inflection is hard to mimic because you have to have a contextual understanding of what is being said, and what has been said in the story up to that time frame. No TTS model is capable of that today, and probably not for a long time.
> Or maybe a form of extreme compression, voice is converted to text with Whisper, sent over the wire as text, and re-created with the same voice on the receiver.
I don't understand the value prop here. The number of cases where you have access to the necessary computational resources, but _not_ adequate bandwidth is so small.
The most likely use case is probably scamming. With a small snippet of someone's voice (for example, from answering a robo call) you can now synthesize a completely reasonable sounding phone call for conning people out of their money. By the way, this attack vector was _highly_ effective against the elderly even when the attackers voice sounded quite distinct from the individual that they were posing as. The elderly have no chance against this type of fraud. There is going to be big money in authenticating individuals so that phone calls, etc can happen between trusted parties.