https://www.npr.org/2023/03/22/1165448073/voice-clones-ai-sc...
https://www.forbes.com/sites/thomasbrewster/2021/10/14/huge-...
My thinking is, you might as well get in the habit of defending yourself now. The alternative is to monitor how widespread the attack is and only adjust your policy once it becomes "sufficiently" widespread. But I don't think that's even a labor savings, since defending yourself isn't actually that hard.
With the demos we've seen feels absolutely doable, but for now requires quite some effort.
But even tampering seems pretty easy if the attacker has a more modest objective, of having you and your buddy each talking to one of the attacker's henchmen using voice changers. The emoji verification won't help here -- each henchman just gives the emoji for their respective conversation.
I feel it is implied that the latency is low enough (a few 100s of ms) to not impede the conversation, and that the parties have talked before and would notice if the tone of the conversation was completely different. Or maybe I'm misunderstanding.
To defend, could ask to verify emojis at a random point in the middle of the call to make the attacker's life more difficult. Especially right before discussing sensitive information ;-)
Or drip verify over the course of the call, e.g. "what's your 3rd emoji?", and listen for signs of an attacker cutting in and out.
Recently the French Government required its members to use it, but it's made by a French startup, AFAIK.