I don't know what to say, your demo is really bad. I cannot understand what the person was speaking after you altered the sound, this approach doesn't make sense to me.
the major call providers have settings one can use as well
The noises confuse the ASR models, I presume.
Someone could still record audio and manually transcribe it or take notes with a pen and pencil... just sayin'. Some of this just comes down to having to trust the other party.
I do take some issue with my voice almost certainly being used to train models when using certain services.