Part of the problem is that phone lines are crap for voice quality. If you're on GSM, you're using the AMR codec which - for the time - was brilliant. But it is incredibly low bandwidth. And, effectively, it synthesises speech rather than just transmits it.
So the algorithm is not getting a pure waveform and analysing that - it is looking at a partially reconstructed digital simulacrum.
The pitch that we received was about detecting stress in the human voice so call centre handlers could tell if they were speaking to a fraudster. It didn't work. Oh, sure, you can make a reasonable prediction by listening to pauses and pitches - but that doesn't account for the fact that calling up and being on hold for ages in inherently stressful.
I'm not surprised this system was fooled by AI. But I am surprised anyone bought it in the first place.