The video mentions that the audio was recorded separately, and shows a few other options like text-to-speech (which obviously doesn't match the voice) and some smarter voice matching audio generation (VoCo) which could pass for the original voice sent over heavily compressed, low-bandwidth video conferencing or something like that. I'm guessing that if this is used for actual disinformation, finding a voice actor/audio engineer to try and match the speech would be most effective.
Most of the examples in the video they had the subject record the audio separately from the generated video. In one or two of them they cite some audio-generating thing.