So for voice encryption, you need to obscure all that through artificial jitter and noise, and lack of compression in strategic places. It is a complex topic and I'm not sure the science is settled beyond "skipping compression helps".
There is something to be said about the volume of data being sent during silence; but, uncompressed audio shouldn't have that problem; and, for audio that does, there could be filler data to maintain a given bitrate across the line.
Really transmitting the zeroes as they are is just transmitting uncompressed PCM, which is the trivial solution. Adding filler is undoing most of the compression. The hard part is to add just enough filler, jitter and confusion for an attacker to be sufficiently blinded while maintaining an acceptable compression ratio.
Really? Would I not just get a not quite constant stream of (unencrypted) data that‘s small enough to send at a low bandwidth? And when that data arrives at less than the maximum bandwidth of the channel that I actually use, I just add some filler. And then I encrypt that now really constant stream of data.
So, no. The filler is part of the protocol and is not undoing any of the PCM compression; silence would let you compress the stream more than the plain voice codec would - and THAT would interact with encryption. But that’s quite unusual for real time systems.
Yes, one often does CBR, but even there, variable difficulty of compression often produces variable jitter, which can then be used to infer information about the plaintext stream. There are constant runtime CBR codecs, but one has to take care to use them.
So e.g. maybe you're doing 48kbps, the codec would still consume 48000 bits per second for silence. However if a radio link layer indicates it's struggling to move 48000 bits per second that "adaptive mode" audio codec could shift down to 36kbps instead. This is likely to be a much better user experience than throwing away 25% of your compressed audio due to packet loss and then trying to reconstruct it.
For something like Opus you can gracefully degrade this way from 48kbps (transparent for voice) to 8kbps (non-transparent but easily understandable) without an eavesdropper learning anything about the content, they only get insight into whether you've got link trouble.
Says if you don’t and simply XOR an audio with PRNG output, the resultant entropy will not be constant and transmission sounds like a noisy radio. Something like that.
I haven't tried, but I'd think encrypting and decrypting a packet would take less than 0.1ms, no?
What's the variation in the structure of the data? Wouldn't you just encrypt and send a fixed-length interval of audio each period?
Historically voice encryption was politically only meant for state use, with strict controls, and us plebs not getting any voice encryption or very weak encryption only. Compared to encryption on the internet, this state has persisted for longer in communications. Even in new communication standards the options for encryption generally offer weak/irrelevant security for modern standards (end-to-end encryption).
What's the variation in the structure of the data? Wouldn't you just encrypt and send a fixed-length interval of audio each period?
check out this article https://apkwind.com/netflix-mod-apk/