To explore that analogy: In the case of lossy audio compression, the compressor deliberately introduces quantization noise into the signal. It does so by running a "psychoacoustic model", which attempts to capture a broad quality of human hearing called auditory masking. There's a number of different kinds of masking[1]: a strong tonal sound creates an "umbrella" across nearby frequencies that can mask quieter noise-like sounds. Similarly there's noise-vs-noise masking, as well as forwards- and backwards- temporal masking. (Yes, backwards. A sound can mask perception of a sound that occurred before it.)
In the audio compression case, we've built an algorithm that attempts to characterize exploitable phenomena of human hearing. These masking characteristics aren't perceived quite the same by human listeners as the model, nor even the same between individual human listeners. Thus the need for human listening tests. These, due to the experimental care and human subjects required, are expensive.
Back to speech transcription. Say the end goal is "how well does a human comprehend this transcribed speech"? (e.g. vs some standard, such as the original speech, vs. the original speech transcribed by a skilled specialist, etc.) The problem starts to look pretty similar. We can cite numerical, word-centric error rates, but that fails to capture how well meaning is preserved and transmitted. Imagine a perverse algorithm that did a perfect transcription, but then dropped or altered words for maximum meaning obfuscation. It might equal or even beat the cited error rates but be much harder to actually comprehend.
These classic signal processing analyses tell you nothing about the correct operation of a psychoacoustic model codec design (i.e. MP3, AAC, Vorbis, etc.) An analysis of A (original signal) vs. A' (signal passed through a compression-decompression cycle) is just extracting the quantization noise introduced by the codec. That provides no information about how effective the codec was in masking that noise with the original signal content.
To illustrate, imagine a perversely designed codec: it runs two models, the first "good model" is a normal psychoacoustic model. The second "bad model" is the one the compressor uses: it applies the total amount of quantization noise allowed by the good model, but applies it in ways that are maximally annoying to human listeners. This isn't just avoiding masking, it's things like using noise correlated to the original signal, which is generally more obtrusive than uncorrelated noise, etc.
A codec using just the "good model" and one using the "good + bad model" would have (by definition) exactly the same introduced noise, but the latter would sound FAR worse to a human listener.
No joy.
I can't really follow your thinking above; sorry.
In audio, there is distortion, which is correlated with the signal. Noise is uncorrelated. Codec error would seem very much to me to be at least much more correlated than uncorrelated. MP3 artifacts sound at least more "phase-ey" than they sound anything like quantitization noise ( at least to me ). This may be because I have heard badly aligned tape machines make that sort of error happen. "Phase-ey" also triggers my (feeble) mind into thinking about allpass filters as the model for the error.
In terms of intelligibility, it's possible to improve intelligibility by adding noise, and by adding clipping - aviation comms does this at times. What destroys intelligibility is phenomes being destroyed by phase changes and bad amplitide errors ( where there are actually good amplitude errors ).
I've run an ABACUS voice quality analyzer several times, and I don't think it's purely a distortion analyzer - adding clipping at least can improve MOS/PESQ score surprisingly. Even more surprisingly, there's no mechanism available for calibrating gain staging on one.
If I understand you correctly, your discussion around intelligibility primarily applies to signal processing of voice vs general audio. Codecs such as MP3, AAC, etc. don't have the luxury of making assumptions about the signal content, and so aren't designed along those principles. E.g. speech codecs can generally run well at much lower bitrates than general audio codecs because they operate on a constrained domain of audio (i.e. speech).
Regarding distortion management, see Rate-distortion optimization[1] for lossy audio codecs: where the purpose is to manage distortion within the limits of the bit rate supported by a communication channel or storage medium.
[1] https://en.wikipedia.org/wiki/Quantization_(signal_processin...
Google Voice transcription often makes easily detectable mistakes that I can mentally correct using context. Smarter, context-aware mistakes might be worse, if they are also more difficult to detect.
We find that the artificial errors are substantially the same as human ones with one large exception confusions between backchannel words [acknowledgment words like “uh-huh”] and hesitations.
The difference they found, but suspect might be a result of the different transcription guidelines of the training corpus: we see that by far the most common error in the ASR system is the confusion of a hesitation in the reference for a backchannel in the hypothesis. People do not seem to have this problem.
The term 'human parity' refers to adding up all the bits in a human and taking the residue to a convenient modulus to detect human error.