AI outperforms humans in speech recognition
techxplore.com
techxplore.com
That humans can still communicate effectively with each other despite a higher transcription error rate is just due to the the way humans do transcriptions: listen to a section of audio, then write down what they heard. If they understood the meaning, they'll likely write down a transcription that retains this meaning, but they might drop an article or duplicate it or switch the order of two words. It's easier for an algorithm to avoid those mistakes of inattention, but harder to prevent errors that change the meaning.
So WER isn't a perfect indicator of transcription quality. It also matters which words are affected. I wouldn't be surprised if people preferred a human-written transcription that reads well despite a high WER over a machine-generated one with fewer errors, but more obvious ones.
Bonus: find the word duplication in this comment.
Part of it is that human communication has something similar to error correction built in. We recognize patterns of words and can guess what the next word is based on that. This allows us to discern the meaning without always understanding every word. We effectively fill in the blanks.
I'm not sure if these models do that.
Gur ercrngrq jbeq vf gur. V unq gb hfr n cebtenz gb svaq vg: crey -yar 'juvyr (/\o(\j+)\o\f+\o\1\o/t) {cevag $1}'
They highlight here the aspect of low latency, so this work might be current SOTA under this condition.
Anyway, yes, we have reached human parity, or even surpassed it, for this dataset, under the same conditions. The conditions are:
* English telephone speech, mostly no overlapping speech (also 2 speakers max), mostly no background noise (although quality is not great either), speakers are directly next to the microphone (telephone).
* You don't know the speakers beforehand, specifically you have never heard their voice or accent beforehand!
Both are crucial points.
Automatic speech recognition (ASR) models are quite good when there is not too much noise, and when there is not too much speech overlap. In this case, we usually have super human performance. But it's very different when you are somewhere in a room not directly next to the microphone, when there is some background noise, or when multiple people speak at the same time. Then we have not really reached human performance yet.
Also, the condition about not having heard the speaker beforehand is special, and maybe artificial. Test that for yourself. You will see that your recognition rate gets quite bad. But once you adapt to it, you will get much better. ASR also can do adaptation, but I think it's not as good at it, i.e. it will get some improvement, but not as much as a human.
Humans can cope with errors in transcription/recognition by using their comprehension
In so far as we can understand people with different accents it's because we have been trained on them. Even if they are not common around us we've had some exposure, from occasional visitors, travels, or media. When we hear an accent we've really never been exposed to we aren't likely to understand it. A good example is foreign speakers trying to speak our native language... even if they've learned our language in school for years, their even slightly off pronunciation can make it very difficult to understand what they are saying.
So the benchmarks say how well model X does on this exact transcription taks given this exact training data, and no other knowledge.
Even basic things, like female/male voices in train vs test set don't match.
I’m surprised that the human error rate is that high; is this the error rate from native to second language, or from second language to native? And what is the average second language skill level of the translators?
(My translation accuracy feels like it would be about 50-75% of words in freeform outdoors conversation at normal speed, but because of this I wouldn’t even consider a course taught in German at this time).
However, it's worth remembering that this refers only to the transcription accuracy, i.e. mapping sound to a particular word in a permanent record (e.g. a written word, in a human transcriber's case). Humans can cope with error rates around 5% (and even higher) when recognizing words because we can use context to deduce the intended meaning, or ask for clarification when the word doesn't match the context. AI has some way to go until there.
This isn't to say this is useless, quite the contrary. Being able to transcribe words in conversational language is hugely useful (for both nefarious and non-nefarious purposes). It's just that it refers to understanding the words, not the sentences :).
But I would like to see where the 5.5% figure comes from.
A lot of communication happens non-verbally. A lot of redundancy in language means that parsing the exact word that was said doesn't matter: the meaning is still understood.
50% might mean you are only picking up common functional words, 75% means you are mostly getting the gist of the conversation, but 95% is really plenty to follow everything that's going on.
Also "Covid", for example, doesn't seem to be in the vocab.