For the Mandarin system the human performance was obtained from people in our office, not random Turkers. 4% WER for a group of 5 humans vs 3.7% for the system.
Perhaps you're raising the bar - human level performance no longer consists of an average or median level, but the top 1% or better. I'm not sure that's fair.
This makes me ask two questions: #1- Do systems like this need the court-reported word recognition rate in order to be useful? Or can they compensate for mistakes by using the context? #2- Could we improve these systems by also feeding them video of the speakers lips?
Maybe I should go do a masters to figure out the answers.
https://www.uni-ulm.de/fileadmin/website_uni_ulm/allgemein/2...
Noisy environments are exactly when seeing the lips is a huge deal for me. I have a friend who has a tendency to absent-mindedly place his hand in front of his mouth. In a quiet office or home, no issue. In a bar? He pressed mute, as far as I'm concerned.
The current task they are performing is extremely difficult: I think it is analogous to having a listener receive anonymous random phone calls from people they have never heard before, speaking in a random accent, in a difficult/noisy environment, who speak for a few seconds about a random topic including proper nouns that the listener has never heard before and then promptly hang up and the listener is asked for an exact transcription with no mistakes.
I don't think it is surprising that the error rates produced by Mechanical Turk workers seem high for some of these tasks, and it actually seems like a big accomplishment that a speech recognition system can do nearly as well under such difficult conditions.
Context about the speaker or the current subject of conversation would clearly help, as would visual cues, and the ability to ask the user to repeat or clarify something.
But I guess Mechanical Turk workers are probably not very committed to getting every detail right (for what they're payed I guess I wouldn't worry too much).
Surely, you can get a good transcription from a crappy audio recording, but that's going to cost you more.