Also, major issue with this kind of research is that they combined several systems in order to get best results. Most practical systems don't use combinations, they are too slow.
Also, major issue with this kind of research is that they combined several systems in order to get best results. Most practical systems don't use combinations, they are too slow.
Also I'm note sure an error reduction of 20% (1-6.3/7.8) is to be considered small; depends on the particular challenge really. Like, sentiment analysis only starts to get interesting above 80% on some dataset, as much can be guessed correctly in very naive ways..
Human lvl on this task is estimated to be ~4% so we have quite a lot of ground to cover still..
In Microsoft paper http://arxiv.org/pdf/1609.03528v1.pdf Table 5 it's a line "Povey et al. [19] LSTM". http://www.isca-speech.org/archive/Interspeech_2016/pdfs/059... Interpolate it to RNNLM column and you'll get 7.8. See also http://www.isca-speech.org/archive/Interspeech_2016/pdfs/047...