All the speech recognition engines I've interacted with so far were awful. Not just bad, awful.
>Collecting such data sets could be very difficult and prohibitively expensive.
Uh, movie subtitles?
All the speech recognition engines I've interacted with so far were awful. Not just bad, awful.
>Collecting such data sets could be very difficult and prohibitively expensive.
Uh, movie subtitles?
Movie subtitles have poor alignment, usually contain multiple speakers (sometimes talking at the same time) and often contain sounds or other things which are not dialogue. Cleaning this is expensive, probably much more than just getting transcriptions of single speaker samples. It corresponds much closer to "real human life" but that is not where papers are published, unfortunately.
We have thousands of hours of ebooks (librivox - librispeech is a 1K hour subset used by many) that are used in open source speech recognition systems, and Baidu has a direct line on many more hours than that.
The better than human line is (in almost every paper - Baidu is not alone in this at all!) bullshit though - while the system is quite good, they only compare to humans given a fragment of a statement without context. Humans are very, very good at inferring given context (would you recognize speech the same way at a funeral/dinner party/college lecture/hospital?), and this is where most models fall flat.
Conditional language models can help this to an extent, but human adaptation and the ability to generalize will take a while to catch up with.
So, deep learning to clearing background noises?
But rolling out a paid service like that comes with a lot of expectations that being part of a free app doesn't.
Movie subtitles are very rarely actual transcriptions of what is spoken; they are instead summaries, edited for brevity and quick comprehension.
I don't know much about what kind of corpora is required for training this kind of model, but subtitles don't seem appropriate.
And for reasons that have always escaped me, subtitles frequently just don't say the same thing the soundtrack does.
It's a lot better than nothing, but subtitles make for a pretty low-quality dataset.