Whither Speech Recognition? (1969) [pdf]
pdfs.semanticscholar.org
pdfs.semanticscholar.org
> How little we actually hear, when we listen to speech, we realize when we go to a foreign theatre - for there what troubles us is not so much that we cannot understand what the actors say as that we cannot hear their words.
As a non-native speaker of English, I started watching English movies with subtitles when I was a teenager. This had an interesting effect: after a few years of doing this, I am now used to knowing each word that is spoken in a movie exactly, on its own - after all, it is clearly printed on the screen.
I now get nervous watching movies in my native language (German) without subtitles, simply because I am not able to extract each word precisely. Somehow I trained myself to expect an exact "acoustic" understanding from movies, as opposed to a "semantic" understanding. It is incredibly how the human brain is able to extract the meaning of a spoken sentence by context, facial expressions and gesture, even if we only understand half of the sentence acoustically.
So in retrospect I'd guess the level of funding at the time was closer to right than this critique, even if most of the work was flimflam. (The author wrote a popular book about information theory which I liked, so this is disappointing.)
When am I going to wake up and be able to dictate into my phone and make corrections with my voice? We must be close.
I don’t mind the mistakes made when dictating but having to pull up the keyboard takes away from the “magic“.
Dictating with corrections is a UI issue. I already dictate messages to my phone running Android Auto, and it confirms the entire message. To make corrections, I have to say the entire message again. It would be better if I could just restate the part that was heard wrong and let the application figure out which part to replace, and this is entirely doable today with the available building blocks. Just don't expect a Google product manager to figure that out.
Going from text stream of phonemes (“mmaaaayyynaammeezzzbahhhb”) to text stream of sequence of words (“my name is bob”)is vastly more limited.
One of them is speech recognition, the other is language modeling.
language comprehension? From experience, it’s cheaper and faster to get results by making a new human :)