Whatever Happened to Voice Recognition?
codinghorror.com
codinghorror.com
MAYBE it is as bad as he describes when you do not train the system and just start speaking. But that is like complaining about Emacs after using it for only 10 minutes. Retarded.
I wouldn't be able to code with it though, but for producing prose in English. It. Actually. Works.
(I recognize that the problems described in the post is about recognition "in the wild" and not with one person in one quiet room with one microphone like I use it)
As have I. The dictation capability built in to Vista / Windows 7 is really quite amazing. (They both also have voice control.)
I'm an extremely fast typer, but dictating - even with time spent correcting mistakes - is quite a bit faster.
I'll be ... yep, tucked well away under Accessories/Ease of Access. Who knew? (Six month old laptop w/ Win7, successor to an XP machine.)
Needs setup and the, so-far unused, built-in mic is not cooperating. It does say it allows text dictation, besides OS voice control.
And try this one on for size: Microsoft Office 2003 came with built-in voice control and dictation, which could also be used in other applications!
"In 2007, reports surfaced that Windows Speech Recognition could be used to remotely access and/or control a user's computer. Theoretically, playing a pre-recorded message containing Windows Speech Recognition commands could allow one to execute tasks on another computer remotely."
Why, that reads like a Wikipedia article about a Microsoft product!
(And, yes, the "security" issue is pretty far-fetched but is factual, I remember it being raised.)
I think the reason people haven't adopted voice transcription is because (1) they don't realize that some of the programs out there have reached the 99% accuracy level, (2) we all have developed the habit of working with a computer with our fingers - breaking habits is hard and (3) -related to 2- our brains have been trained to think as we type in ways that allow us to get certain work done, and if we moved over to getting our work done via voice, there are neural connections we'd have to re-create through training.
For me, writing with a word processor is not just about setting down my words, but even before that, setting down the structure of the essay, then going back into it non-sequentially to flesh it out.
I can't imagine how that would be workable, even if the speech recognition were perfect. Things like "go back to the first paragraph of the section, now after the first sentence insert X" is unwieldy, and bound to take far longer than traditional cursor control with keys or mouse.
Of course we're a long way from that goal. Firstly, people don't understand just how difficult it is to express clearly what it is that you want. I have a good friend that had bought an iPhone whilst she was living abroad for 6 months. On returning back home, she wanted to be able to add new songs to her iPhone without losing the stuff already on there. That's how she described the task to me, as the go-to person for computer problems. So I started asking a bunch of questions - do you still have the same computer that you were using abroad? Do you want copy the music already on the iPhone onto your computer? What about other data on the iPhone, do you need to recover that too, or can I blow it away? She got very frustrated with all of these questions and finally just snapped "Oh, you know what I mean, I just want to be able to use my iPhone normally, including adding new stuff to it!"
Sigh. She really didn't (doesn't) understand that this just isn't enough information for me to work with - and that's me, as a walking talking human being with a strong understanding of what the computer is doing. How is a computer, with problems in transcribing the spoken words, let alone understanding the underlying meaning of those words, and not having any idea of the real world context of the problem, supposed to figure out what she wanted?
In the end they went with an inferior solution from one of the big hardware manufacturers who coincidentally were also mentioned by the legal team as a patent holder.
So maybe real research is being hampered?
I think the key with speech recognition and natural language processing is to forget the HAL9000/StarTrek 'hello computer' nonsense, and focus on constrained domains augmented by other technologies.
Touchpad keyboards are blatantly less usable than hard keyboards but by the time you factor in a) portability and b) auto-completion of words/urls/search-strings, then they become a more than viable alternative.
People have incredibly adaptive linguistic skills, developers should harness this more IMO rather than be overly ambitious with NLP.
That said I was very impressed with the accuracy of Google's voice search app. But like Jeff Atwood says, it's easier just to search normally.
I am not an expert on speech recognition, but I have done my PhD in a speech lab, so I have at least an idea on the issues. Almost every current speech recognizer is divided into two parts: the acoustical modeling (from audio file to "phonemes"), and language models (from "phonemes" to words).
Most people working on acoustical modeling are interested in difficult situations (multi speakers, noisy speech, etc... where performances are nowhere near human in general).
The language modeling part has to deal with sparsity: the idea is to model the probability of getting a work w_n given that the previous words were (w_n-1, w_n-2, ...). Given that vocabularies (numbers of possible words) are of the order of millions for large vocabulary recognition, you can imagine that many (most) combinations are never seen in the training data (trillions and more combinations), so you have to somewhat "smooth" the data, remove early "impossible" combinations., etc... There are a lot of heuristics.
Also, even though the author is a bit off on the timing, he is right that the basic methods are the same for a long time (statistical, data-driven, HMM for the acoustical model). Sure, we now have (somewhat) speaker-independent models so that you don't have to train the model with your own voice for hours anymore, and language models can handle large vocabularies, but the basics are really the same. A researcher who would have hibernated 20 years ago and just woke up would be able to update in no time.
Proposing a new system is difficult, because the current systems are extremely hard to beat: the state of the art requires thousand of hours of training, which means you are almost required to use a lot of programs outside your own lab. For example, almost everyone uses the same software for acoustical modelling (HTK), same labeled data, etc... The published improvements are often tiny (less than one point, e.g. going from 79 % to 80 %), and I actually wonder if those are scientifically significant.
Using things like pitch and other non verbal cues (prosody is the actual term used in the literature) is often suggested, but estimating those reliably is extremely difficult. I would guess this will improve because some of languages which are currently funded require prosodic information. Chinese is the obvious example: Chinese has the notion of tones, where the pitch of phonemes may radically change the meaning of the word.
On the bright side, I think speech recognition was one of the first field which pushed the idea of data-driven models (at least 3 decades when the first usage of HMM appear), and as such, has been on the forefront of the current explosion of non trivial statistical models trained on big datasets. In that sense, I think it had non trivial contribution outside its own application.
That seems a little off to me. It seems like you would want a list of phoneme probabilities to work with along with how long each sound took etc. Your language model might say “I saw the see past the dune” is less likely than “I saw the sea past the dune” but the more information the language model has to work with the better. For example, dialect impacts both the sounds people use and the words they use.
Also sperating background sounds from a message is easer if you are following the content of the message etc.
The article Jeff linked talks about how how the performance has pretty much levelled off except for small incremental improvements and most of the money and research was shifted away from the area. I guess the trend will continue, gaining small improvements through applying more and more data as training for current methods until someone comes up with a completely different approach to push it forward.
Until that basic reality changes, I don't think mainstream computing needs will change. Along the same lines, I don't think the reverse is feasible, with computing interface changes bringing about social changes. The shifts we've seen from hand-writing to type-writing to computer-based word-processing has been happening in roughly the same environment for a very long time.
As soon as I hear "If you want X, please say...", I start pounding the zero button. The thing is, relying on buttons is not only faster in itself, but also doesn't force me through so many "you said 'X'; is that correct?" questions.
If it were something straightforward I would already have helped myself on your website. The fact that I'm on the phone means that it's an open-ended question requiring a real human on the other end (or that your website is useless).
And we're probably not much into it just because of how impenetrably tortuous it is. Almost everything related to it is variable -- culture is always evolving language, slangs die of usage over time, and there are accents to worry about. With all of these problems, there is the demand in tandem for progress in nlp areas to deal with -- where really some of the most difficult challenges lie ( http://en.wikipedia.org/wiki/Natural_language_processing#Con... ). With all that said, I have my money down on Google. They seem to be doing a lot of work in areas where the variability of these tasks is required. Presently, the voice transcription feature on youtube seems fairly impressive, and I've noticed in the past google search's nlp abilities to be curiously good, certainly more far ahead than any other search engine today.
Remember grade school, with parts of speech, Conjunction Junction, and diagramming sentences? Well, think back to the last (non-trivial) sentence you actually said out loud. I challenge you to try to diagram that sentence.
Although there's quite a lot of research that's been done, surprisingly little of it deals with the way most of us really communicate.
That's not to say that current approaches to linguistics are correct; I don't know enough to make that judgement. But I think the problem lies more in the complexity of the domain - not because modern linguists are so foolish as to base theory on written language.
not because modern linguists are so foolish as to base theory on written language
That was actually what I was trying to get at. While there are clearly some rules we follow when speaking, they are much more open than those governing our writing.
It's odd, though, or at least counterintuitive, that our comprehensive of these less-structured spoken communications is higher than that of more-structured written ones.
Undoubtedly there is much room for improvement on these higher-level features but computers are still well behind humans in large vocabulary isolated keyword spotting: this is a task where one word from a very large corpus of words is spoken and the human or computer has to guess what that word was. Computers do poorly relative to humans (particularly in noise), which suggests that many of the mistakes that computers make is in not being able to interpret the acoustics correctly.