Having stated my bias: Speech recognition systems are actually not that complex at their core. It's a blending of statistical models. Getting good data is a problem. You need a good acoustic model that's adapted to your users and the environment in which they will be using your application. Everything from the fluency of speakers, to physical environment, to the characteristics of the channel over which the speech is sent needs to be considered.
If you have a good acoustic model, now you have to worry about your language model. Are you going to try to accept all words in a language, or just restrict your users to a particular domain of language? If you have a good language model, then you need to worry about the dialog management. How do you keep context in a conversation? It's not an easy problem.
The primary problem with speech recognition systems is that human beings set their expectations of them too high. It's a psychological factor. When those expectations are not met, the user is frustrated and angry. Consider this. Whenever you call AT&T, your health insurance company, or credit card company, do you enjoy the experience of the IVR system that routes your call? Probably not. You probably don't even talk to it and resort to pressing the buttons instead. Unfortunately that's the experience most people have with speech recognition. I think it's the worst possible application of it.
If you're making a small, toy application whose vocabulary is pretty restricted and whose functionality set is small, then you're probably okay. If you venture into full dialog/anything-goes type applications, the chances are high that your app will be a bomb.
These researchers can swap out all the lower-level statistical models they want, but it won't fundamentally improve the technology. There are systems out there with word error rates very close to that of humans, but the systems higher up in the stack that interpret what is recognized are still very crude.