This is what I don't get. I used to play with Microsoft's SpeechAPI back in 2007 and it was pretty decent at real-time speech recognition - almost perfect, if you limited your program to a pre-defined command grammar. It was all done completely off-line, real-time on a machine that had much less processing power than your average smartphone of today. I don't see anything inherent in speech recognition that would require so much processing power that the phone can't handle it real-time. The only reason I see is that continuous moving of everything to cloud because it's easier to capture revenue stream that way.