Speech recognition worked fine in the late 90s. You could optionally train it for 15min with your own voive to improve the results. The Nuance software was better back then than what shipped with Win Vista inbuilt by default in 2006 and later. The same Nuance software runs nowadays on servers to power Siri, Google Now, Cortana. And who in the right mind would think every user gets its on dedicated super-computer. In reality, cloud based speech recognition and NLP has the advantage of a central database to collect different voice samples (to train) and a multi Gigabyte central NLP database, etc. But you won't get more CPU cycles than what would be available on a modern smartphone or a Smart TV or (what would fit in) an Echo like device. An offline-available software assistant as kind of premium feature for people who care.
So true - I vividly recall my attempt to dictate "Dear Sir" that ended up appearing on the screen as "Down Server" no joke.