Facebook has open sourced some pre-trained models: https://github.com/facebookresearch/wav2letter
Picovoice has some smaller, more efficient models capable of running on edge devices: https://github.com/Picovoice
Full ASR does require quite large models and datasets, but you don't need nearly that much power or data to fine-tune a model for your own domain.
They have an example that accepts streaming from the microphone: https://github.com/mozilla/DeepSpeech/tree/master/examples/m...
See the last full release here: https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1
https://github.com/facebookresearch/wav2letter/issues/327
Someone in that thread ported it to linux. This demo is just "acoustic model emissions", which are character level predictions (with no repeated characters), but I have "decoding" (turning into english sentences) working locally as well and I'll post a new demo at some point.