Obviously the big cloud players offer their own APIs and SDKs (for a price), but there are a few other solutions worth looking at.
Facebook has open sourced some pre-trained models: https://github.com/facebookresearch/wav2letter
Picovoice has some smaller, more efficient models capable of running on edge devices: https://github.com/Picovoice
Full ASR does require quite large models and datasets, but you don't need nearly that much power or data to fine-tune a model for your own domain.