Speech recognition is terrible on specialised vocabulary, strong accents, low quality audio, etc. You can train it, but only if you have transcriptions of your audio. Steno is a good way to make transcriptions.
I have mathematics education videos I would like subtitles for, all in non-US, non-UK accents, full of specialised terminology. In my tests, Amazon and Google's APIs were completely useless for this, as was Dragon. If I find the time, I would like to learn steno of some sort to make the subtitles and provide training data for future videos.