No mention of DNN based ASR like DeepSpeech? There’s even open source python implementations available from Mozilla and Paddle.
These models are way easier to train, have surprisingly good accuracy, and are robust to noise.
These models are way easier to train, have surprisingly good accuracy, and are robust to noise.
From prior use, Google's speech API (at least the "video" model) is freakishly accurate compared to DeepSpeech to where I wondered if they used closed captioning to help train their model. But I haven't seen rest of these at work: https://i.imgur.com/cdOlARO.png