A 2019 Guide for Automatic Speech Recognition
heartbeat.fritz.ai
heartbeat.fritz.ai
Facebook has open sourced some pre-trained models: https://github.com/facebookresearch/wav2letter
Picovoice has some smaller, more efficient models capable of running on edge devices: https://github.com/Picovoice
Full ASR does require quite large models and datasets, but you don't need nearly that much power or data to fine-tune a model for your own domain.
https://github.com/facebookresearch/wav2letter/issues/327
Someone in that thread ported it to linux. This demo is just "acoustic model emissions", which are character level predictions (with no repeated characters), but I have "decoding" (turning into english sentences) working locally as well and I'll post a new demo at some point.
They have an example that accepts streaming from the microphone: https://github.com/mozilla/DeepSpeech/tree/master/examples/m...
See the last full release here: https://github.com/mozilla/DeepSpeech/releases/tag/v0.5.1
It's also too bad this doesn't mention any traditional HMM-based ASR techniques, as HMMs continue to be used on many SOTA systems, particularly those that can be reproduced publicly: https://github.com/syhw/wer_are_we
The article quotes DeepSpeech, wavenets, LSTMs of all sorts; essentially, all the neural networks that scale terribly. DeepSpeech for example is pretty heavy and requires a decent GPU to get anywhere near a realtime factor of 1.
Meanwhile ASR through HMM's consistently hits realtime factors sub-1 and can run on small CPUs. e.g. the default models Kaldi ships with outperform DeepSpeech on a lot of modern examples, and are exponentially faster.
Moreover, training these HMM's is something that is feasible for a normal developer. Training newer models requires data of scale and quality (iirc Mozilla's models are trained on Common Speech which is an enormous crowd sourced dataset, and Google's wavenet models use an internal dataset of very high quality and quantity).
Until the models get more practically achievable, ASR for average people will probably continue to be dominated by Kaldi, Sphinx etc
^^ If you're getting started, just following the steps here can get you set up really fast. IIRC there's also an HTTP endpoint at /recognize you can use instead of WebSocket if you're transcribing audio files so it's pretty cool!
(Wikipedia doesn't make the best LM, I just wanted to test with something that knew about a lot of interesting english words)
1. I don't have the quad core CPU we hit <0.01x on, I only have a dual core laptop CPU. (The quad core numbers were with an i7 7700k or so, they came from the other person I'm working on this with).
2. I'm not running a parallel decode. That's in a branch from my collaborator I haven't merged/built myself yet since I'm only on a dual core.
3. Screen capture has a CPU hit and seems to have slightly increased my RTF during recording.
Here's your comment read with 0.05x-0.10x on a dual core CPU: https://youtu.be/jIgUKwR-LaA
Is that enough to convince you that with a stronger CPU and parallel decode we can hit 0.01x?
This video can be partially reproduced with the code and models from here: https://github.com/facebookresearch/wav2letter/issues/327 but most of our decoder optimizations post-date when I compiled that demo.
Some of the later code is here:
https://github.com/talonvoice/wav2letter/tree/w2lapi_static
https://github.com/ckamm/wav2letter/tree/wip-parallel
Here is the exact timing loop code I used in that demo: https://bochs.info/p/nhz5cv
I definitely think it's a big step in the right direction; it's easily 100x faster than DeepSpeech for us.
If I could have anything I wanted for xmas, I'd ask for a speech to text system that is fast enough to work in browser thru wasm or something.
Is there a SIMD.js / WASM equivalent optimized convolution / GEMM? That's pretty much all we'd need to port this to web... well, that and maybe a language model that isn't 1GB. The wav2letter acoustic model I'm using is based on the librispeech conv_glu, which is almost entirely served by conv1d layers.
I've honestly already been considering a demo for my main project (which is mixed english / command decoding) that runs entirely in a web page, if you have engineering time to throw at your christmas wish, we should talk :P
I really love the idea of hacking something together for development so I don't need to use my arms and hands so much, or could lessen my mouse use (for accessibility reasons), but I don't know how realistic that is.
Edit: I looked into cloud-based ASR, such as that provided by Azure and AWS, but that would mean network latency on top of the recognition latency, and that would drive me nuts!
So, if your project can be built on that, good news; you can build now.
But I'm unsure of what Snips actually is - I had a look at the website, but I don't know if this is OSS, commercial software, a library or what?
https://www.youtube.com/watch?v=OWyMA_bT7UI
It used the windows version of Dragon Dictate that had a python interface which was then hooked up to emacs, iirc. I attempted to replicate his system at some point but never got it working -- seems that the libraries he referenced weren't well maintained and/or running windows through a VM on my mac introduced additional issues.
I'm not totally convinved that the language needs to be designed specifically for voice coding though - I can see how having an IDE designed for it would be a huge bonus though... damn, another interesting side project to add to the list!