DeepSpeech 0.6
hacks.mozilla.org
hacks.mozilla.org
1. I want to teach it ten words. How do I do this?
2. I want to speak into my microphone (available as a Pulseaudio device) and recognise the words and output the words as a text stream on stdout. How do I do this?
This is the documentation:
https://deepspeech.readthedocs.io/en/v0.6.0/Python-Examples.... https://deepspeech.readthedocs.io/en/v0.6.0/Python-API.html
It does not answer the questions I have.
https://deepspeech.readthedocs.io/en/v0.6.0/DeepSpeech.html
The introduction page is full of incomprehensible jargon.
I'm assuming you mean "my target vocabulary is only ten words". In that case, you have a couple of options:
1. Just use it as is and see how it does.
2. Create a language model containing just those 10 words so that the engine is constrained to that vocabulary, which should improve accuracy.
3. Train a (smaller) model from scratch, which requires (audio, transcript) pairs. This is basically research work, there's no recipe you can just follow. The advantage is that if you can make it work, it can be a much smaller model than a model trained for general English dictation.
> 2. I want to speak into my microphone (available as a Pulseaudio device) and recognise the words and output the words as a text stream on stdout. How do I do this?
We have a few examples with microphone input, but they're all based on voice activity detection: https://github.com/mozilla/DeepSpeech/tree/master/examples
I should write an example that is just continuous streaming and output without any voice activity detection.
I'd be happy to discuss further on Discourse: https://discourse.mozilla.org/c/deep-speech
How? The docs seem to be suffering from the same problem as the git docs: they assume the reader is already a domain expert, which would seem to run counter to the stated goal of being simple (and ubiquitous).
I tried to figure out how to "just use" deepspeech and failed.
There are also examples transcribing from the microphone as I mentioned above: https://github.com/mozilla/DeepSpeech/tree/master/examples/
I also have a small GUI example here [0], click once to start recording, once again to stop and show transcript. It receives the same command line arguments as the ones shown in the README link above (namely --model, --lm and --trie).
The README starts with a very basic installation and usage guide, I don't see how that's assuming any expertise.
Fyi, the other 4 example links there are 404. Also in the root README. The linked README's are missing, or diversely suffixed and mislinked.
I was able to find the right links though, for example the WAV transcriber is at : https://github.com/mozilla/DeepSpeech/blob/v0.6.0/examples/v...
1. Train it using the training dataset but change all word labels that aren't in your 10 word list to "OTHER_WORD" or whatever. There's probably not much point doing this. Alternatively I guess you can restrict the beam search to only look for your 10 words. Yes this is really complicated - it's a speech recognition engine - they are really complicated!
2. The API docs look pretty clear to me:
> createStream()[source]
> Create a new streaming inference state. The streaming state returned by this function can then be passed to feedAudioContent() and finishStream().
> feedAudioContent(args, kwargs)[source]
> Feed audio samples to an ongoing streaming inference.
> intermediateDecode(args, kwargs)[source]
> Compute the intermediate decoding of an ongoing streaming inference. This is an expensive process as the decoder implementation isn’t currently capable of streaming, so it always starts from the beginning of the audio.
Huh well that's a flaw the neglected to mention! I guess you can't really do what you want yet.
The documentation is outdated. That is no longer true, intermediateDecode is cheap. Thanks for noticing, I'll fix it.
It's been a little while since I got it running, but I basically got a siri clone working. If you want to test it out, I can try to answer questions / whatever problems pop up.
The code is here: https://github.com/shawwn/DeepSpeech/commit/01f5cf8d39c356ae...
As far as I know, you can simply run speech_to_text.sh. It will connect to your microphone and start dumping out transcribed audio to stdout.
It wasn't super easy, but once you spend a little time with the code you can sort of figure out ways to get it to do what you want.
EDIT: By the way, people will try to convince you that you're nuts and that the documentation is crystal clear and so on. Know this: It's not just you. I had the exact same experience. It's a recurring theme in AI programming.
The only thing to do is to either find someone else who feels similarly, or roll up your sleeves and dive into the code.
No it really is simple: the original implementation had a base of prefabulated amulition grammar, this was surmounted by a malleable logarithmic function in such a way that the two main spurving vocabularies were in a direct line with the panametric grammar fields. The latter simply consists of marzlevanes fitted to the ambifacient morpheme waneshaft to eliminate side utterances. This main winding is of the normal lotus-o-deltoid type, just placed in panendermic semi-boloid stators. Basically every second conductor is connected by a nonreversible tremmie pipe to the differential grammeters.
All these words are in: https://en.wikipedia.org/wiki/Turboencabulator
Anyone have a comparison for how good/bad that is compared to other solutions, and what it means for practical usage, if that can be guessed at from a single number?
It seems like sota is 2.20% word error rate
For real world applications, this is absolutely crucial, users want latency numbers on the order of milliseconds, not seconds. This is why, if you run a standard test set like LibriSpeech on, say, a commercial offering from Google, it will perform considerably worse than state of the art according to Google papers.
This repository [0] has a benchmark of some commercial offerings. Our model beats all of those on Librispeech clean and other (except for Speechmatics on Librispeech clean), as well as on Common Voice. But note that the Common Voice corpus used in that benchmark is very old.
In sum, I would compare this against solutions that go for the same space: fast, client-side ASR, rather than state of the art.
[0] https://github.com/Franck-Dernoncourt/ASR_benchmark#benchmar...
I guess you prefer an "end-to-end" model over a hybrid HMM/NN model, for simplicity, right? As far as I remember, you use CTC, right? I always wondered why you have chosen CTC, and not some better model, like RNN-T, RNA, or some of the streaming attention variants. They should give you all the same properties (online capable, low latency, simple, end-to-end), but much better WER performance. Or is this simply because there currently is no simple ready-to-use implementation for those? Note that we published some TF code recently for some streaming attention variants, and plan to publish some RNN-T/RNA code soon.
> I guess you prefer an "end-to-end" model over a hybrid HMM/NN model, for simplicity, right?
Simplicity and ease of targeting other languages, yes. We're a small team.
> As far as I remember, you use CTC, right? I always wondered why you have chosen CTC, and not some better model, like RNN-T, RNA, or some of the streaming attention variants.
We started DeepSpeech in 2016, before these recent developments for end-to-end ASR were mainstream/SotA.
> Or is this simply because there currently is no simple ready-to-use implementation for those?
Implementing the model architecture for training is only part of the problem for us. We have a hand-crafted inference graph to allow for small and efficient client code and inference models, and the more complex the architecture is, the trickier it gets to make sure it all works on all platforms, including TFLite, with quantization, etc.
We're investigating alternative architectures as well as mixed CTC/RNN LMs to deal with language model size, but no final decisions made yet.
> Note that we published some TF code recently for some streaming attention variants, and plan to publish some RNN-T/RNA code soon.
Nice! Can you share a link to the streaming attention code?
IMHO the WER is more important than latency improvements in the millisecond range. The most frustrating thing is having to dictate over and over and the transcription is incorrect each time.
Consider that the time to a correct transcription is the latency plus error correction. If error correction is manual it will be orders of magnitude slower, so optimize for WER.
I’m terms of competition, Siri has latency in the 5+ second range due to the network call especially in area with poor data rates. I think a client side model like yours will easily win in this category. If you’re already ahead here, why not focus on WER next?
Another great capability is to generate alternative transcriptions for words with low confidence values to allow for quick error correction. Do you offer something like this today?
Also, consider the long term view that new models are constantly being released and refined. It’d be best to have an architecture that allows quick replacement without a lot of hand tuning, or where the tuning can be automated to a greater extent.
+1000 to @mostlyjason's comment - Great latency figures mean nothing if the word error rate is high, since it dents confidence in the output (so why use DeepSpeech?) and (as the parent comment notes) necessitates manual error correction.
I would love to see a future release focus on optimizing WER for these reasons.
It uses our TensorFlow framework Returnn (https://github.com/rwth-i6/returnn). We currently only have some configs/code online for hard attention variants, or segmental models. They can be found here: https://github.com/rwth-i6/returnn-experiments/tree/master/2...
Note that the configs are maybe not so cleaned up, as this is very much research. This is for a paper we submitted to ICASSP. We did not publish the paper yet elsewhere, but I can send you a copy by mail (just contact me: albzey@gmail.com).
Also, as this is research focused, the encoder here is also a BLSTM, because we wanted to compare this work to other global soft attention models, and have the comparison mostly focused on the streaming attention modeling aspect. And also there is some lookahead, which is currently unlimited. So it would need a few modifications to really be used for online streaming. I'm also not sure whether this is the best model, or whether some of the many other variants (MoChA, RNN-T, etc) are maybe better.
Edit: I forgot, we also have some local windowed attention variants, which can also be applied for streaming: https://github.com/rwth-i6/returnn-experiments/tree/master/2...
A bit misleading given that decoding can't be streamed yet.
Edit: never mind - the API documentation was just out of date.
You should follow Google's approach - give fast live results that don't depend on data from the future, but also go back and correct old words when you do have that data. It's kind of how humans work really.
My wav2letter research does not use streaming, despite supporting streaming, because I get lower average latency (roughly 0.02x RTF including decoding, or 80ms for a 4000ms input audio like you describe in the article) by running the entire encoder CNN at once instead of running it in chunks.
This is also nice because the CPU usage is practically nothing (0.1%) until you stop talking.
I have a web demo here [1], with a relatively terrible language model (mostly just been working on acoustic modeling so far, as it’s just me). Most of the latency is waiting for the JavaScript VAD (which is much slower than webrtcvad and I couldn’t figure out how to tune it) and waiting for the network. If you look at the network inspector, the server should report its encode and decode times.
Besides being a great resource for speech analysis, this could be a real game changer for acquiring listening comprehension in a foreign language.
I feel that even after a few years of learning a new language I still have trouble with listening. Part of that is that it's often all or nothing, even one or two unknown words in a sentence means I can't understand the sentence. But worse is that most language teaching materials use a very small set of native speakers, which deprives the learner's brain of being able to generalize.
For that, it's important to include accents and even non-native speakers, especially in English.
I think that those wavenet models for TTS were able to be conditioned by gender and maybe accent, if they had the data.
[1] https://storage.googleapis.com/tfjs-models/demos/posenet/cam...
It all seems very doable based on what I see in the technology today, I just don't have the skills to do it.
There's definitely some ways it can go bad, but even basic fluency monitoring without any active remediation attempts would be a good addition.
How do you define fluency? I guess it involves some combination of speed and accuracy. Speed shouldn't be too hard to measure, but for accuracy you'll probably need a model that's more accurate than 7.5% WER and can handle the difference in vocal range between children and adults. Otherwise the speech disfluencies you want to detect will be drowned out by the model not correctly recognizing actually pretty clear speech.
You definitely would have to use a model trained from content in the target audience (probably down to the grade level as things change so dramatically from year to year) as well as probably some labeled examples of students with various reading/speech deficiencies. This of course would lead to a lot of challenges from a privacy/regulatory/etc perspective and the entire thing would be challenging from an optics perspective (AI teachers taking over, etc etc).
Edit: One thing to keep in mind with accuracy is that there is no ambiguity what the word should be, the question is how closely the utterance matches the expected sounds of the word.
Certainly seems like an area with huge potential application.
This is great news.
I'm not very familiar with the deep learning framework ecosystem. Does anyone know what the simplest way would be to incorporate a DeepSpeech model into a WebAssembly project?
For instance, is it straightforward to compile Tensorflow Lite with emscripten? If not, can this TensorFlow model be run with tensorflow.js? Or can it be converted to some other format that is easier to use with WASM?
But the whole point of the TFLite model is that it has low resource requirements so an accelerator for inference shouldn’t be necessary.
The English acoustic model that we released does not have a fixed vocabulary. What determines the vocabulary is the language model, which can be created from text. So if you have in-domain text, and would like to try it out, I would say the first step is to create a language model using your text and then experiment with it.
We have some documentation on how to create the LM here: https://github.com/mozilla/DeepSpeech/tree/v0.6.0/data/lm
It's not super detailed, but I'd be happy to answer questions on our Discourse: https://discourse.mozilla.org/c/deep-speech
https://2r4s9p1yi1fa2jd7j43zph8r-wpengine.netdna-ssl.com/fil...
Or, if you prefer the code: https://github.com/mozilla/DeepSpeech/blob/0427c1572ac8f7253...
(well, I suppose technically it's in https://github.com/mozilla/DeepSpeech/blob/master/data/alpha... through https://github.com/mozilla/DeepSpeech/blob/master/util/text...., but I mean come on)
In practice ... meh. It gets some things predictably wrong. They're readable, just wrong. So the good news is that WER is words it has 100% correct. The remaining errors aren't usually totally unreadable, just irritating, going from "totext" instead of "to text", but of course names go pretty wrong and don't get better. It butchers my name, of course.
If you test it yourself, please do realize it really doesn't deal very well with accents.