Show HN: Text-to-speech and speech-to-text open-source software stack
github.com
github.com
The original task was to automate the testing of a voice-enabled IVR system. While we started with real audio recordings, very soon it was clear that this approach is not feasible for a non-trivial app and it will be impossible to reach a satisfying test coverage. On the other hand, we had to find a way to transcribe the voice app response to text for doing our automated assertions.
As cloud-based solutions where not an option (company policy), we very quickly got frustrated as there was no "get shit done" Open Source stack available for doing medium-quality text-to-speech and speech-to-text conversions. We learned how to train and use Kaldi, which is according to some benchmarks the best available system out there, but mainly targeting academic users and research. We made heavy-weight MaryTTS work to synthesize speech in reasonable quality.
And finally, we packaged all of this in a DevOps-friendly HTTP/JSON API with a Swagger definition.
As always, feedback and contributions are welcome!
Was this before the streaming API was added to DeepSpeech? I recently did some testing with it, and it provides text within ~100ms of last audio block on my PC.
edit: that is, the most significant latency I had was from having to wait a bit to detect end of speech.
Fortunately for me I only needed it for command recognition, so the process was quite quick, and results were very good. No need to retrain the net.
I could contribute towards it since I have done it before.
Thank you for building this!
Why does the voice/pronunciation have such drastic volume spikes and dips?
- https://github.com/gooofy/zamia-speech#asr-models
- https://github.com/mpuels/docker-py-kaldi-asr-and-model
in regards of speech recognition except the fact that its easier to use?
the other one is an example for packaging kaldi in a docker container.
in the past to provide Speech training data, but they were not really interested. I thought it would be great to improve STT having a real good and HUGE set of german audiobooks based on Text, that is publicly available... unfortunately i had no success trying to script something for this purpose (mainly lack of time).
It basically describes the thing you mentioned - matching freely available audio books with the source text and using some tools to preprocess the data suitable for ASR training (alignment, splitting).
1. Compiling it is hit and miss. Sometimes it works, sometimes it doesn't. There is no official package in any Linux distribution, so packaging anything with it is incredibly painful. There's no easy way to cross-compile your project, so you'll end up working around the build process.
2. The documentation is woeful.
> Recent CMUSphinx code has noise cancellation featur. In sphinxbase/pocketsphinx/sphinxtrain it’s ‘remove_noise’ option. In sphinx4 it’s Denoise frontend component. So if you are using latest version you should be robust to noise in some degree already. Most modern models are trained with noise cancellation already, if you have your own model you need to retrain it.
> The algorithm impelmented is spectral subtraction in on mel filterbank. There are more advanced algorithms for sure, if needed you can extend the current implementation.
Inconsistent methods across the codebases prevents you knowing where to look, and if it is documented, it may involve spelling errors which you have to guess around (like above). Also plenty of vague references to other documents that may or may not even exist.
I maintain a platform which features live video events we'd like to add captioning and so far can only see IBM Watson providing a websockets interface for near real time stt.
we are already using it for a callcenter with around 50 parallel audio streams.
Good example: https://cloud.google.com/text-to-speech/
Of course it is not a competitor to Google in any sense.
high quality with google cloud speech and amazon polly.
Also I suspect these guys aren't using the latest kaldi architectures.
Aren’t some of the peak Kaldi numbers also with _huge_ language models that aren’t so practical to deploy? The best kaldi results I’ve found here both use “fglarge” LMs, but a quick search didn’t tell me how big that actually is.
The best I can find for the test set are (from main RESULTS):
test-clean: 4.31% WER
test-other: 10.62% WER
From here (not sure why buried): https://github.com/kaldi-asr/kaldi/blob/master/egs/librispee... test-clean: 3.80% WER
test-other: 8.76% WER
This is the 2019 wav2letter SOTA:
https://github.com/facebookresearch/wav2letter/blob/master/r...Sure, one of their techniques is to semi supervised train on 60k hours of audio, so I’ll post their best numbers both with and without that step (as you said “if kaldi is trained on the same data”)
Without 60k unsupervised audio:
test-clean: 2.31% WER
test-other: 5.18% WER
With 60k audio: test-clean: 2.03% WER
test-other: 4.11% WER
That’s so far ahead of kaldi’s librispeech RESULTS, let’s go up the list and look at their acoustic model tests that don’t even bother with a language model:Without 60k or language model:
(test-clean / test-other WER)
3.05 / 7.01
With 60k without language model: 2.30 / 5.29
Am I missing a kaldi whitepaper here?I found “State-of-the-Art Speech Recognition Using Multi-Stream Self-Attention With Dilated 1D Convolutions” on wer_are_we claiming 2.2 / 5.8 using Kaldi as a base, which is impressive, but they changed the network architecture and I didn’t find this in any of the open-source kaldi librispeech subdirectories. (If arbitrary model changes count, at some point you might as well treat both wav2letter and kaldi as “tensorflow with some speech helpers” because I’m sure you can implement either framework’s best neural architecture in the other framework.)
If there’s a better paper on kaldi than 2.2 / 5.8 you should comment with it and PR it to wer_are_we, because wav2letter is currently beating every other LS result posted there.
Point being if you put some more effort I believe you can get better results with kaldi. Remember we're discussing with which toolkit one can achieve better performance (given a certain amount of effort), which is not the same as just taking the results from different teams some of which have spent an order of magnitude more compute in gaming the results.
For SOTA results with a hybrid approach check out RWTH's stuff. The kaldi people have recently been busy rewriting their backend to work with pytorch.
In general using librispeech to gauge performance is a bad idea. It's read speech. I prefer to base my opinions on real world datasets.
Have to say I am impressed by how good their (FB) results are without using an LM though.
RWTH has an impressive entry on the wer_are_we list for librispeech (2.3 / 5.0), but unfortunately I couldn’t find a way to actually use that research without a lot of work. Facebook has easy recipes posted that reproduce their results, which have been quite nice to work with. RWTH, Google, and Capio held an impressive top spot in the results for a while, which might be where you’re basing your judgement.
I don’t think the biggest benefit of wav2letter was the architecture (until their SOTA paper), I think it was the speed and simplicity. I managed to port the original wav2letter inference to numpy/pytorch in two small files (same model weights, same inference output) because it was such a simple convolutional architecture. (The only fiddly part was matching the output of their featurization code)
I agree the data and training overhead is a big aspect, which is why I’ve been training and releasing freely available wav2letter models on far more data than what Facebook has published using librispeech.
My upcoming goal is to reproduce the SOTA results but with stronger training data. My baseline data is around 4000h currently and is far more diverse than librispeech. Unfortunately I’m independent and can’t afford the LDC datasets, though I have some neat ideas for extending FB’s semi supervised work to even more data.
Cool project! I'll try out your models.
https://github.com/facebookresearch/wav2letter/blob/master/r...
My models are here: https://talonvoice.com/research/ I haven’t yet posted the model I’ve been working on most recently. I’m 120 epochs into a large size model trained on all of my datasets. I also have another 1000-1500h of audio I haven’t finished prepping to train on.
Here is a web demo. It’s currently running a slightly older checkpoint of my WIP large model, and the deepspeech LM: https://web2letter-west-1.talonvoice.com/
Do you have any experience with online decoding in wav2letter ? Is there something like a Websocket API available somewhere ?
the english is from the tedlium recipe with WER of 7%.
room for improvement, but for our original purpose it was sufficient.
just an example, kaldi is a weird mixture of c++, python2, python3, shell scripts, java, perl. hard to oversee. deepspeech is python. wav2letter is an exe file.
I was under the impression DeepSpeech was native (C++), with bindings for Python and others. Personally I've used it with Node.js so far, and I couldn't see any dependencies on Python.
edit: I was talking client, you're talking training I guess.
in short: not surprising the blockbuster cloud services provide better results as they have way more training data. tradeoff between price, privacy, quality.
just a small server, hope that it wont crash when posting the link here
The language model is built from sentences and is rather quick to build (~seconds), at least for the small number of sentences I used. This can then be combined with the pre-trained neural net.
Though I hear DeepSpeech is currently fairly US-influenced when it comes to recognizing accents. So if you're not a native speaker, consider contributing to the open-source dataset over at https://voice.mozilla.org/
training an asr model is totally different requirement than using it as a client
> the link here
In all honesty I was just looking for some pre-processed examples.
The API itself seems intuitive enough, good job.
What really stands out to me is that open-source text-to-speech is really awful compared to commercial solutions - which surprises me. Does anybody know why this is the case? (And not just "money", I'm talking technology)
i guess thats for the same reason that google is dominating the speech recogniction world: they have tons of training data available. not smarter algorithms, just more data.