A TensorFlow implementation of Baidu's DeepSpeech architecture
github.com
github.com
[1] https://www.kaggle.com/c/tensorflow-speech-recognition-chall...
> There are only 12 possible labels for the Test set: yes, no, up, down, left, right, on, off, stop, go, silence, unknown.
This won’t work for large vocabulary continuous speech recognition, which is what you want if you want to transcribe podcasts, phone calls, or generally human-to-human, spoken interactions.
Has anyone tested this out? Impressions of its usability?
What makes it less suitable for production compared to TF?
Is it just the Kafkaesque build process (admittedly I last tried a year ago), or is model-fitting or prediction especially buggy?
2017: Kids can run real time E1+ voice transcription systems made exclusively of free software on commodity gaming hardware. Bizarrely, the dominant implementation is based upon the "free" browser community Mozilla, based upon work released by a "don't be evil" global megacorporation, but they are reduced to imitating China to get there.
https://www.wired.com/2009/08/dayintech_0806/
Not everything changes. 1997 - 2017: Microsoft has 90% desktop market share, but the Year of Linux on the desktop was fast approaching.
Linux not gaining a least some traction was a huge disappointment.
yes indeed, like a few thousand kilometers.
PS. With respect to "There’s nothing “dominant” about this implementation or the DeepSpeech architecture in general." the use of dominant was really poetic license in support of the line of amusement, but in fact I'm not aware of a more popular open source transcription system by Github stars ... are you? As for China != Baidu SV, I actually find it even more bizarre that they would develop such algorithms in such a foreign environment.
However, you can use voice activity detection (VAD), for example webrtcvad from PyPI, to chop long audio into smaller bits that are able to be digested.
Maybe we should just put VAD in the client and have this occur automatically?
Just this week I started looking into how I could generate transcripts for a bunch of videos. Even if the transcripts aren't perfect, it helps with tagging and searching through large video collections that include certain keywords.
Sadly, I didn't have any luck with local solutions. I managed to generate a few transcripts using GCP's Cloud Speech API with minimal hassle, but I'd much prefer to do it locally.
I was planning on trying this out later today, and had already downloaded the Common Voice corpus. Having to add another step to break up the input into smaller chunks probably isn't a huge deal, but I wouldn't have known what tool to use in order to achieve that.
Do you know of any comparisons between various speech-to-text tools? I've avoided commercial tools so far because I'm hesitant to drop $250+ just for playing around, but I'd be interested in seeing if they're truly superior to existing open alternatives.
I can't promise when we'll get to it, as from now until new year is a bit of a wash.
I don't know of any detailed comparisons of commercial solutions. However, with respect to pure word error rate, the article[2] does a comparison of several engines as of circa 2015.
And thanks for you Mozilla peep’s hard work!
https://google.github.io/tacotron/publications/tacotron/inde...
They are some open source implementations.
Edit: Another interesting one: http://research.baidu.com/deep-voice-3-2000-speaker-neural-t...
I don't have an ETA, but it's in the works.