http://cmusphinx.sourceforge.net/
http://julius.osdn.jp/en_index.php
It is amazingly easy to create speech recognition without going out to any API these days.
This was 2 years ago, so maybe it's simple now, but I didn't find it "amazingly easy" back then.
I do have a number of projects where I could definitely use a local speech recognition library. I have used [Python SpeechRecognition](https://github.com/Uberi/speech_recognition/blob/master/exam...) to essentially record and transcribe from a scanner. I wanted to take it further, but google at the time limited the number of requests per day. Today's announcement seems to indicate they will be expanding their free usage, but a local setup would be much better. I'd like to deploy this in a place that might not have reliable Internet.
But these days, if you go all the way through their tutorial, and give it a proper read, it's very doable to set up.
We've written detailed, up-to-date instructions [1] for installing CMU Sphinx, and now also provide prebuilt binaries [2]!
If you're interested in not sending your audio to Google, CMU Sphinx and other libraries (like Kaldi and Julius), are definitely worth a second look.
[1] https://github.com/Uberi/speech_recognition/blob/master/refe... [2] https://github.com/Uberi/speech_recognition/tree/master/thir...
CMUSphinx is not a neural network based system, they do use language and acoustic modeling.
Also, they give you the tools and knowledge to build better models (and explain the theory), which is where most of the competitive advantage is IMHO.
Googles engine also works fine (have been trying it with the phones), but the pricing may or may not be a deal breaker.
Not really. The hard part is not the algorithm, it is the millions of samples of training data that have gone behind Google's system. They pretty much have every accent and way of speaking covered in their system which is what allows them to deliver such a high-accuracy speaker-independent system.
CMUSphinx is remarkable as an academic milestone, but in all honesty it's basically unusuable from a product standpoint. If your speech recognition is only 95% accurate, you're going to have a lot of very unhappy users. Average Joes are used to things like microwave ovens, which work 99.99% of the time, and expect new technology to "just work".
CMUSphinx is also an old algorithm; AFAIK Google is neural-network based.
https://github.com/yajiemiao/eesen
Baidu open sourced their CTC implementation
https://github.com/baidu-research/warp-ctc
I think we will have an easy to install OSS speech recognition library and accurate pretrained networks not far off from Google/Alexa/Baidu, running locally rather than in the cloud, within 1-2 years. Can't wait.
Speech Intent Recognition
... the server returns structured information about the incoming speech so that apps can easily parse the intent of the speaker, and subsequently drive further action. Models trained by the Project Oxford LUIS service are used to generate the intent.
Do others offer something like this?