Open Source Speech Recognition
chrislord.net
chrislord.net
I want to take this opportunity to plug an Alexa Skills Kit API (in Python) I just dropped:
https://github.com/johnwheeler/flask-ask
It does a lot of things the JS API doesn't: Jinja templates, decorator-based intent routing, slot defaults/conversions, and request signature verification to name a few.
Please check out the samples directory for client code:
https://github.com/johnwheeler/flask-ask/tree/master/samples
These are direct ports from the Java samples:
https://github.com/amzn/alexa-skills-kit-java/tree/master/sa...
Yours look much simpler to get started, would you care to comment on what the differences are between yours and the project I just posted?
I just put mine up a month ago - not discoverable for "Python Alexa" or any keywords. Working on it :-)
So for the differences -
Alexa Skills are deployable as AWS Lambda functions or behind HTTPS. Currently, ask-alexa-pykit works on Lambda and Flask-Ask implements the signature verification required for HTTPS deployments. (i.e. Flask-Ask works on your own HTTPS server or Lambda).
Another difference is in the intent mapping design. Flask-Ask is based on the same architectural patterns of Flask with context locals, parameter mapping / conversion, and of course, Jinja templates!
For example, Mapping an intent with ask-alexa-pykit looks like this:
http://pastebin.com/raw/hQJLKnHL
Flask-Ask is like this:
http://pastebin.com/raw/9fWrGNYY
Flask-Ask also converts slots like firstname from the example above into arbitrary datatypes, and has stock conversions for AMAZON.DURATION (e.g. 'P2YT3H10M' into a Python datetime.timedelta). Full parameter mapping docs here: https://johnwheeler.org/flask-ask/requests.html#mapping-inte...
Flask-Ask templates are grouped together in the same files since utterances are typically small phrases--to make them easier to manage. Templates are of course an optional feature but are encouraged!
It's still very early, but I'm working my butt off, full-time on it! I have a 5-min tutorial that shows how to get up and running with Flask-Ask and ngrok: https://www.youtube.com/watch?v=eC2zi4WIFX0 - The API has changed a little, if you try it out and have any questions, you can do an issue or hit me up! john at ! johnwheeler.org
Thank you!
For example, I've been playing with home automation and speech recognition, and have been able to get any Sphinx based recognizer working in a single sitting, in a few hours or less. But I've yet to get Kaldi working yet after a several nights of effort. It seems much more powerful, and based on my reading, it's more accurate than Sphinx. But that doesn't do me any good if I can't get it to run, haha.
IMHO, there is a big opportunity for someone to come along and repackage it in a user friendly way, but the people who actually understand it are too busy doing "real work" to bother with such frivolity.
We use Speechmatics (and sphinx) for caption alignment and timestamps. We certainly recommend them if that’s the service you need.
Raw processing power will be the bottleneck on a Raspberry Pi.
The trick with pocketsphinx is to limit the vocabulary you want to recognize, and create a corpus of the types of things you want to be able to recognize and feed it through here: http://www.speech.cs.cmu.edu/tools/lmtool-new.html
If you try to use pocketsphinx to recognize arbitrary English (e.g. dictation) it's not going to work very well in my experience.
For example that has the ability to recognise anger / happy / questioning to relative accuracy?
But seriously, I've seen emotion recognition, and speech recognition, but haven't come across anything that provides both in one package.
Perhaps, could use wrapper library to combine functionality. It's a very interesting area.
https://www.informatik.uni-augsburg.de/lehrstuehle/hcm/proje...
It does recognize separate words, but, IMO, the biggest flaw is that it only recognizes pre-defined commands.
I'd love to mix and match NPL libraries, voice synthesis, voice identification, and speech recognition to make a comfortable "User Interface" to some systems in my house.
I think it'd be a fun project, but nothing seems to be able to take arbitrary audio streams and give me a "User identification" based on voice patterns and also arbitrary spoken text.
I know, yes, this is a VERY tall order, but it something that should be possible. At the very least, the identification part isn't needed. It's just important that it works offline and provides a text stream.
Kaldi is much better, but very difficult to set up.
None of the open source speech recognition systems (or commercial for that matter) come close to Google.
Takes the pain out of it
Is that because of the data they have, or because of their superior algorithms?
If, suddenly, someone would apply the fact that copyright bans remixes to training of neural networks, and apply the fact that licenses for this have to be granted explicitly, Google would lose 90% of their advantage over other companies.
Personally, I’d be for making a requirement that companies open source their trained models if the training data contained data supplied by users, not paid employees.
I guess we need something new and shiny here in the open source space, probably based on neural networks