We believe that people can have both cloud-grade Voice AI (look at our accuracy at https://medium.com/snips-ai/snips-brings-cloud-level-spoken-...) which is privacy-aware
Everything runs 100% on-device and is private-by-design, and we are open-sourcing the platform over time
It works in english, french, german, japanese, italian, spanish (and a lot more languages coming soon!)
Take a look at our blog, and we'd be delighted to know what you build with our platform!
Browsing your site it is hard to see what you will charge for, and I am confused by the inclusion of some sort of token system. (It seems directly in conflict with the desire not to be beholden to an outside company to use the hardware if they are tied together?)
Mycroft does most of the Alexa parlor tricks such as weather and wikipedia lookups out of the box. If you use the plugins you can integrate it with Home Assistant, Kodi but I've had mixed luck.
They send the voice to Google for speech to text, but I believe that is configurable and they are working on a personal server which theoretically could operate entirely locally.
It'd be simple enough to hook this up to Philips Hue or whatever to do what you want.
There are numerous Sphinx language bindings. I went for Ruby via Isabella https://github.com/chrisvfritz/isabella. I used this because it provided a framework for what I wanted: define a simple grammar (JSGF, Java Speech Grammar Format) and call specified script(s) with the parsed results. The hardest part was probably mapping out the phonemes for the grammar atoms. (If your target language isn't English, you may be out of luck).
This worked really well for what I needed (directing band-in-a-box from the other side of the room) but still needs a little tuning. Even with a leading activation token ("Hey Isabella...") she sometimes gets confused and thinks she's been summoned when it's just some random song playing. Choosing a concise, simple, unambiguous grammar was helpful, along with sensitivity adjustment. There are other knobs to twiddle -- as evidenced by the academic mailing list activity -- but I didn't need to look closer for my simple use case.
It was a fun little project and the kids liked it, especially paired with a text-to-speech module (tts gem under Ruby): "Hey Isabella, am I <adjective>?" (or "Is <sibling> <adjective>?"), and a randomly generated response :)
I ended up using a set of JSGF grammars (one per intent) to generate a statistical language model for use with pocketsphinx. Rhasspy also features a web-based interface for creating custom words -- I have a mapping from Sphinx phonemes to eSpeak phonemes so you can iterate over a pronunciation until it sounds right.
As you mentioned, the wake/hotword stuff with Sphinx isn't terribly robust. I've been Docker-izing Mycroft Precise (https://github.com/MycroftAI/mycroft-precise) to address this.
I haven't started building anything yet but I will in the new year (fingers crossed).
This is an external service I wrote that does offline speech/intent recognition with pocketsphinx, and forwards structured JSON events into Home Assistant for use in automation scripts.
Clap on, clap off, THE CLAPPER.