[1] http://snips.ai/
[1] http://snips.ai/
With Picovoice you can use the voice control engine to accomplish this. Maybe something similar to this demo?
https://picovoice.ai/#voice-control-demo
The cool thing about this engine is that it is tiny. It uses less than 8% CPU on RPi3 and altogether it is less than 2MB (code, model, etc). Technically you can run it on something much smaller and cheaper than RPi.
Alternatively, the speech-to-intent engine could be a good candidate. More information on this along with an interactive demo will be released this weekend.
We do work with a couple of SoC manufacturers and will disclose some of the results when our partners are ready. In general, we can run on any MCU with a C compiler and 200KB of RAM (maybe less if there is fast FLASH available). We already of models working on ARM Cortex-M and Cadence's HiFi4.
> Runs in real-time with only 5.6 MB of memory and 25% CPU usage on a Raspberry Pi 3.
The voice control comes in two variations standard and tiny. The tiny one consumes even fewer resources. I provided metrics for the standard one. You can check the benchmark repo on benchmark it yourself as well :) https://github.com/Picovoice/wakeword-benchmark
Snips looks like a promising prospect!