Rhasspy is an open source, fully offline voice assistant toolkit
rhasspy.readthedocs.io
rhasspy.readthedocs.io
Bit of background: Rhasspy was originally designed for Home Assistant (https://www.home-assistant.io), but now works with lots of home automation projects (Hass.io, Node-RED, OpenHAB, Jeedom). Its sister project, voicej2son (http://voice2json.org), is for command-line use and has fewer options.
With Snips.ai being bought by Sonos, we're now focusing on compatibility with its MQTT protocol (https://docs.snips.ai/reference/hermes) so existing plugins/skills will just work. Supporting Snips-like number/duration/dateTime slots across over a dozen languages is going to be a major challenge, so please reach out if you speak a language besides English* :)
* Also consider donating to the Common Voice project: https://voice.mozilla.org
They’re pretty decent. Solid range. The setup is sorta ok. Requires a bit of googling and tweaking
* Debian-Based Linux System
* SDK for Speech Algorithms with Full Documents
* C++ SDK and Python Wrapper
* Speech Algorithms and Features
* Keyword Spotting (Wake-Up)
* BF (Beamforming)
* DoA (Direction of Arrival)
* NS (Noise Suppression)
* AEC (Acoustic Echo Cancellation) and AGC (Automatic Gain Control)
* All-in-One Solution with High Performance SoC
* 8 Channel ADC for 6 Microphone Array and 2 Loopbacks (Hardware Loopback)
Does anyone know if there's a way to hack an Echo Dot and use it as the speaker/mic for Rhasspy? Rolling out our own hardware that is as effective as a Dot would probably be very difficult?
Long Answer: Not even kinda close to a way to do this, that hardware is locked down good. If you can brute force the key that locks the adb/fastboot then you have a chance.
Some of the echo's run android, so you'd need to make Rasspy run on android for those versions of the echo. Alternatively, you'd have to find the versions of the echo that runs a proper linux flavor, you might have a chance there.
If you really start messing with the operating system or system software, you have to make sure that you can access the mics in software. (some of) their mic array and ADC array feature FPGAs that handle audio manipulation, so you'd need drivers/whatever to interface with those.
Website: https://voice2json.org
GitHub: https://github.com/synesthesian/voice2json
(I am not affiliated but am using it in my own pet project)
For those wondering, Rhasspy and voice2json are from the same author (me). If you want a command-line tool for voice assistant tasks (wake word detection, speech to text, intent recognition, etc.), check out voice2json.
See the recipes for some interesting things you can do with voice2json: http://voice2json.org/recipes.html
This should make it much easier to swap out parts, and distribute the computing across multiple devices.
https://hacks.mozilla.org/2019/12/deepspeech-0-6-mozillas-sp...
It would still be much better than pocketsphinx.
Rhasspy is designed to recognize user-specified voice commands, so the accuracy will highly depend on the complexity of your commands. If needed, you can also try doing open transcription: https://rhasspy.readthedocs.io/en/latest/speech-to-text/#ope...
Sounds like it's not too far off.
I guess something similar should be possible with Rhasspy?
I was part way through a "smart speaker" project and planned to use Snips.ai, but I see now that they've been bought by Sonos so Rhasspy is looking pretty tempting now.
However my plan was to use pi zero's at the speaker end, with my beefier HA machine doing the speech recognition.
The recommended use of a Pi Zero is as a recording/wake word detection/audio playback satellite. Other functions, like speech/intent recognition can be done remotely (e.g., https://rhasspy.readthedocs.io/en/latest/speech-to-text/#rem...)
I'm curious about extensibility - would it be possible integrate with a C# app running on Windows, for example?
I'm particularly interested for accessibility reasons, looking for ways to control tools like JetBrains Rider without shifting my hands from keyboard to mouse.
I think this project would really benefit from taking one of the excellent existing voice assistant/speakers on the market (Google home, echo dots, etc., and flashing them with some custom firmware.
Would love to have great hardware for custom tinkering though. With more development time and money behind this project, I hope it could grow into a great tool.
Arguably, being offline and keeping your recordings off the cloud, it is already superior to commercially available options.
I've just looked over the docs. I'll probably be playing with this very soon.
If the primary issue is the same as me, that you download too much crap without paying attention to available storage, just use the terminal emulator to create a 25-100mb file you can delete when necessary.
[1] storage ️ internal ️ apps [2] If you use PWAs to cache data and avoid high data bills, keep in mind that clearing browser cache clears PWA cache as well (my most-used PWAs are for hn and xkcd)