Rhasspy – Offline private voice assistant for many human languages
github.com
github.com
Wanted to post this as currently Alexa and Mycroft are being discussed, but I have not seen a mention of this project. It is offline first, compared to Mycroft which relies on cloud services by default, even if you seem to be able to run it completely offline.
To share my personal experience with it, I am impressed with how well it can handle the basic commands I need (lights on/off, music control). It also features a reliable and easy satellite integration, so I can do the heavy lifting (STT and TTS) on my home server and only keep wakeword and sound input/output on the RPi. Different than Alexa, you define each command ("sentence") yourself, which seems to really help with the recognition accuracy. It also means you have no problem of discovering all available commands.
I have it integrated with Home Assistant for some basic automation. Essentially each sentence sends off a command to Home Assistant with the recognized text and an ID. You can also send back text that is then spoken by Rhasspy. For reference, I am using Porcupine as wake word engine ("Computer"), Kaldi as STT, Fsticuffs as intent recognition, and Larynx as TTS (voice "harvard"). The home server is a desktop PC, the satellite currently runs on an RPi 4 with about 2-10% load.
[1] https://news.ycombinator.com/item?id=21926027 [2] https://rhasspy.readthedocs.io/en/latest/ [3] https://community.rhasspy.org/
Ideally I can find a HAT with mics and speakers with sufficient quality/volume that I can build self-contained I/O devices for each room; the key really being that I want no more than a single power plug, and no more bulk than ~ a Pi + Housing.
I do not use the Rhasspy integration in HASS, I run the main node in a docker container. Rhasspy (main and satellite) is connected to the Mosquitto broker exposed by HAOS. The satellite pushes the STT, Intent Recognition and TTS via the Hermes MQTT handler to the main node over this connection. The acronym overload should become clearer when you start tweaking things in the Rhasspy web interface, it is very well done for the amount of customization you get.
The HASS side is done via yaml scripting. You can find my mpd.yaml at [2]. I select music to play via an ampd [3] interface hosted at the RPi connected to my HiFi (not the satellite, dedicated RPi). The sentences I use are at [4]. The names for HassTurnOff are the same as in HASS, this gets handled by built-in intents in HASS. I am considering trying out Node-RED, because I am a bit unhappy with the yaml.
I am also looking for a case. Currently I use [5] just hooked up to the ReSpeaker HAT 3.5mm and USB on the RPi with cables in a chaos on my desk. I also took apart [6] (smaller than you would expect, does not go back together nicely) and it might fit an RPi0W2 with HAT, if you take of its top. Both speakers are USB powered, so you can plug it directly into the RPi. I would prefer to stuff everything in the case of [5] and hook up the blue LED to a GPIO, but it is quite cramped in there, maybe it can "fit" by making a hole in the back.
[1] https://rhasspy.readthedocs.io/en/latest/tutorials/ [2] https://pastebin.com/M7KmQSiM [3] https://github.com/rain0r/ampd [4] https://pastebin.com/tsUqWwpy [5] https://www.amazon.de/gp/product/B07DDK3W5D [6] https://www.amazon.de/gp/product/B08GC8K8ZR
Most likely I'll do this the exact way you have, not using the HASS integration (I've found it limiting in several ways; particularly when using Frigate, as the HASS Appliance doesn't provide a way to mount storage).
I use Node-RED for a couple of things now and it's a really nice interface for creating complex interactions that would just be a bit of a nightmare to write up in YAML.
I've tried a couple of things so far, like the M5 Stack Atom speaker with ESPHome Media Player which meets the compactness requirement but audio quality is awful and the volume is so quiet it's practically rendered useless just by fan noise from computers in my home office.
Using a decent HAT, I might build something little out of wood that won't look out of place in the living room; my wife is ~~unlikely~~ not going to approve anything like (6). The ReSpeaker 2 with it's JST for a little speaker might be ideal!
Related and also cool: While exploring PeerTube last week, I stumbled across someone using a new voice assistant, Numen, on their PinePhone running SXMO. The code repository link is in the video description: https://peertube.plasmatrap.com/w/dYo8QcVVFFpzS7PwddFPMV
Offhand, I can't think of why this should be too hard, if it generates JSON or whatnot I can set up a watch folder or something?
[1] https://rhasspy.readthedocs.io/en/latest/intent-handling/#co...
I used voice2json to build a voice-controlled car audio player, it works amazingly well: https://github.com/lukifer/voicetunes
I personally like to run the satellite on an RPi, because I can easily use a microphone HAT and a cheap pair of speakers with it. I can also place it anywhere I want, not only at my desk. It also means I have no other sound sources and sinks on the same hardware, which makes the setup a bit easier. I do the main processing in a docker container on a desktop PC serving as my home server (Ryzen 5 3600X, but it handles far more than Rhasspy).
[1] https://rhasspy.readthedocs.io/en/latest/installation/ [2] https://rhasspy.readthedocs.io/en/latest/intent-handling/