Year of the Voice – Chapter 2: Let's talk
home-assistant.io
home-assistant.io
Edit: if you want to keep in the loop of the work we're doing, subscribe to our free monthly newsletter @ https://building.open-home.io/
I hope you saw the recent demo of a ChatGPT plug-in for HA that might add a lot of functionality to HA if there is the ability to use it to construct action plans and then execute them.
https://www.reddit.com/r/homeassistant/comments/12l10se/intr...
Fun fact: we supported Almond (later Genie), the IoT AI from Stanford which they released a couple of years ago. It would convert natural language into their ThingTalk language, which could do things like you mention. It was not popular at all and Stanford ended up shutting it down.
I also like that you are directly supporting ESPHome and I'll definitely take that route and build something nice.
However, for the more casual users, are you thinking about releasing an off-the-shelf "voice assistant" device?
We're still exploring different ML on the edge chips to help with wake word and audio processing. One hard requirement that we have is that the tooling is openly available, so that anyone can convert an AI model and run it on the chip. We don't accept tooling to be available as a website or under NDA. That challenge is harder than we expected!
For now, though, you may want to check out OVOS: https://openvoiceos.com/
I have a general question about HA: the main issues I’ve had with it have had to do with specifically only two points.
1. The recorder module has brought the system to a crawl when a single rogue device on my network kept spamming MQTT with status updates about power consumption and eventually killed the SD card it was running on. Some aggressive ignore statements in the config fixed the issue but I still see this as a major pain point as there is no indication of anything happening. Is there any plan to introduce any kind of housekeeping into the recorder module besides periodic purges (which in my case were not enough to actually keep the system running)?
2. I like to have control over my computers so I did not want to use a pre-installed image of HA for RPi. But the alternative is a cumbersome process of installing it in a Python virtual env and keeping things updated which is not ideal. Are there any plans for improving this installation path? Or alternatively some sort of advanced version of the HA image which would allow someone like me full control over the base OS while keeping the benefits of the cohesive self-updating HA distribution?
2. We offer VM images if you want to keep flexibility but also benefit from all our work on an integrated system. You can also do a Supervised installation, but then it's your own responsibility to keep the host OS running the requirements of HA.
Works really well, updating is low effort (but not automatic) but you could probably set something up if you wanted to pretty easily.
Wondering if you all are thinking about multiple remote microphones, and choosing which microphone will respond to the speaker?
Ideally, I'd love to have multiple microphones/speakers throughout the house, all of which can listen and answer, however, only the one that hears me the clearest/loudest/etc. actually answers.
Make sense?
My entity names can contain the English or other text that just won't be picked up by the speech to text, so if there was a way to record a couple of pronunciations for an entity name, and home assistant would fine-tune whisper on them in the background it would be wonderful !
If I toggle a light I wouldn’t want it to execute the command twice haha
I love the VOIP integration shown off that can hook up to an old phone. One of my guilty pleasures is using peak forms of technology from the 20th century when things were more analog. It could be a lot of fun to bring an old phone into the mix to complement my turntable and PVM.
The quality, variety & diversity of voices that synesthesiam's "Larynx" TTS project (https://github.com/rhasspy/larynx/) made available, completely transformed the Free/Open Source Text To Speech landscape.
In addition "OpenTTS" (https://github.com/synesthesiam/opentts) provided a common API for interacting with multiple FLOSS TTS projects which showed great promise for actually enabling "standing on the shoulders of" rather than re-inventing the same basic functionality every time.
The new "Piper" TTS project mentioned in the article is the apparent successor to Larynx and, along with the accompanying LibriTTS/LibriVox-based voice models, brings to FLOSS TTS something it's never had before:
* Too many voices! :)
Seriously, the current LibriTTS voice model version has 900+ voices (of varying quality levels), how do you even navigate that many?![0]
And that's not even considering the even higher quality single speaker models based on other audio recording sources.
Offline TTS while immensely valuable for individuals, doesn't seem to be attractive domain for most commercial entities due to lack of lock-in/telemetry opportunities so I was concerned that we might end up missing out on further valuable contributions from synesthesiam's specialised skills & experience due to financial realities & the human need for food. :)
I'm glad we instead get to see what happens next.
[0] See my follow-up comment about this.
I'm interested in Text to Speech for creative pursuits, such as video game voice dialogue and animated videos.
This is one of the reasons why the range & quantity of available voices is particularly important to me.
After all, you can't really have scene set in a board room with nine characters[3] if you've only got three voices to go around. :)
I've actually been spending time this week on updating my "Dialogue Tool"[1] application (originally created to work with Larynx to help with narrative dialogue workflows such as voice "auditioning", intelligent caching & multiple voice recordings) to work with Piper.
Which is where I ran into the question of how to navigate/curate a collection of more than 900+ voices.
The main approaches I'm using so far are:
(1) Random luck--just audition a bunch of different voices with your sample dialogue & see what you like.
(2) Curation/sorting based on quality-related meta-data from the original dataset.
(3) Generating a different dialogue line for each voice that includes their speaker number for identification purposes that also (hopefully) isn't tedious to listen to for 900+ voices. :)
I haven't quite finished/uploaded results from (3) yet but example output based on approaches (3) & (2) can be heard here: https://rancidbacon.gitlab.io/piper-tts-demos/
The recording has two sets of 10 voices which had the lowest Word Error Rate scores in the original dataset--which doesn't mean the resulting voice model is necessary good but is at least a starting point for exploring.
I'd also like to explore more analysis-based approaches for grouping/curation (e.g. vocal characteristics such "softer", "lower", "older") but as I'm not getting paid for this[2], that's likely a longer term thing.
A different approach which I've previously found really interesting is to use voices as a prompt for writing narrative dialogue. It really helps to hear the dialogue as you write it and the nuances of different voices can help spur ideas for where a conversation goes next...
[1] See: https://rancidbacon.itch.io/dialogue-tool-for-larynx-text-to... & https://gitlab.com/RancidBacon/larynx-dialogue/-/tree/featur...
[2] Am currently available/open to be though. :D
[3] Will try to upload some example audio of this scene because I found it pretty funny. :)
Where can I find all these voices? https://github.com/rhasspy/piper/releases/tag/v0.0.2 lists "only" ~50 files.
I'm the author of Piper; it is a successor to Larynx (originally named Larynx 2). Piper uses the same underlying model as Mimic 3, which I developed before joining Mycroft. However, Piper uses a different library to get word pronunciations, so the voices aren't compatible between the two projects.
It's been an awesome year so far with Nabu Casa, and I'm very fortunate to be able to work on something I love. I hope to contribute to the open source voice space for many years to come :)
> On a Raspberry Pi 4, voice commands can take around 7 seconds to process with about 200 MB of RAM used.
Have you looked into supporting something like the Coral Accelerator[0], which can drastically speed up machine learning inference on a Raspberry Pi?[1]
It used to be available for $60, but it is hard to find in stock at the moment except for way over MSRP.
[0]: https://coral.ai/products/accelerator
[1]: https://www.hackster.io/news/benchmarking-machine-learning-o...
Is there better consumer stuff available now?
Amazon and Google have both done pretty good here, and I hope there's a good option to replace them in the future. Maybe it's possible to transplant an esp32 into one of those devices? Maybe sometime will start selling something.
https://salsa.debian.org/deeplearning-team/ml-policy https://blog.opensource.org/episode-5-why-debian-wont-distri... https://deepdive.opensource.org/podcast/why-debian-wont-dist...
I'm also going to take the opportunity to poo-poo this pet peeve:
>More powerful CPUs, such as the Intel Core i5, can generate 17 seconds of audio in the same amount of time.
Oh, really?
...Intel Core i5 brand microprocessors... Introduced in 2009...
About as accurate as saying "users of four-wheeled vehicles" when carriages were still around.
At the very least, you need to provide the node.
A pi 4 has a very specific performance to power profile. An "i5" has no specificity.
An underclocked i5 could either be a whisper or an inferno compared to this ARM chip.
It also mentioned useful agents like GPT 3.5 as a sort of distant afterthought, which is both paid and highly non-local.
I super recommend it, if you can afford the extra 10 to 15 watts of power.
The concept itself is great, and if I can finally have a good local only assistant, then that's fantastic.
And yes, one of the reasons people are more excited than ever is that the latest versions ChatGPT are actually really good.