Offline Voice Assistant on a Microcontroller with 192KB RAM
picovoice.ai
picovoice.ai
Good luck for your company, it sure is hard to make one grow.
Edit: removed the word misleading before advertisent. And explained what i meant after the word product.
> Picovoice engines call home servers to stay active and report the consumption for billing purposes only.
That's not offline.
[1] https://www.st.com/en/evaluation-tools/stm32f4discovery.html
I use a picovoice wake-word on my DIY offline Assistant myself (RPI-based) and I was tempted to dig up a STM devkit from somewhere but then I remembered that I probably would want a microphone array for it if it works well and I don't see a good way to integrate that.
When you bootstrap a small business, most likely nobody knows about you. Everyone boasts about how awful ads are so you try to please them and get your product in the eyes of users by doing something useful: write technical articles, provide free/open-source solutions to real problems etc.
And then someone comes in the comments just to point out what a bad person you are for doing this.
What is one supposed to do? Publish a website and just hope that people will end there somehow?
Good job in getting this working on the STM32 board! I have mostly given up on voice assistants because of the latency and rigid phrasing required.
Your solution might seem rigid too for others but if we can use our own phrases and make it respond instantly because the model is local, the interaction can finally become as easy as pressing a button without touching it.
You've taken on a tremendous task, I really wish you succeed!
Edit: Just to be clear, the poster _literally_ only ever posts about their product and only comments on their own product posts.
Picovoice here solves both the issues of hotword/wake-word detection and intent extraction. This looks like something you could build on top of ARM's [keyword spotting program](https://github.com/ARM-software/ML-KWS-for-MCU) and the wake word services listed in [Rhasspy's docs](https://rhasspy.readthedocs.io/en/latest/wake-word/#raven)
But, to implement something like this from scratch would take a good while.
Also, here are some Automatic Speech Recognition toolkits (which won't run offline on a microcontroller) out there. These are useful to pipe the data into a program that deals with intents (something like [RASA](https://rasa.com)
(Require Internet) * [Deepgram](https://deepgram.com) - I believe they build upon OpenAI's Whisper model and have their own custom models too * Google Cloud / Microsoft Azure / AWS / IBM Watson
(Can be run Offline) * [OpenAI's Whisper](https://github.com/openai/whisper) * [nVidia's NEMO](https://github.com/NVIDIA/NeMo) * [PaddleSpeech](https://github.com/PaddlePaddle/PaddleSpeech)
When you see how complicated the space is and how many ways you can actually shoot yourself in the foot. This post starts to look a wee bit better.
[1] https://www.st.com/en/evaluation-tools/stm32f4discovery.html
19. Why do Picovoice engines require an internet connection?
While data is processed offline, locally on-device, Picovoice engines call home servers to stay active and report the consumption for billing purposes only.
[1] https://picovoice.ai/pricing/I read the article and was interested but after clicking around the site I was thoroughly confused about billing/pricing/metering.
> Picovoice engines call home servers to stay active and report the consumption for billing purposes only.
[1] https://www.st.com/en/evaluation-tools/stm32f4discovery.html
[1] https://picovoice.ai/blog/end-to-end-intent-inference-from-s...
You might not understand but there is a huge amount of work behind this simple demo.
0. Thanks, but no thanks.
No offense, but I expect an offline service to be really offline from start to finish.