I just wonder what system requirements Whisper has and whether there are open source voice recognition models that are specifically built for embedded devices.
I just wonder what system requirements Whisper has and whether there are open source voice recognition models that are specifically built for embedded devices.
Edit: According to this comment[0] the base model runs in real time on an M1 CPU. The tiny model apparently decodes an audio file twice as fast. These are promising results.
You could use really small chunk sizes and process them in a streaming fashion, but that would impact accuracy, as you're significantly limiting available context.
https://mycroft-ai.gitbook.io/docs/using-mycroft-ai/customiz...
I'm currently trying to setup a deepspeech server on my raspberry pi to see if it works ok for commanding spotify.
Edit: just realised you said `TTS` not `STT`
The Mycroft has done a lot of cool and important work in the field to ship an actual personal assistant product (stuff like wake word detection).
[0]: https://github.com/Sheepybloke2-0/trashbot - It was called trashbot because the final implementation was going to look like oscar the grouch in a trashcan displaying the reminders.
Almond is also interesting as a voice assistant, though I think it doesn't perform speech recognition itself.