Problem with whisper is its not really optimized for command recognition versus general dictation.
- Whisper processes 30 second audio chunks. So if you process 5 seconds of audio you have to pad it out with 25 seconds of silence. Hence a loss of efficiency with wasted CPU / GPU cycles on 25 seconds per chunk in the case above.
- Whisper most likely can't handle hundreds of commands much less than a thousand performantly.
- Whisper doesn't handle short commands very well with a degree of accuracy post processing commands from free dictation utterances.
Command dictation should be weighted higher than general dictation when decoding.
I work with a little under 1500 of commands dragon naturally speaking. DNS is hot garbage as a program despite it has the best accuracy to date with the feature of commands and dictation in one utterance. You get to pay $750 for the privilege m
I've yet to see a free and open source speech recognition engine that can handle both dictation and commands with a high degree of accuracy.
Please please let me know if there's alternatives out there. I would definitely pay to support an open source project like this that focuses on command and dictation.
Most solutions out there that are open source nowadays focus so much on iot command recognition with intents. That's not well suited for controlling your computer with grammars containing voice commands.