[1]: https://github.com/alphacep/vosk-api/issues/55
[2]: https://github.com/outlines-dev/outlines?tab=readme-ov-file#...
[1]: https://github.com/alphacep/vosk-api/issues/55
[2]: https://github.com/outlines-dev/outlines?tab=readme-ov-file#...
It's interesting as speech recognition has become more popular than ever through services like Alexa, and other iot devices support for OS speech recognition has very little development. Don't get me started with accessibility apis either...
Unfortunately most implementations (especially those that are iot focused) don't have very important features for robust speech recognition.
1. Ability to enable and disable a grammar
2. Modify grammars while the engine is loaded
3. Scoped grammars that are context-specific
4. Recognition callbacks
5. Multiple grammars active simultaneously.
Unfortunately I don't think vosk api will ever support those features. I know there's a few PRs that address a few of those points but have not been merged for years.
Given the criteria above there's very little open source that allows for complex grammars that's easy to run for an end user locally.