Pure C, CPU-only inference with Mistral Voxtral Realtime 4B speech to text model
github.com
github.com
I tried incorporating this Voxtral C implementation into Handy but got very slow transcriptions on my M1 Max MacBook 64GB.
[1] https://github.com/cjpais/Handy
I’ll have to try the other implementations mentioned here.
The MacOS built-in dictation streams in real time and also does some cleanup, but it does awkward things, like the streaming text shows up at the bottom of the screen. Also I don’t think it’s as accurate as Parakeet V3, and there’s a start up lag of 1-2 secs after hitting the dictation shortcut, which kills it for me.
I feel like this is the best of both worlds.
Perhaps a little janky with backspaces, but still technically feasible.
https://github.com/kitlangton/Hex
Faster than handy and uses way less memory.
In the end Omarchy's new support for voxtype.io provided the nicest UX, followed by Whisper.cpp, and despite being slower, OpenAI's Whisper is still a solid local transcription option.
Also very impressed with both the performance and price of Mistral's new Voxtral Transcription API [2] - really fast/instant and really cheap ($0.003/min), IMO best option in CPU/disk-constrained environments.
[1] https://llmspy.org/docs/features/voice-input
[2] https://docs.mistral.ai/models/voxtral-mini-transcribe-26-02
(I wasn't able to find anything at glance)
Handy claims to have an overlay, but it seems to not work on my system.
Although as llms-py is a local web App I had to build my own visual indicator [2] which also displays a red microphone next to the prompt when it's recording. It also supports both Tap On/Off and hold down for recording modes. When using voxtype I'm just using the tool for transcription (i.e. not Omarchy OS-wide dictation feature) like:
$ voxtype transcribe /path/to/audio.wav
If you're interested the Python source code to support multiple voice transcription backends is at: [3]
[1] https://learn.omacom.io/2/the-omarchy-manual/107/ai
[2] https://llmspy.org/docs/features/voice-input
[3] https://github.com/ServiceStack/llms/blob/main/llms/extensio...
(I keep coming back to this one so I've got half a dozen messages on HN asking for the exact same thing!).
It's a shame, whisper is so prevalent, but not great at actual streaming, but everyone uses it.
I'm hoping one of these might become a realtime de facto standard so we can actually get our realtime streaming api (and yep, I'd be perfectly happy with something just writing to stdout. But all the tools always end up just batching it because it's simpler!)
[1] https://github.com/peteonrails/voxtype/blob/main/docs/WAYBAR...
And not to forget, for many use cases more than just English is needed. Unfortunately right now most STT/ASR and TTS focus on English plus 0-10 other languages. Thus being able to add with reasonable effort more languages or domain specific vocabulary would be a huge plus for any STT and TTS.
--from-mic only supports Mac. I'm able to capture audio with ffmpeg, but adapting the ffmpeg example to use mic capture hasn't worked yet:
ffmpeg -f pulse -channels 1 -i 1 -f s16le - 2>/dev/null | ./voxtral -d voxtral-model --stdin
It's possible my system is simply under spec for the default model.
I'd like to be able to use this with the voxtral-q4.gguf quantized model from here: https://huggingface.co/TrevorJS/voxtral-mini-realtime-gguf
(But take with grain of salt; I haven't tried yet)
Given that it took 19.64 mins to transcribe the 11 second sample wav, it’s possible I just didn’t wait long enough :)
I can, for example, capture audio from that with Audacity or OBS Studio and do it later, so it should be possible to do it in real time too assuming my machine can keep up.
Cool project!
Any ideas from the HN crowd currently involved in speech 2 text models?