[1]: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
[1]: https://huggingface.co/spaces/hf-audio/open_asr_leaderboard
I'm actually a little surprised they haven't added model size to that chart.
And just to be clear, 500MB is even enough for a raspberry Pi. Then your problem is not memory, is FLOPS. It might run real-time in a RPi 5, since it has around 50 GFLOPS of FP32, i.e. 100 GFLOPS of FP16. So about 20-50 times less than a modern iPhone. I don't think it will be able to keep it real time, TBF, but close.
regardless, this model with such quantization strategy runs real time at +10x real-time factor even in 6-year old iPhones (which you can acquire for under $200) and offline at a reasonable speed, essentially anywhere.
You get the best of both worlds: the accuracy of a whisper transformer at the speed and footprint of a small model.
Oh and I type this in handy with just my voice and parakeet version three, which is absolutely crazy.
And handy even takes care of all the punctuation, which is really nice.
Thanks a lot for suggesting it to me. I actually wanted something like this, and I was using something like Google Docs, and it required me to use Chrome to get the speech to text version, and I actually ended up using Orion for that because Orion can actually work as a Chrome for some reason while still having both Firefox and Chrome extension support. So and I had it installed, but yeah.
This is really amazing and actually a sort of lifesaver actually, so thanks a lot, man.
Now I can actually just speak and this can convert this to text without having to go through any non-local model or Google Docs or whatever anything else.
Why is this so good man? It's so good
man, I actually now am thinking that I had like fully maxed out my typing speed to like hundred-120. But like this can actually write it faster. you know it's pretty amazing actually.
Have a nice day, or as I abbreviate it, HAND, smiley face. :D
That's unfortunate. I think I can update my version but I have heard some bad things about performance from the newer update from my elder brother.
Newer than Sequoia, you mean?
The brew recipe [1] says macOS >= 15.
Anyway, I'm on Sequoia — it's mostly better than Ventura, which was what my M2 MacBook Pro came with. I'm holding off upgrading to Tahoe (macOS 26), hoping they fix liquid glAss.
I can tell that this is now definitely going to be my go-to model and app on all my clients.
The one built in is much faster, and you only have to toggle it on.
Are these so much more accurate? I definitely have to correct stuff, but pretty good experience.
Also use speech to text on my iphone which seems to be the same accuracy.
One note for anyone using Handy with codex-cli on macOS: the default "Option + Space" shortcut inserts spaces mid-speech. "Left Ctrl + Fn" works cleanly instead. I'm curious to know which shortcuts you're using.
edit: holy shit parakeet is good.... Moonshine impressive too and it is half the param
Now if only there was something just as quick as Parakeet v3 for TTS ! Then I can talk to codex all day long!!!
Very lightweight and good quality
I tried comparing Parakeet streaming with Moonshine streaming. Moonshine is smaller, and I felt it was subjectively faster with about the same level of accuracy.
If e.g. parakeet can be run on my phone in real time showing the transcript live:
- with latency low enough to be "comfortable enough" for the instructor to keep an eye on and approve the transcribed instructions
[not necessarily every word of the transcript, i.e., a commanded "edit" doesn't need to be applied in the outcome as long as it's nature is otherwise clear enough to not add meaningful amounts of ambiguity to the final "written" instructions]
by glancing at the screen while dictating the explanation (and blurting out any transcription complaints as soon as that's possible without breaking one's own string-of-thought or spoken grammar too much)
, I'd very happily switch to that approach instead of what I was doing.
Bonus if there's a no-bulky-or-expensive-hardware way to accommodate us both speaking over each other so I won't have to _interrupt_ his speaking just to put a clarifying comment (on what he just said) in the transcript for him to see and sign off, where the at least "only" briefly interrupts his thoughts right while he actually reads my transcribed words (he doesn't have to hear them, and it's better if he won't; I can probably get him to put on earmuffs to not hear me louder than he hears his thoughts, and a sufficiently-smoothed SNR meter for specifically his voice should take care him regulating his volume while the earmuffs mute it and I occasionally talk over him)...
i was using assmeblyAI but this is fast and accurate and offline wtf!
On Mac, I've been using VoiceInk and it's even better. VoiceInk (and MacWhisper too, IIRC) use the neural engine and the delay between dictation and appearance of the typed text is almost imperceptible.
I think most apps that use Parakeet tend to use this version of the model?
See if Parakeet (Nemotron) still uses 4GB+ with my implementation: https://rift-transcription.vercel.app/local-setup