Too bad that tool no longer seems to be developed. Looking for something similar. But it's really nice to see what's possible with local models.
By "any GPU" you mean a physical, dedicated GPU card, right?
That's not a small requirement, especially on Macs.
https://news.ycombinator.com/item?id=46640855
I find the model works surprisingly well and in my opinion surpasses all other models I've tried. Finally a model that can mostly understand my not-so-perfect English and handle language switching mid sentence (compare that to Gemini's voice input, which is literally THE WORST, always trying to transcribe in the wrong language and even if the language is correct produces the uttermost crap imaginable).
I run 120M Parakeet model formt STT thing. Even that tiny model works much better than macos dictation these days.
The downside is that couldn't get it to segment for different speakers. The concensus seemed to be to use a separate tool.
(and it did it perfectly without any edits required for me at all.)