Depends what you mean by “fast”.
I’ve tested WhisperLive, it’s basically real-time (i.e. low latency).
I’ve tested WhisperLive, it’s basically real-time (i.e. low latency).
Faster-whisper can process audio faster than real-time, but AFAIK vanilla Whisper needs a few seconds long audio "frame" to do inference/STT.
Whisper Live fixes that, and reduces latency to a few tens/hundreds of ms.
My Apple M1 MacBook (2021) can infer whisper-medium at roughly 10x realtime, for comparison. Takes about 20min to process three hours of audio.