Is insanely-fast-whisper fast enough to actually run on the CPU and still trascribe in realtime? I see that none of these are running quantized models, it's still fp16. Seems like there's more speed left to be found.
Edit: I see it doesn't yet support CPU inference, should be interesting once it's added.