I gave it a spin a little bit ago. Per usual, install docs didn't quite work OOTB, here's how I got it working: https://llm-tracker.info/books/howto-guides/page/speech-to-t...
One limitation that seems undocumented, the current code only supports relatively short clips so isn't suitable for long transcriptions:
> ValueError: The input sequence length must be less than or equal to the maximum sequence length (4096), but is 99945 instead.