Ideally it also exposes that parameter to the user.
Speed comparisons seem moot when quality is sacrificed for me, I'm working with very poor audio quality so transcription quality matters.
Ideally it also exposes that parameter to the user.
Speed comparisons seem moot when quality is sacrificed for me, I'm working with very poor audio quality so transcription quality matters.
Side note, the insanely fast whisper readme gives benchmarks on an A100 but only the FA2 lines were. The rest were on a T4 looking at the notebooks/history. Turing doesn't support FA2 so the gap should be smaller with it, but based on the distil-whisper paper CTranslate2 is probably still faster.
TensorRT-LLM might be faster but I haven't looked into it yet.
It's enabled by default with the latest Transformers version, so just make sure you have:
* torch>=2.1.1
* transformers>=4.36.0
I just reran the notebook with 4.36.1 (minus the to_bettertransformer line) but it was slower (the batch size 24 section took 8 vs 5 min). Is there something I need to change? Going back to 4.35.2 gives the old numbers so the T4 instance seems fine.
Insanely fast whisper (god I hate the name) is really a CLI around Transformers’ whisper pipeline, so you can just use that and use any of the settings Transformers exposes, which includes beam size.
We also deal with very poor audio, which is one of the reasons we went with faster whisper. However, we have identified failure modes in faster whisper that are only present because of the conditioning on the previous segment, so everything is really a trade off.
Just call the pipeline with:
result = pipe(sample, generate_kwargs={"num_beams": 5})