But for other languages, especially those with a variety of dialects (like Arabic), the output can be really hard to read. Interestingly, gpt-3.5-turbo can turn a garbled Arabic call transcript into a decent English summary. I guess this is due to the amount of redundancy/repetition in phone calls.
I wonder whether anyone knows of something better than Whisper Large v2/v3 for Arabic ASR?
I had a look recently, and what I found was:
- Among published multilingual models (i.e. ignoring English-only models), it seems that Whisper Large has the lowest (best) overall WER (word error rate).
- Although Facebook published a multilingual model (mms-1b-all) that outperforms Whisper on many less common languages, its Arabic WER on standard benchmarks is much worse than that of Whisper, so I don't think it's worth trying on my data.
- There's a paper from ~6 months ago by some researchers who claimed to get slightly better performance than Whisper vanilla, but their code is not public, and I can't find any blog posts or articles talking about using their work.
- Googling turned up some attempts to finetune mms-1b-all with Arabic, but nothing I found included any WER data, so I assume these attempts didn't work out.