This looks huge. Anyone know how this compares with Whisper in terms of quality and speed?
[1] https://ai.facebook.com/blog/multilingual-model-speech-recog...
Edit: Just checked the paper, it seems to be worse[1][2] but feel free to correct me.
I feel like they should've just taken the Whipser architecture, scaled it, and scaled the dataset as they did.
[1] Page: https://i.imgur.com/bq15Tno.png
[2] Paper: https://scontent.fcai19-5.fna.fbcdn.net/v/t39.8562-6/3488279...