It runs but any audio input (you will need to provide wav not mp3's) I tried (tried 20s/40s/300s) I get just one short sentence returned in target language that seems not related at all to my audio input (i.e. Tous les humains sont créés égaux).
Seems like some default text but it runs on full GPU for 10 minutes. Tons of bug reports in GitHub as well.
Text Translate works but not sure what is the context length of the model. Seems short at first glance (haven't looked into it).
Oh and why is Whisper a dependency? Seems not need if FB has their own model?