51 karma · joined January 10, 2026
It can support up to 99 but the subtitle generation quality tends to decreases with more obscure languages as they wouldn’t have been in the training data
The other thing is the speech recognition models are very sensitive to the audio. If the audio quality is good then the model would do a good job, I’m not sure how much experience you have with Whisper Large but it is very capable on normal speech. The issues arise when there are many competing sounds overriding each other
ASR in LR is down often from my experience.
Yeah that is my suspicion too, I am considering adding an option to transcribe system audio too so it could hook into whatever video or audio is playing on the computer. That would be quite a big change so would like feedback first! hahah