Using the large model, it works really well, even in low volume settings/speakers mumbling. Some of my transcripts are pharma related and Whisper stumbles on the drug names, but I’m pretty understanding of that.
fine-tuning whisper is a nightmare, I don't know what the interviews are for, but again most enterprise STTs offer customization. you can add medical terminology.
---Google, Amazon and Nuance have medical models but either expensive or not available for personal projects.