Microsoft has a tool that accepts wav or mp3 and transcribes it.
But I do not think it can distinguish between speakers.
How well does Whisper work in terms of correctness for single speakers?
But I do not think it can distinguish between speakers.
How well does Whisper work in terms of correctness for single speakers?
fine-tuning whisper is a nightmare, I don't know what the interviews are for, but again most enterprise STTs offer customization. you can add medical terminology.
---Google, Amazon and Nuance have medical models but either expensive or not available for personal projects.