Transcription speed and accuracy keeps going up and it’s delightful to see the progress, I wish though more effort was dedicated to creating integrated solutions that could accurately transcribe with speaker diarization.
Author of Wordcab-Transcribe here. We use faster-whisper + NeMo for diarization, if you want to take a look.
https://docs.nvidia.com/deeplearning/nemo/user-guide/docs/en...