The only thing Whisper misses is speaker diarization. I'm currently working on a model that uses Whisper + pyannote to transcribe Interviews and also detects who is speaking. It's working but damn it takes so long
basically whisper.cpp has some support but its not great (based on my own testing)
- https://huggingface.co/spaces/vumichien/whisper-speaker-diar...
- https://github.com/Majdoddin/nlp pyannote diarization
- whisperX with diarization https://twitter.com/maxhbain/status/1619698716914622466 https://github.com/m-bain/whisperX
Edit: Looked at your link and I misunderstood. I think I understand you're waiting for the ChatGPT specific model now?
That's incorrect