For example, instead of:
Hello there!
Hi
How are you?
Good, and you?
You get something like: Voice A: Hello there!
Voice B: Hi
Voice A: How are you?
Voice B: Good, and you?For example, instead of:
Hello there!
Hi
How are you?
Good, and you?
You get something like: Voice A: Hello there!
Voice B: Hi
Voice A: How are you?
Voice B: Good, and you?You can do this pretty conveniently using pyannote-audio[0].
Coincidentally I did a small presentation on this at a university seminar yesterday :). I could post a Jupyter notebook if you're interested.
PS: Bai & Zhang (2020) is a great review on the literature [1]
This is the best AI diarization and transcription I’ve been able to get so far: https://github.com/zachlatta/openai-whisper-speaker-identifi...
> The Google Recorder app (...) transcribes meetings and interviews to text, instantly giving you a searchable transcription that is synced to the recorded audio (...) With this new update, the recorder can now identify and label each speaker automatically—an impressive feat. Google Recorder is exclusive to the Pixel 6 and newer Pixel devices.
https://arstechnica.com/gadgets/2022/12/pixel-7-update-adds-...
would be lovely to see this feature open sourced.
It’s ok, but the quality of speaker identification is nowhere near as good as the transcription itself.
I’d love to see models which try and use stereo information in recordings to solve the problem. Or, given a fixed camera and static speakers, I thought it should even be possible to use video to add information about who is speaking. There doesn’t seem to be anything like that right now tho.