Thanks for the reply and explanation. It is very helpful. Our app is old, started well prior to whisper. But we have updated it regularly as useful tech came along. I will check out CTC Viterbi and error-align !
6 karma · joined February 9, 2019
I think its similar because we both seek to assign audio time stamps to sentences and words.
But I wonder why your forced alignment algorithm is so heavy duty. (My head started to spin at CTC emissions. ) Probably yours is just way more thorough than mine,
My simplistic approach would have been to transcribe the audio. And then run a differencing script chapter by chapter matching the book text with the audio transcript. And then do something similar intra chapter to get sentence and word level time stamps.