Single Speaker Speech Recognition with Hidden Markov Models
kastnerkyle.github.io
kastnerkyle.github.io
I've got an audiobook with a particularly obnoxious commentator interjecting dumb things every couple of minutes and I'd like to slice those out.
I'd do it manually if it were only a couple of edits, but there are easily 150+ comments spread through 6 hrs of material. Should be easy for a speech recognizer, just two voices with different genders.
Alternatively, you might look to see if the speakers are in separate channels (if recording is stereo). Then it would be really simple - just take one channel out and resave as mono!
If you have a sample I'd be curious to take a look - sounds like an interesting problem.
- Divide all examples (observation sequences) uniformly into as many segments as the number of states.
- Cluster the observations corresponding to each state, and estimate the GMM using the cluster set so that each cluster corresponds to one multivariate Gaussian.
- Do this repeatedly until convergence: get the Viterbi alignment of all examples, use it to get new segments, and estimate the parameters again using the previous step.
Please correct me if I'm wrong. Also, I have two questions:
- What kind of accuracy increase should be expected if using both Viterbi training and Baum-Welch re-estimation, instead of just the latter?
- What kind of accuracy should be expected if only using Viterbi training?