Audiogrep transcribes audio files and then creates “audio supercuts”
antiboredom.github.io
antiboredom.github.io
Existing trained models can be downloaded from here: http://www.kaldi-asr.org/ (via: http://www.openslr.org/12/ http://www.clsp.jhu.edu/~guoguo/papers/icassp2015_librispeec...)
I second the opinion that Kaldi is more advanced, but it is also way, way more complicated to do anything with a custom dataset. There are a few examples of decoding with existing models though, so maybe that is a start. These lectures may help: http://www.danielpovey.com/kaldi-lectures.html
There is also an interesting toolbox here that you can train, though getting access to TIMIT, WSJ, etc. is pretty annoying. http://www.cs.cmu.edu/~ymiao/kaldipdnn.html
You may also get mileage out of some kind of post-transcription NLP/cleaning if you haven't done that yet.
And this related project: https://github.com/antiboredom/videogrep
He already states the obvious idea to integrate audiogrep into videogrep, which at the moment just uses subtitle files.
> All the instances of the phrase "time" in the movie "In Time": https://www.youtube.com/watch?v=PQMzOUeprlk
> All the one to two second silences in "Total Recall": https://www.youtube.com/watch?v=qEtEbXVbYJQ
> The President's former press secretary telling us what he can tell us: https://www.youtube.com/watch?v=D7pymdCU5NQ
I assume this will be useful to data scientists who want to process lyrics? what other intended/near-at-hand use cases are there?
Actors hate doing ADR and it's time-consuming and annoying for editors. This wouldn't automatically solve the problem because you wouldn't have a good match between dialog recorded in different acoustic environments, but it does have the potential to save a lot of grunt work, especially for background dialog where you can compromise on quality a bit.
Also, in post-production you often find yourself wanting to edit just one or two words in a scene and you'd rather not bring the actor back for such a small problem, so you look for other scenes and other takes of the same scene where the same word or syllable appears, and do a little cut-and-paste and blending, the audio equivalent of photoshop retouching. It would be very useful for that.