How I OCR hundreds of hours of video (2011)
waldo.jaquith.org
waldo.jaquith.org
On a more national scale, could something like this be done for Congress? C-SPAN already does most of the hard work of filming and uploading to the web, so perhaps it won't be too difficult. I think it would certainly attract a lot of interest... maybe I'll give it a go.
Feel free to check it out at https://archive.org/details/tv and search around for some fun terms/words.
If you really like it and if you like the Internet Archive, feel free to donate a one-time sum or set up a subscription at http://archive.org/donate/ - they're a US-based 501(c)(3) non-profit organisation - so donations are tax deductable if you're US based.
While a spellchecker might fix Jenn1fer -> Jennifer, at the OCR stage there is much more information to do it properly; but it obviously doesn't know that McClellan is valid word and thus a much more likely alternative than i\1cCie1ian, and it needs to be told that. The list of speakers on those videos is limited, and their surnames can be added to the appropriate dictionaries to improve their recognition.
Google cache of the site if it's unavailable (I'm getting a database error).
Handbrake has a CLI in addition to the GUI, ffmpeg is another good CLI option.