Paircast automatically transcribe the video's audio to text.
I think that an audio (or video) version would be superior, but more because, assuming there are captions, the audio might be able to describe something beyond the code alone more easily than typing out a note would.
You don't just take a transcript of the dialogue in the movie to make a book.