They manage to do it (just sound, not keypresses) with a 60fps DSLR camera at the end by examining the rows of the video - does this mean that for any videos already in existence, sound can be decoded from the images?
Its not a complicated paper, give it a go.