Cool project (legal/ethical issues aside)! My big question is why bother with OCR at all? Presumably you can construct a representation of the screen for each question ahead of time, and you already solve the projective transform due to the camera, so at that point why not pick some appropriate distance metric and pick the question/answer pair corresponding to the virtual image that is closest to the camera image?