This looks nice. What I'd really like to see, along these lines, is a python library for automated document metadata extraction with confidence assessment, like this:
./autometa.py --author --verbose academic-paper.pdf
Author: "Edward Witten" Confidence: High (matches template "amslatex")