Wouldn't it be easier and more generic to have an OCR solution for this task?
If you have any pointers in the OCR route, do suggest them here, or on this GitHub issue! https://github.com/socialcopsdev/camelot/issues/101
I'd bet that commercial OCR packages that are long in the game have unified code for these functions between regular OCR and PDF processing.