I just started on a project with about 5000 pages worth of government supplied documentation in PDF form. I wish I could just throw your tool at it.
The hard part would be parsing tables and other layout-dependent semantics. You usually start with text coordinates (like HTML elements with absolute position) and have to work backwards from that. I worked for some years in a project for a client that was full of edge cases, because whenever the input PDF (from a government agency) would have a slight layout change the parser would break. It took multiple iterations to make it robust enough.
Would love to chat with you if you're up for it - you can test demo run our tool and contact us through the interface