Using Textract from Amazon to do the OCR. I dump the raw json in a SQLite database.
Information extraction is trickier. I extract useful things like lines and pages, locations of lines, set up various columns and now can search the SQLite db either with python code or SQL queries.
There's of course cost for running textract OCR on AWS. There are open source solutions like Tesseract
Lots of missing context from these sheets that has to be interpreted (ie, how do you taxonomize each field of information?). Then asking questions on top of these documents is a step on top: "is the allegation about sexual violence?", "What is the name and rank of the person being accused?", "Is anything anomalous in the review process?", "Has this person's rank changed in the past 5 years?" etc etc.
Now expand this problem to hundreds of thousands of different types of document.
Meaning, the information that formed the PDFs very likely come from a relational database and an inversion back to its original relational form is probably the convenient form. Whether that means it turns to 40 tables, that's fine, so long as it's relational and a query can be written.
This right there is the difficult part - what do you mean exactly? I cannot come up with anything better than search, as in like Google search. And they did it for books already, it's seriously good.
The public good of having a resource like this available to the public for free is beyond unimaginable as far as I'm concerned.
Mind you, I'm not exactly looking for advice here. It's a supremely difficult problem and gut-ideas more often than not don't pan out.