"Source --> PDF --> docx --> HTML --> json" is just nuts.
"Source --> PDF --> docx --> HTML --> json" is just nuts.
It's likely that the original source for the PDF/paper copies was lost. This necessitated the process of using OCR to get it into Word.
Obviously maintaining such a large and important piece of documentation in Word is wrong, but that's apparently how it's being done.
This is an important statement. in that it is a reasonable goal, but nearly impossible.
Even using word, once one gets past a simple document, and starts having TOC, index, etc. and is "four five-inch binders" full of pages, then not just ANY CLERK will be able to maintain the document.
It has been my experience that while almost everybody can create a passable document in Word, something technical/government-ish with 500 pages and all of the publishing accouterments that come with a document of size, purpose, and complexity, you're getting into specialist territory, and almost all of the DTP specialists I know start by importing Word files in to InDesign.
This reads like someone who doesn't have an understanding of the enormous power that word has hiding under it's shell.