From a machine-parsing perspective PDF files are a nightmare. Chunks of text may be broken anywhere, mid-sentence, mid-word. These chunks may appear in the document in any order.
Spaces may be encoded as spaces, or they may be created a number of other ways, like by positioning chunks, or setting character spacing per character.
The mapping from code point to glyph does not need to be pure Unicode, a PDF document may contain a custom font with additional glyphs.
This is all stuff I learned by trying to parse a limited set of PDFs found in the wild.
All of these gotchas are by the way completely PDF/A compliant.