I've done a fair bit of it in the past - the issues with missing spaces or substitute characters used for spaces certainly rings a bell. I know quite a few libraries in the .NET space for processing PDFs, including at least one open-source one (PDFSharp from memory) which we had in our source code repo, and made up something like 75% of the total LOC for our product, despite PDF processing being a relatively minor feature). We ended up abandoning that though, can't quite remember why.
The real question is how did such a torturous (once proprietary) format manage to become the only ubiquitous standard for paginated, self-contained, formatted text suitable for printing. In principle HTML could actually do the job, but arguably it's not the ideal tool either, and I've not seen an example of an HTML document that just prints as expected across multiple pages either.
Odd that this got posted almost the same time as https://news.ycombinator.com/item?id=33145498 btw...