As an experiment, we once tried converting an OCRed dictionary (this one: https://www.sil.org/resources/archives/10969) into an XML dictionary database. (There are probably better ways to get an XML version of that particular dictionary, but as I say, this was an experiment.)
Despite the fact that it's a clean PDF, and uses a Latin script whose characters are quite similar to Spanish (and the glosses are in Spanish), the OCR was a major cause of problems: Treating the upside down exclamation as an 'i', failing to separate kerned characters, confusion between '1' and 'l', misinterpreting accented characters, and so on and so on. And for some reason the OCR was completely unable to distinguish bold from normal text, even though a human could do so standing several feet away.
So I did think of extracting the characters from the PDF. If it had been a real use case, instead of an experiment, I might have done so.
Write-up here: https://www.aclweb.org/anthology/W17-0112/