> Why not OCR all the time? > Running OCR on a PDF scan usually takes at least an order of magnitude longer than extracting the text directly from the PDF.
... so? Google can OCR video and translate it in something that feels like real-time; what PDF processing are they doing that is so performance bound?
> Difficulties with non-standard characters and glyphs OCR algorithms have a hard time dealing with novel characters, such as smiley faces, stars/circles/squares (used in bullet point lists), superscripts, complex mathematical symbols etc.
Sure, but more than the random shit you find in PDFs anyway?
> Extracting text from images offers no such hints
Finding an algorithm that approximates how a human approaches a page layout doesn’t feel like it would be all that hard.
Obviously it’s very easy to stand on the sidelines and throw stones, but parsing PDFs using anything other than OCR + some machine learning models to work out what the type of a piece of text feels like pretending we are still constrained by the processing costs of 5 years ago