Does anyone know of any attempt at this?
Blah blah blah transformer something BERT handwave handwave. I should ask the research folks. :-)
Does anyone know of any attempt at this?
Blah blah blah transformer something BERT handwave handwave. I should ask the research folks. :-)
As an experiment, we once tried converting an OCRed dictionary (this one: https://www.sil.org/resources/archives/10969) into an XML dictionary database. (There are probably better ways to get an XML version of that particular dictionary, but as I say, this was an experiment.)
Despite the fact that it's a clean PDF, and uses a Latin script whose characters are quite similar to Spanish (and the glosses are in Spanish), the OCR was a major cause of problems: Treating the upside down exclamation as an 'i', failing to separate kerned characters, confusion between '1' and 'l', misinterpreting accented characters, and so on and so on. And for some reason the OCR was completely unable to distinguish bold from normal text, even though a human could do so standing several feet away.
So I did think of extracting the characters from the PDF. If it had been a real use case, instead of an experiment, I might have done so.
Write-up here: https://www.aclweb.org/anthology/W17-0112/
Many menus are also available in PDF form, so we're trying to figure out if it's worth bothering with the PDF itself, or if we should just render to image and thus reduce the problem to the menu-photo one.
And on top of this many times there are no text data in the PDF just a JBIG2 image per page.