Why not looking for stretches characters with spaces between them, then concatenate, check against a dictionary and if a match is found, remove the spaces.
> “On_April_7,_2013,_the_competent_authorities”
Same here.
Why not looking for stretches characters with spaces between them, then concatenate, check against a dictionary and if a match is found, remove the spaces.
> “On_April_7,_2013,_the_competent_authorities”
Same here.
If it's supposed to be a somewhat final result, ran against a dataset you have little control over (people sending you PDFs they made, vs. PDFs coming from a know automated generator) you'll hit all the other not so edge cases very fast.
Like, otherwise invisible characters inserted in the middle of your text, layout that makes no logical sense and puts the text in a weird order but was fine when it was displayed on the page, characters missing because of weird typographic optimization (ligatures, characters only in specific embeded fonts etc.). Basically everything in the article is pretty easy to find in the wild.
Where dictionary lookup really becomes infeasible is in languages with a great deal of inflectional morphology, like Finnish. For languages like that, you need some kind of morphological analyzer. Hunspell provides a fairly simple one, while a number of finite state transducers (xfst, foma, sfst...) provide more sophisticated mechanisms to build a morphological parser with.