Do you have some concrete examples of corrupted texts? I'd very much like them to use as test cases in a document processing pipeline.
Do you have some concrete examples of corrupted texts? I'd very much like them to use as test cases in a document processing pipeline.
DjVU: https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
PDF: https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
I don't have an easy way to find all of the errors I observed but typical ones are on page 58 at the end of the third line, 'Germanic' with a little bit of bleed between the serifs at the bottom of 'n' gets transformed into 'Germauic', and on page 420 at the end of line 7, 'compass' with a small dot in the middle of the 'o' becomes 'cempass'.
FWIW I've transcribed a bunch of other DjVu files produced by the Internet Archive and not noticed anything like this.
If I can reliably flag problem cases then at least I will know which files are going to have a lot of manual work.
How would you know? Maybe if a document had an unusual misspelling. But how would you know if DjVu munged a 3 into an 8?
It seems also that the problem happens more often with cyrillic glyphs which have less descending/accending elements that typical latin ones.
I totally understand how this could theoretically happen but it all hinges on it actually happening in the documents that I'm working with and to have a test case where it happens for sure would make all the difference in determining whether or not the documents I'm working with are at risk or not and if so if I can detect which ones are at risk (that would be half the battle won).
It would have to be a case where a human would see a 3 or an 8 (or a similar transposition with other glyphs) resulting in a document that is corrupted in such a way that afterwards the human would see the (invalid) alternative.
Even a single instance of this happening would be very relevant.
To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. The compression algorithm could not only mistake '3' as '8' with small gaps left, it would replace this dirty '3' with the image of reference '8' image, so that human reader think that he sees a scan, i.e. an image of page scanned, while in fact it sees 'edited' image, not corresponding to actual page.
I take your word, there is a technical reason behind this, not that I don't believe you.
> To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity.
Yes, I totally get that. It's about DjVu's compressor replacing the image of one character with the image of another.
You would obtain more credibility by actually providing such an example.
By comparing it with the original, surely.