PaddleOCR: Lightweight, 80 Langauge OCR
huggingface.co
huggingface.co
[0]: https://arxiv.org/abs/2109.03144 [1]: https://github.com/PaddlePaddle/PaddleOCR [2]: https://gradio.app/ [3]: https://huggingface.co/spaces
I entered a simple text "John kissed Mary in 2010. Two years later Adam was born." and a question "When was Adam born?" The answer was. . . "Two years."
This is one of the reasons I hate the current AI hype. People would do much better being honest and realistic about our progress in the field - it's enormous, but not near what the marketing teams claim it to be.
SQuAD is the prototypical example of a dataset for this task, see https://rajpurkar.github.io/SQuAD-explorer/
Now, this is an artificial divide: there's nothing in the defintion of what OCR is that means "oh it only works on scanned documents". It's just what it used to mean back in the day.
Still, if that's your kind of problem, you may look into something more along the lines of a deep convnet to do character detection and then recognize each bounding box into an alphabet.
But I see that (like tesseract) it cannot recognize different styles (italics, bold, monospace...) - not only it seems to just translate into pure (UTF) characters, it shows confusion on terms in alternative styling.
- for instance, take the cropped bitmap of the term and compare it with the renderings of the styled variants of the recognized word and see what is more similar, and/or, check the thicknesses, spacing, slanting in the bitmap of the word...
I think it would be intriguing to develop it.
Better techniques are required than the raw one I used: for example, finding the best overlapping of the two bitmaps, maybe with some sort of gradient descent over a few pixels distance in panning and scaling - this should give a near to 100% correspondence in the correct case (regular vs italics vs bold vs monospaced vs BI, BM, IM, BIM), but only if the font used for the comparison is the same.
By the way: in the fact that while adjusting the overlapping the computed difference should increase with the gradient when the style corresponds (R on R), but may be random in other cases (R on I, R on M, though not R on B), there could be another key in the heuristic.
This packages the tesseract very nicely: https://kebekus.gitlab.io/scantools/
I actually would buy Abbyy OCR but pricing for Linux ist just insane for private use. I just saw, the CLI is even discontinued: https://www.ocr4linux.com/
What may drive this decisions?
I mean, you provided the link:
"Why we are sunsetting FineReader Online
The entire ABBYY FineReader product family is getting a new look and feel. Our online recognition tools will be reworked and introduced at a later date to demonstrate the power of ABBYY's OCR technologies."
Tried to OCR some russian text from an image and got absolute nonsense.
The new version was botched last I tried.
Edit: I finally got it to work. The result looks good! https://i.imgur.com/hoS4oMP.png
Though it looks like yet another OCR program that doesn't understand archaic lexical paradigms like the long S or ligatures.
To me good results is like 99%+ correct, and the ability to highlight where it’s confused.
This kind of blobby faded printing is still challenging for OCR. The fact that it decided to just skip entire sections is the most troubling part for me (like seriously wtf). But the parts it didn't skip I think are quite good compared to when I use other software on the same kind of material.
I wish these things had a bit more...sanity...for lack of a better word. t769 is just ridiculous. TEcole isn't a word. Beaucoupde is clearly two words that shouldn't be smushed together. etc.
Interestingly, Apple quietly bakes high quality free OCR into macOS as a library that developers can invoke in their own software that works better than this in some ways and they just don't advertise it to end users or do anything with it at all themselves (they do on iOS but Preview.app could have had an OCR option since Catalina). It also doesn't recognize things like long S, though, so it's still annoying for old texts.
Tesseract versions 4+ (when they started supporting LTSM) is pretty easy to train on obscure fonts and rare languages.
The downside is that you need labeling...a lot of labeling.
We used about 10,000 lines of humanly crowdsourced results to supplement an existing model of about 50000 lines.
Accuracy is over 99% which is sufficient for most uses.
Those were the only two relative deficiencies I noticed.
It does seem to beat tesseract on samples with mixed dark-on-light and light-on-dark text, but that was the only big win I saw in my brief look at it.