Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now.
Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now.
Tesseract is more like getting a pretty good motor for free (recognizing text), but it's up to you to build the rest of the car around it (preprocessing images, handling errors, dealing with the output, potentially training it to your task, and various other issues).
Yes! Can anyone comment on why this is the case, since OCR is proclaimed to be a solved problem?
I've always wondered why Google Lens works "out of the box" and shows great accuracy on extracting text from images taken using a phone camera, but open-source OCR software (Tesseract, Ocropy etc.) needs a lot of tweaking to extract text from standard documents with standard fonts, even after heavily pre-processing the images.
PS: Has Google released any paper on Google Lens?
We were basically trying to transcribe message screenshots which should have been relatively straightforward given the homogeneity of the font. But this was not the case as tesseract was not trained in the layout of msg screenshots. The accuracy of raw tesseract on our test dataset was somehwere about 0.5-0.6 BLEU.
Once we were able to isolate individual parts of the image and feed it to tesseract, we were able to get around 0.9 BLEU on the same dataset.
TLDR;Some nifty image processing is required to make tesseract perform as expected.
[0] (https://www.askgoose.com) [1] (https://github.com/tesseract-ocr/tesseract)
I'm surprised, too. After all, if you can train an AI to recognize a cat, why can't it be trained to recognize a letter?
Mine, for example, works well on clean laser-printed text. It fails on anything written with a typewriter, though. (My definition of "failure" is it's quicker to retype it from scratch than fix the OCR's errors.)
I'd also love to have one that worked on cursive handwriting.