EasyOCR: Ready-to-use OCR with 40 languages
github.com
github.com
Notice the faded text from the printer running out of ink and the slanted text. From limited experience each of these are thorny problems and the state of the art CV algorithms won't help you escape from having to learn how to algorithmicly pre-process images and clean them up prior to feeding them into a CV algorithm. You might be able to use Google's Cloud OCR but that charges per image, although it is pretty good. Even if you use that you've graduated to the next super difficult problem which is Natural Language Processing.
Once you have the text you need to determine if it has meaning to your application. That's basically what NLP is about. For the receipts example, how do you know you're looking at a receipt? What if its a receipt on top of a pile of other receipts? How do you extract transactions from the receipt? Does a transaction span multiple lines? How can you tell? etc etc etc.
Honestly I was kind of surprised that good basic OCR isn't a totally solved issue with an ecosystem of fully open-source solutions by now.
Yes! Can anyone comment on why this is the case, since OCR is proclaimed to be a solved problem?
I've always wondered why Google Lens works "out of the box" and shows great accuracy on extracting text from images taken using a phone camera, but open-source OCR software (Tesseract, Ocropy etc.) needs a lot of tweaking to extract text from standard documents with standard fonts, even after heavily pre-processing the images.
PS: Has Google released any paper on Google Lens?
We were basically trying to transcribe message screenshots which should have been relatively straightforward given the homogeneity of the font. But this was not the case as tesseract was not trained in the layout of msg screenshots. The accuracy of raw tesseract on our test dataset was somehwere about 0.5-0.6 BLEU.
Once we were able to isolate individual parts of the image and feed it to tesseract, we were able to get around 0.9 BLEU on the same dataset.
TLDR;Some nifty image processing is required to make tesseract perform as expected.
[0] (https://www.askgoose.com) [1] (https://github.com/tesseract-ocr/tesseract)
Tesseract is more like getting a pretty good motor for free (recognizing text), but it's up to you to build the rest of the car around it (preprocessing images, handling errors, dealing with the output, potentially training it to your task, and various other issues).
I'm surprised, too. After all, if you can train an AI to recognize a cat, why can't it be trained to recognize a letter?
Mine, for example, works well on clean laser-printed text. It fails on anything written with a typewriter, though. (My definition of "failure" is it's quicker to retype it from scratch than fix the OCR's errors.)
I'd also love to have one that worked on cursive handwriting.
I've just tried easyocr on a receipt, and it's pretty bad. I've also just noticed that ASDA have a "mojibake" problem and print ú instead of £ on the entire receipt ...
What's actually happening seems to be that the ch_tra model can recognize simplified too and output the corresponding traditional version if the character isn't in the traditional "alphabet"; it doesn't work so well in the other direction.
Example recognizing a partial screenshot of https://chinese.stackexchange.com/a/38707 (anyone can try this on Google Colab, no hardware required; remember to turn on GPU in Runtime -> Change runtime type):
import easyocr
import requests
zhs_reader = easyocr.Reader(['en', 'ch_sim'])
zht_reader = easyocr.Reader(['en', 'ch_tra'])
image = requests.get('https://i.imgur.com/HtrpZCZ.png').content
print('ch_sim:', ' '.join(text for _, text, _ in zhs_reader.readtext(image)))
print('ch_tra:', ' '.join(text for _, text, _ in zht_reader.readtext(image)))
Results: ch_sim: One simplified character may mapping to multiple traditional ones: 皇后->皇后,後夭->后夭 豌鬟->头发,骏财->发财 As reversed, one traditional character may mapping to multiple simplified ones too: 乾燥->干燥, 乾隆->乾隆 嘹望->嘹望,嘹解->了解
ch_tra: One simplified character may mapping to multiple traditional ones: 皇后->皇后,後天->后天 頭髮->頭發,發財->發財 As reversed, one traditional character may mapping to multiple simplified ones too: 乾燥->干燥, 乾隆->乾隆 瞭望->瞭望, 瞭解->了解
Compare to the original text: One simplified character may mapping to multiple traditional ones:
- 皇后 -> 皇后,後天 -> 后天
- 頭髮 -> 头发,發財 -> 发财
As reversed, one traditional character may mapping to multiple simplified ones too:
- 乾燥 -> 干燥,乾隆 -> 乾隆
- 瞭望 -> 瞭望,瞭解 -> 了解
Of course, automatic character-to-character conversion from simplified to traditional can be wrong due to ambiguities; excellent examples from above: 头发 => 頭發 (should be 頭髮), 了解 => 了解 (should be 瞭解).It doesn't seem to have a dictionary to do word level matching, only character level.
Yes, the simplified model is not that great at recognizing simplified either, at least in this case.
E.g. poor-quality images of ID cards or credit cards, where the position of data is known.
https://developer.apple.com/documentation/vision/recognizing...
In my job as a support engineer I sometimes get screenshots of complex technical configurations and end up having to type them in one character at a time, so this would be really handy.
Looks like maybe I could just create a wrapper around EasyOCR.
http://askubuntu.com/a/280713/81372
I put it in a custom keyboard shortcut, so I just press it, draw an onscreen rectangle around any non-selectable text, and in a few seconds it goes to the clipboard.
It means that it performs not-so-good when for example image contains black text and white text on green background since this is not "normalized" through image preparation steps and it cannot detect white text on green background (but you can do it yourself)
This does involve creating your own labeled data.
I got this 99% accuracy by performing incremental training using latest Manheim model as a base. I added about 20k lines which is not really that much. https://github.com/tesseract-ocr/tesseract/wiki
The hard part was crowd sourcing those 20k lines :)
Tesseract might not be best for photos as you said but I did not have major problems.
Of course some documents the source is so bad that a human can't achieve 99%.
Tesseract used to be quite average before they moved onto LTSM models a few years ago.
Also this: https://github.com/UB-Mannheim/tesseract/wiki
The original data was here: https://github.com/tesseract-ocr/langdata_lstm
I did use another data source from Manheim but can't locate it right now.
Using vanilla Ubuntu 18.04
I looked at the example training files and made a small script to convert my own labeled data to fit the format that tesseract requires.
I did do a bit of pre-processing adjusting contrast.
All the data munging was done on Python (Pillow for image processing, Flask for collecting data into a simple SQLite DB before converting back to format that Tesseract requires).
Python was not necessary just something that felt most comfortable to me. I am sure someone could do it using bash scripts or node.js or anything else.
EDIT: To make life easier for my curators I did run Tesseract first to generate prelabeled data for my training set. It was about 90% accurate to start with.
So the process was: Tesseract OCR on some documents to be trained -> hand curation (2 months)-> train (took about 12 hours) -> 99% (on completely separate test set)
It works for sparse text on images, and for that specific use case it is better than Tesseract.
a bit out of topic - but does anyone happen to know if there is an open-source, new school OCR library for music notation?
Cost: AWS is not free vs Open Sourced
Time: AWS averages under 10 seconds vs 140 seconds on a standard Dell 7480 & 9 seconds on a GPU Google colab
Character Accuracy: Almost same on a high quality input. No comparison with AWS on a blurred camera photo like this https://github.com/ExtractTable/ExtractTable-py/blob/master/...
However, it looks like my simple example of an old "S note" export (like a lowish resolution phone screenshot) confused it a bit:
Reglementation -> Reglemantation
km -> kn
illimitée -> illiritée
limite -> liite
baptême -> bapteme
etc.
Overall, it works, and it is quite easy to install and use. I'd have to compare it with tesseract, but I think it's a bit better. A lot slower, though (I only have AMD devices, no CUDA). It's underusing my CPU, and maybe leaking memory a bit, though I didn't clean up.Take that with a grain of salt, that was a quick try, I haven't tried to tune anything.