[0] https://github.com/tesseract-ocr/tesseract [1] https://github.com/tesseract-ocr/langdata
- OCR for French is very good
- OCR for Vietnamese (latin characters) was dispointing
- OCR for Chinese indeed outout Chinese characters but not the good ones
- OCR for Chinese didn't handle letters in Chinese script
- no support for Vietnamese Nôm (not suprising)
The good news is: - you can train new language model
- the software is built on a (clean?) API you can use yourself
- there is some usable documentaition to get started & also for training models
- headers files are well commented
So this enable one to build a solution on Tesseract, but this is by no way an out-of-the-box solution especially if you have very specific needs. All in one it is the best free (libre, 0$) OCR engine out there.
[1] http://crim.fr/sites/default/files/memoire_m2_Lecailliez_fin... [French]
There is Cuneiform, a former main competitor to ABBYY Finereader. CuneiForm got open sourced a view years ago, though in a sad state (project files where in VS C++ 6 ('98), comments in Russian), but a community fixed that and ported it to Linux. It's also probably the best one for Russian language. It also has an UI and some advanced features that only ABBYY amd Cuneiform have, but non of the other competitors (certainly no other open spurce OCR package). https://en.wikipedia.org/wiki/CuneiForm_(software)
If you need to do OCR that also preserves table structure -- which is what I bought FineReader for in the first place, I don't think there's any open source alternative, and FineReader does a very capable job.
Here's an example of FineReader in action: OCRing the docs released by the FBI on Clinton's email system. I've also included the pdftotext output showing how FineReader's text conversion also attempts to preserve the physical layout of the text characters:
https://github.com/dannguyen/clinton-hillary-email-fbi-inves...
I used to use Evernote quite heavily and I've never been able to replace their excellent OCR. I would search for some text and was always blown away when it would find a photo of a whiteboard or a sketch of mine.
Any idea what Evernote uses?
Everything about it sucks, cost, anual page limit license, I can only run it as root for some reason, but the OCR quality is superb.
here's the scripts I use if it's useful: https://github.com/cove/scanbd
https://gist.github.com/dannguyen/a0b69c84ebc00c54c94d
Unfortunately, it doesn't do well on large chunks of text if you're trying to OCR a scanned document.
What I've tried so far (including Tesseract) is either bad for Russian texts or cannot work with mixed texts (e.g.Russian with some English words). Or both.
Programming languages/platform don't matter, but smth Linux-compatible is better of course.
OCR isn't limited to language usually unless you are doing some really high end stuff when it does linguistic prediction but you only need that if you are working with really poor (image) quality sources.
But overall OCR is "language" agnostic, it is however usually not type set agnostic so what you would want to do is train it for whatever fonts are common for a particular language.
This gets slightly tricky if you have to do handwritten transcription or very stylized fonts but in those cases the "language" again is not an issue because your OCR program doesn't understand language to begin with.
https://christopher5106.github.io/computer/vision/2015/09/14...
A lot of information about current Tesseract is there - https://github.com/tesseract-ocr/docs/tree/master/das_tutori... .
Tesseract is trainable, even though the bulk of capabilities came from algorithms designed well before deep learning became popular.
One of slides mentions that it's puzzling that Tesseract is "winning" over modern ML attempts to solve OCR. However, latest developments - adding LSTM networks to Tesseract - are reported to be promising. Wonder when they'll become available on the Github...