Tesseract OCR
github.com
github.com
(1): https://towardsdatascience.com/pre-processing-in-ocr-fc231c6... (2): https://github.com/chriswolfvision/local_adaptive_binarizati...
I once wrote a bookscanner app in Java (https://boofcv.org), where everything was done automatically (preprocessing, object detection / book extraction, skin detection / finger removal, deskewing, line-slope-correction and so on). It was very difficult to adjust the parameters, that at least most of the books looked good.
Otherwise, the C++ code on Github requires converting images to PGM format.
---
The page is in French, so I will mention that the Python script is here: https://www.vvpix.com/gmp_Telecharger_script.php?sFichier_a_...
To install it, copy the file to: `C:\Program Files\GIMP 2\lib\gimp\2.0\plug-ins`
Then call the script via the `Python-fu > Color > Binarize` tab in Gimp.
The algorithm is quite slow for large images. Aim for 1440p at most.
That being said, the results on my quick experiments look great, so it saves my time compared to other more manual methods in the end!
Sadly I can't find any open-source vectorizer & OCR for repair scanned technical drawings, and Tesseract has a lot of issues with rotated text labels specific to CAD.[1]
I learned about it from this post: https://alexn.org/blog/2020/11/11/organize-index-screenshots...
Firefox ScreenshotGo[beta] https://mzl.la/2NMgD30
Or, just upload your photographs to Google Photos. Google OCRs all images automatically(!) and you can search them for text in the images. This includes text e. g. on posters in the background.
Usefully, the new macOS / iOS releases will do this automatically (although for macOS you'll need to be running Apple Silicon)
I read HN on my kindle[1] to assimilate knowledge from the comments using its highlighting, clipping features.
But commenting is painful on Kindle's 'forever experimental browser' so I take a screenshot and when I connect it to the computer the to-comment stack on a network to-do list is updated with a ready to visit HN story URL using query from the text on the screenshot parsed using Tesseract.
[1] https://hntokindle.com/ (Disclaimer: I built this)
OwlOCR is the one I use, with shortcuts set up next to the cmd-shift-3/4/5 screenshot keys.
TextSniper is another one, and I think there was a third that I can't remember the name of.
I have used Tesseract for OCRing scanned books and it was great. I had no idea it was so old, nor that it had been through so many maintainers. To all of them past and present, thank you.
One point of friction has been selecting an OCR workflow. Any chance you would share what you’ve been successful with?
Most of time was spent in field parsing and validating ocr output (is it valid date). At one point I realized that playing with tess config was giving marginal improvement, and investment in post-ocr parsing/wrangling was more valuable e.g. in date column, if ocr says b, consider it 6 and flag low confidence record.
One new nice-to-have use case customer asked was varying orientation of pages, that I couldn't hack together quickly.
I use https://github.com/4lex4/scantailor-advanced to deskew the images and generate the PDF.
It isn't perfect but my purposes are more around research than publication, so, YMMV!
You pretty much need black text on white background at 300-600 dpi. (Not sure the exact size but I’ve had crappy scans do better by scaling the file.)
I’ve had reasonable success with photos of printed pages run through text cleaner.
Doesn't OCRopus qualify as well (it does look unmaintaned, or less actively maintained than Tesseract)?
I hacked something together with "Capture2Text". Basically taking the screenshot, saving to jpg, shelling to the exe and getting the text back. Works pretty good.
One approach would be to say language doesn't matter, just train on converting any character from any language alphabet from image to text. The problem is that higher accuracy can be achieved by isolating characters from each language from each other. I imagine that particularly for Latin alphabet languages, accuracy must improve dramatically by splitting out any kanji or hanzi.
I'd recommend https://scantailor.org/ for this (OSS, but unmaintained)
Scan Tailor forum: https://forum.diybookscanner.org/viewforum.php?f=21
ScanTailor official repo also archived on November 29, 2020.[0]
> This project is no longer maintained, and has not been maintained for a while.
As alternative to ScanTailor (and its forks) there is gImageReader (Tesseract Qt/GTK GUI), but also seems like unmaintained since 2019.[1,2]
[0] https://github.com/scantailor/scantailor/commit/e881b30b6ed1...
[1] https://github.com/manisandro/gImageReader
[2] https://github.com/probonopd/gImageReader/releases/tag/conti...
Agree it’s best to skip tesseract unless the free cost is important. We spent a lot of time trying to preprocess and tune tesseract before realizing cloud OCR solutions are much better and fairly cheap.
I haven't seen it make any mistakes at all and responses take less than 3 seconds usually.
If you're looking to add a text layer to a PDF (for search purposes for instance) I can highly recommend OCRmyPDF: https://github.com/jbarlow83/OCRmyPDF/
It uses Tesseract and works quite well for most PDFs, I made a semi-functional script before I discovered it and it would have saved a lot of hassle.
Overall I thought it was great and I wonder how good it would perform these days with 10 years of improvements!
https://kebekus.gitlab.io/scantools/
Make SURE to select the correct OCR language
https://www.forbbodiesonly.com/moparforum/threads/fender-tag...
I'd like to capture each item that is delimited by whitespace, convert to text, and store its position and line in a database.
The same code may appear more than once with different meaning, so position is important.
The tags are often different colors as well.
Anyone know which technology may be best or simplest to implement?
This is for a historical search function.
Any open-source solution you'd recommend for handwriting recognition?
I'm wondering how actively it is being developed. I see that the last release is from 2019. I also see, however, that there have been some version 5.x alpha releases published this year. Does anyone know what is happening inside the project?
"Tesseract 4 adds a new neural net (LSTM) based OCR engine which is focused on line recognition, but also still supports the legacy Tesseract OCR engine of Tesseract 3 which works by recognizing character patterns."
I believe Apple are adding this feature to the OS soon.