Tesseract.js: Pure JavaScript OCR for 100 Languages
tesseract.projectnaptha.com
tesseract.projectnaptha.com
As far as I know, it powers all OCR at Google (e.g. in Keep, Docs, etc.).
This (Tesseract.js) is a WASM port of the project by a separate group of people.
I investigated using this port a couple years ago, but as you can see from the demo, it's fairly slow to initialize and run, so I never found a practical use for OCR client-side rather than server-side, but I still think it's tremendously cool.
In case anyone's interested (shameless plug), because I do a lot of academic research that involves tons of copying from webpages, PDF's and screenshots and pasting into notes documents, I created a tool at https://pastemagic.com that helps selectively remove rich text formatting, remove line breaks and does OCR on screenshots and camera photos. Setting up Tesseract on my server and creating a simple HTTP endpoint for it took less than an hour, and for free I had OCR as powerful as Google's. Pretty cool I thought.
Afaik Google no longer uses Tesseract for any of its products. Googles Clould OCR is much better than Tesseract.
I think Google devs still work on Tesseract, but only as their side project (not sure about this, obviously)
Tesseract 4.0 has a brand-new neural engine that totally supersedes the earlier engine, however -- I wonder if there's any relation between that and Cloud OCR?
Fun fact: it actually started as a US national defence initiative in the last big AI hype bubble in the 1980s.
While it isn't in the wiki, probably for intellectual property reasons, I'm almost positive Cuneiform originated as the Soviet version of the same thing (I have evidence in notebooks somewhere): OCR was something needed for "AI" of the 1980s in the Soviet system as well.
Had it been further developed by Microsoft ... or Apple, in contrast to Tesseract, it would have continued to be the official OCR of an opposing world-historical system.
The only semi-tricky bits were parsing stderr if anything went wrong (distinguish warnings from actual errors), and the fact that Tesseract doesn't respect the JPEG orientation bit (big problem with iPhone camera images), so checking that and manually rotating the JPEG first if necessary (gibberish otherwise).
The goal is to draw a box using GUI, then use those coordinates to extract text from several homogeneous pages.
I also have a different goal of trying to interpret structure of a PDF that has visual structure (headers, sections and subsections all numbered). But that seems to lend itself to some sort of text parsing.
Some reading here: https://stackoverflow.com/questions/53219016/detecting-secti...
Here's how to extract text from a PDF based on coordinates (this explains how to do it on web, but it's also possible using other platforms):
https://groups.google.com/d/msg/pdfnet-webviewer/h2W3VksbQUI...
Here's how to extract a PDF's logical structure:
https://www.pdftron.com/documentation/samples/#logicalstruct...
https://github.com/tesseract-ocr/tesseract/wiki/Command-Line...
It's not quite what you want, but I think you could probably filter the output based on the selected region and pretty quickly get what you want.
Not sure if it only reads at those coordinates vs. OCRing the whole thing (for example if you were legally prohibited from OCRing content outside a certain coordinate space), but it is selectable.
There's a very popular and minimalist CLI called scrot that I think would be ideal... well scratch that, I made a search and our question has already been asked and answered:
https://askubuntu.com/questions/280475/how-can-instantaneous...
https://stackoverflow.com/questions/21497447/ocr-on-a-screen...
It is opensource and runs on Java.You can also extract the areas of interest in the pdf and run it via cmdline[1].You can get more details if required on my blog[2]
[1]https://github.com/tabulapdf/tabula-java/wiki/Using-the-comm...
Tesseract is acceptable only if the text is neatly laid out, in more-or-less straight parallel lines or at the very least consistent orientation that's close enough to being straight horizontal lines.
Google Cloud Vision, however, can read any orientation, any font, through perspective distortion and does not need the different text blobs in the image to be consistent in any way. Superior in every way to plain Tesseract (and if it is Tesseract after preprocessing, the magic is in that preprocessing more than it is in Tesseract)
I would actually be very surprised to hear GCV uses Tesseract; and if they don't, why would they use something inferior for other products?
Tesseract is the shittiest OCR and Google doesn't use it internally. Their cloud OCR offering is much more performant.
By all of that, what I mean to say is that I've learned a decent amount of fun OCR trivia over the past few years.
Firstly, the engine that powers Google Cloud Vision is almost certainly an entirely independent code base from Tesseract built on neural networks. In fact, the most recent major version of Tesseract (version 4.0) was a sort of rewrite of the core of Tesseract to use bidirectional LSTMs to seem a bit more like the modern OCR pipeline that systems like GCV use.
The original Tesseract algorithm dates back a previous AI spring— in the 80s when neural networks were cool (before they were uncool, and then subsequently cool again). The core of the original algorithm involved fitting polygons to character shapes in order generate features which could be matched by a kind of rudimentary neural network.
One of the primary authors of Tesseract is Ray Smith (at Google)— who gave a presentation at some point a few years ago about the history of OCR— though I can't quite find a link to it at the moment.
OCR actually predates electronic computers. In 1929, someone had invented a machine that would take a piece of paper and shine a bright light on a single letter, and pass the letter through a carousel of letter masks, so that it could hit a (effectively single pixel) photo-sensor. When the carousel and the letter mask were in alignment with the printed letter, then the drop in brightness registered that a particular letter was seen!
OCR was used by the US Postal system for sorting mail as early as 1965, but it wasn't until 1976 that any system could reasonably support more than a certain number of hard-coded fonts (fun fact this was invented by Ray Kurzweil, the Singularity is Near guy).
Major question:
Why isn’t Tesseract using neural networks. I know it just introduced LSTM based models but they suck.
Why is GCP vision text recognition API so much better than open source alternatives?!
I was looking for an OCR that can do license plates while the car is moving, for a hobby project. The image quality is less than perfect, the lighting is never very good, and as the camera is mounted on my side window, all plates have a perspective transformation applied (e.g., topline and baseline are essentially never parallel)
Tesseract fails miserably. Trying to help it, I have not found a good open source project that would consistently equalize color pictures to black-and-white - sometimes there's shadow on the plates that foils all simple attempts.
And yet, GCV needs no parameters, and seem to do this perfectly on images I've tried.
So, assuming I'm willing to put in the time - how do I build my own GCV -- even if it's just for the hobby use case of reading license plate (and the next stage: reading house numbers - which GCV does reasonably well, although it is a much much harder problem)
Tesseract was amazingly powerful and accurate, but it seemed to struggle if the page was warped or tilted even a little. I had to preprocess the images heavily to try to dewarp the natural spine curvature, and even then it could only get about 99% accuracy (which sounds like a lot, but consider a book where every 100th letter was wrong - I basically flagged the errors on my kindle as I went along and manually corrected them later).
I guess the point of this comment is that, in my experience, Tesseract.js is probably going to need an accompanying PageDewarp.js for it to be of use scanning books. Not everyone has access to a right angle scanner or can slice the spine and get perfectly straight high-res scans.
Years ago, I dug up a couple of papers on the spine thing, but never got around to implementing it. I think you can estimate the curvature based on the shadow and dewarp.
I was just scanning some recipes to save typing, so it wasn’t really worth the efffort.
[1] https://mzucker.github.io/2016/08/15/page-dewarping.html
Tesseract s is 2 pure Javascript port of the popular Tesseract OCR engine.
This library supports more than 100 languages, automatic text orientation and script detection, a simple interface for reading paragraph, word, and character bounding boxes. Tesseract s can run either in 2 browser znd on & server with NodeJS.
Not bad, but far from useful.Thanks for being interested in tesseract.js, it makes all the work worth. And I have to thank @antimatter15 for creating this library, without him we cannot go this far
I have read all the comments and here I would like to provide my two cents for some questions:
1. Is tesseract.js pure JavaScript?
Yes, it is 100% JavaScript and it leverages Webassembly port of original tesseract-ocr. (means we compile the C sorce code to JavaScript Webassembly code, powered by Emscripten)
2. The accuracy of tesseract.js is poor.
In my experience, it is hard to get perfect results without applying additional techniques to your source images. You may need to some preprocessing and sometimes train a custom traineddata. It is not easy, but it is the price of high accuracy.
3. Cloud OCR service is much more accurate
Yes, that's true. But tesseract.js provides an in browser offline option to do your OCR, it is useful for scenarios like PWA and high confidential image content (which you don't want to send to server). Tesseract.js is not a silver bullet, but it is handy sometimes.
Hope you enjoy this library and feel free to leave any comment to us!
Why isn’t there an open source OCR engine even half as powerful as the Google Cloud Platform API?!
Another extension (not using Tesseract.JS): https://chrome.google.com/webstore/detail/copyfish-%F0%9F%90...
A few weeks ago, I tried writing something that was long-running, CPU intense, ect, in a webworker. It was so darn slow that I switched to a native language. (I hope I didn't do something silly that made my code run more slowly than it should.)
I see some mention about running in WASM. Does this do something like have ordinary Tesseract compiled for WebAssembly and then fallback to Javascript?
冬 日 平 泉 路 晚 归
But the OCR reports:
冬 日 平 柳 路 晚 归
(Note the different character right in the middle)
I tried 4 or 5 different OCR programs, and none of them worked well enough for my case.
I was actually surprised, I thought OCR was a solved problem with ridiculously low error rates.
Seems so, lists at least Hindi, Urdu, Bengali, Sanskrit, Urdu, Nepali, Marathi, Sinhala and Punjabi!
ec
ij
tf
Il|1
hb
OQ
etc.
Even the tiniest addition or subtraction of printed ink can transform one character in one of the above rows to any other character in the row. Throw in page tilt/warp/etc. and OCR can frequently confuse them unless you train it specifically on your text. The pipeline I've found that works best is:
image -> upscale -> dewarp -> OCR -> spellcheck -> grammar check
For communication, Turbo Codes[1] for example have the decoder produce an integer value for each bit, rather than just a bit. The value is a measure of how likely the value is 0 or 1.
This is then used with previous bit values, which includes parity data, to make a "hard" decision.
I wonder if something similar has been tried for OCR? I imagine the OCR front-end could feed a number of probable hits, along with confidence, into a spellchecker. Or something along those lines.
A license or ID will almost certainly have medium-contrast elements in the background that will show up as dark. But if you were able to manipulate the contrast/brightness appropriately in advance, you could probably get it to work.
I'll test and migrate to this soon depending on accuracy. Great job so far.
So it is a wrapper library of a C++ project, which is cool. But saying it is a "Pure" JavaScript is purely misleading.