Tesseract.js – A Javascript port of the Tesseract OCR engine
tesseract.projectnaptha.com
tesseract.projectnaptha.com
First issue I've encountered was the text recognition performance. Depending on the camera input (if the image contained something that looked like the text or not) I've got 2-20+ seconds per 640x640px image for text recognition on iPhone X. Not so fast as you may see. But the recognition was pretty accurate though.
The performance, as expected, improves when the image size is getting smaller and the amount of text on the image is also smaller.
Since I did't want to recognise the whole text, but only the links, I've used the TensorFlow Object Detection model to quickly find the areas with the text http://**. Then, instead of recognising the whole image I needed to do it only for smaller parts of the image. This gave some improvements to the performance: from the variable 2-20 seconds per frame I've got more stable 0.5-1 seconds. Also not good, but several times faster.
I've described the challenges in more details here https://trekhleb.dev/blog/2020/printed-links-detection/. But to sum up, I had a good recognition quality with an arguable performance with Tesseract.js
The sad thing is most of the state of the art models and algorithms are open research, they just are usually not written by software engineers and need to be rewritten to be deployable. Usually you just get some shell script like "run_eval.sh" that generates the figures in the paper through a bunch of spaghetti code, and most of the time it will depend on a specific old version of Tensorflow, that probably isn't available for your CUDA version, and probably won't compile on your system without hours of Googling.
There's a mode where you can increase the number of worker threads. Tesseract is also designed for text documents and the preprocessing filter I made to convert the images to look more like a text document was pretty naive.
I'm taking an online computer vision class next semester and hope to pick the project back up after learning a bit more.
Sad as it is to say, it's just not up to snuff for any application I've tried it on.
I found an error in the chinese demo, with the example you provided (4th character wasn't the same). I know no OCR is perfect, but IMHO at least your own demo should be free of errors.
:) That would be a dishonest demo.
You try to show how well it works, not that it works perfectly well (which is false). Edit: especially since we know that OCR is hardly perfect - we expect errors to be minimized, not absent, and the first interest is to see where the engine fails.
Calling this “pure JavaScript” seems misleading
Theretically, cross platform support would be another possibility. But one could argue native C code could be bundled as well, albeit with separate integration being needed. (Android and iOS do support such extensions).
Also privacy: running OCR in someone's browser rather than sending the images back to the server keeps them fully in control of the data they are working with.