Show HN: Tesseract.js – Pure JavaScript OCR for 60 Languages
github.com
github.com
A much better example that works quite well is a picture of someone holding a book: http://i.imgur.com/3JWs64x.jpg
Magic .
Read this to yourself. Read it silently
Don't move your lips. Don’t make a suund
Listen to yourself. Listen without hearing
What a wonderfully weird thing, huh?
NOW MAKE THIS PART LOUD!
SCREAM IT IN YOUR MIND!
DROWN EVERYTHING OUT.
Now, hear a whisper. A tiny whisper.
New, read this next line with your best crotchety—
old-man voice:
“Hello there, sonny. Does your town have apost 0
Awesome! Who was that? Whose voice was that?
It sure wasn’t yours!
How do you do that?
How?!
Must be magic.
Problems with this text: misspelled 'sound' as 'suund', didn't recognize the word 'anything', and mis-recognized 'a post office' as 'apost 0'.Not bad. Especially since two of three mistakes are on the edge of the page.
Can someone try and see how it would perform if you simply upscale the image using normal bicubic interpolation? And if it performs much better, I feel like that should be a preprocessing option to scale up the image since it seems to do so poorly on small resolutions.
https://github.com/naptha/tesseract.js#tesseractrecognizeima...
> Note: image should be be sufficiently high resolution. Often, the same image will get much better results if you upscale it before calling recognize.
[0] https://www.microsoft.com/en-us/research/publication/stroke-...
Original text [http://i.imgur.com/CZGhKhn.png]:
> I am also a top professional on Thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well
Google detects [http://i.imgur.com/pSJym1x.png]:
> “ I am also a top professional on Thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well ”
Tesseract detects [http://i.imgur.com/wwbLU6g.png]:
> \ am also a mp pmfesslonzl on Thummack wmcn Is a sue 1m peop‘e \ookmg (or professmna‘ semces We on glg salad P‘ezse see my rewews 1mm my cuems were as weH
Edit:
A screenshot of the same text at a higher resolution: https://imgur.com/a/W7IGu
Tesseract.js output: https://imgur.com/a/niIfM
"I am also a top professional on thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well"
Tesseract.js analysis:
Although Googie's API is certaihiy better,
Tesseract.js should work simiiarly if you
increase the font size.
Screenshots taken
on 'retiha’ devices are around the smailest
text it can handie well.
Edit:
A screenshot of the same text at a higher
resolution:
httgs:[[imgurxomZaN/UGu
Tesseract.js
output: httgs://imguricom[a[hiIfM
This is a neat toy, but not impressive compared to the results from tesseract-ocr/tesseract [0]: $ curl -s http://i.imgur.com/uuFhw90.png \
| tesseract stdin stdout
Although Google's API is certainly better,
Tesseract.js should work similarly if you
increase the font size.
Screenshots taken on 'retina' devices are
around the smallest text it can handle well.
Edit:
A screenshot of the same text at a higher
resolution: https:[ZimguncomlalWHGu
Tesseract.js output:
https:[[imgur.com[a[nilfM
Notice how Tesseract.js results suffer from being unable to differentiate between n's and h's, i's and l's.Edit: In addition to those differences, I think your font size is still a bit too small. On an unedited screenshot from a macbook (https://i.imgur.com/iv4ZdSt.png) I get
Although Google's API is certainly better, Tesseract.js should work similarly if you increase the font size. Screenshots taken on 'retina' devices are around the smallest text it can handle well.
Edit:
A screenshot of the same text at a higher resolution: https:[[imgur.com[a[W7IGu
Tesseract.js output: https:[[imgur.com[a[niIfM
“I am also a top professional on thumbtack which is a site for people looking for professional services like on gig salad. Please see my reviews from my clients there as well" > This might have to do with the way we threshold images,
> with the age of the tesseract version we're using, or
> both. I'll look into it!
I'd be very interested to hear about what is required to make it "match" native functionality. Please do drop me a line if/when you get it figured out! (I'm @jtaylor on twitter [1])https://github.com/naptha/tesseract-emscripten/blob/master/j... specifically the line for lepton
Edit: just tried Google's, and it had one mistake for that entire file. That's pretty impressive.
Upscaled image: https://imgur.com/a/4IQA7
Result on demo page: http://imgur.com/a/A0v5C
The hammerman Tikes flsosushsath: Greetings. My name is Tikes Leafsilk.
You: Rh. hello. I'm Stasbo Murderknower the Craterous Trance of Fins. Don't travel alone at night. or the bogeyman will get you.
You: Tell me about this hall.
Tikes: This is The flccidental Ualley. In 123, Stasho Steamdances ruled from The flccidental Ualley of The Council of Cobras in Ueilapes.
...
I spent a weekend tying Tesseract together with Tekkotsu (the amazing open framework for the Sony AIBO) in an attempt to teach my robot dog to read. The eventual goal being to hook up the output of OCR --> Text To Speech (TTS) and have him read to me.
Alas, the low resolution of the camera was an insurmountable problem. Poor Aibo needed 40-point fonts and I practically had to rub his nose in the book. Not exactly the user experience I was aiming for.
Never got around to the TTS part.
I did something like that a few years ago when making an Eve-Online UI scraper.
Written by a Google employee, see top link at http://www.imjasonh.com/projects
"Price per 1000 units. Unit volumes are based on monthly usage." It's $2.50 per 1000 units, so 0.25 cents per unit.
Edit: And, according to the pricing page [1] the first 1000 units are free.
For higher volumes, there is also the OCR.space api. It offers 25,000 free conversions per month. It is not as good as Google, but works fine on screenshots.
I routinely (daily) need to OCR PDF files. The PDF files are not scans. They are PDF files created from a Word file. The text is 100% clear, the lines are 100% straight, and the type is 100% uniform.
And, yet, Microsoft and Google OCR spits out gibberish that is full of critical errors.
From a problem solving perspective, this seems like an incredibly easy problem to solve in this exact use case. That is, PDFs generated from text files. Identify a uniform font size (prevent o-to-O and o-to-0 errors), identify a font-family (serif, sans-serif, narrow to particular fonts), and OCR the damn thing. And yet, the output is useless in my field.
I am unsure as to why he can't just copy / paste.
I don't even care about perfect formatting, that's easy to fix. I do care about perfect OCR. That's crucial.
So, I do have to use OCR, right?
https://source.opennews.org/en-US/articles/introducing-tabul...
The main advantage of ABBYY is that if you need to do OCR, it is, in my opinion, the best consumer-level package. And it does a pretty good job of doing OCR and conversion to Excel. Here's a Github repo that demonstrates some results:
https://github.com/dannguyen/abbyy-finereader-ocr-senate
But to reemphasize, the above repo demonstrates ABBYY maintaining table structure with PDFs that are scanned images, which is considerably harder than the situation you're in.
I've started a repo that eventually will compare text-to-table tools, which is what you want: https://github.com/dannguyen/pdftotablestable
In principle, text-pdf-to-text is just a matter of parsing PDF (and/or Word) formats and extracting text buried in metadata. (I know it's a lot of work but still).
Even if you forget about what GP said about the source being text PDFs, and when all the sources are png images, as long as those pngs were generated from text documents (Word, PDF, etc) without any scanning or camera involved, it is unacceptable that today's free OCR tools don't get the job done, when in 2016, machine-learning has produced systems that have surpassed human accuracy in much harder tasks like object detection and speech recognition.
I know it's not an unsolved problem. It's just a matter of some knowledgeable machine learning researcher taking a break from working on cutting edge for a few months and putting together a package that gets the image-to-text job done. Once such a base tool is available on github, the community will take over and add features, fix bugs, as needed. (I'm extremely busy with my own degree work ATM, otherwise I would probably do something like that).
EDIT 1: As for tesseract, I hate it with the passion of a thousand fiery suns. It's a kludge, a black-box of traditional-programming karate-chops and overly-complicated bloat that spits out text the way it likes and there is, largely, nothing you can do about it. Compared to machine-learning and modern computer-vision, tesseract belongs to the dark ages. If there is going to be a quality OCR tool, it's has to be written from scratch based on deep-learning from the ground up.
1. Use an existing OCR library to give you the positions of the words, plus a first-cut guess of their content.
2. Take the first word from the OCRed guess, and loop through a set of {font, size, leading} tuples, rendering out the same word at that {font, size, leading} and overlaying it on the image, and measuring error-distance.
3. If your best match isn't within some minimum error-distance, then assume that the OCR misrecognized the first word, and try again with the second, third, etc.
Once you've got a font-settings match:
4. render the rest of the words onto their respective detected bounding boxes;
5. notice which words have a higher error-distance than the rest;
6. for each word, generate candidate mutations of the word (e.g. everything at a Levenstein distance of 1 from the OCRed guess), pick the one that lowers the error-distance, and repeat until the distance for that word won't go down any lower.
7. Return the error-minimized set of words.
You could call this a form of https://en.wikipedia.org/wiki/Code-excited_linear_prediction, with fonts as the pre-trained models.
---
Actually, come to think of it, it'd be a lot easier to detect and unify "identical" sub-regions of the image first (using e.g. https://en.wikipedia.org/wiki/JBIG2 on a lossless setting). Then you could, in parallel to the above, also try to do frequency-analysis to discover which of your image "tiles" would likely form a basic "alphabet" of character-glyphs—and then hill-climb toward aligning that "alphabet" by attempting to produce the most runs of character-glyphs that translate to known dictionary words in whatever language the OCR thinks the text is in.
The font-matching would still be necessary, though, for the rest of the image samples that don't fall into the easily-frequency-analyzed part. (And for languages that aren't alphabetic, like Chinese, where there are no super-common character-glyphs.)
The PDF files that we are dealing with do not have embedded text and are not searchable, but are "digital-native," to use the term that you suggested.
Does this not exist? If not, why does it not exist?!
"Tesseract works best on images which have a DPI of at least 300 dpi, so it may be beneficial to resize images."
- https://github.com/tesseract-ocr/tesseract/wiki/ImproveQuali... Tesseract.recognize(myImage)
.progress(function(message){console.log(message)})
.then(function(result){console.log(result)})
.catch(function(err){console.error(err)});
or Tesseract.recognize(myImage)
.progress(function(message){console.log(message)})
.then(
function(result){console.log(result)},
function(err){console.error(err)}
);
I guess I just still have bad memories of jQuery's old almost-like-real promises. I'd rather never have to think ever again about whether I'm dealing with a real promise or one that's going to surprise me and break at run-time because I tried to use it like a real one. Promise.resolve(Tesseract.recognize(myImage)).then(result => console.log(result))For example, I took a screenshot of this comment and ran it through the demo and got this:
Excited ehent this... but the OCR enenty Seems te be very bad. Maybe it's het Dptimized far recngnizing black text an e white heckgmnhe. EDI example, 1 tank e Screenshnt at this cement ehe teh it. thmneh the den» ehd get this:
It seems to recognize the bounding boxes just fine but mangles the words.
Excited about this... but the OCR quality seems to be very bad. Maybe it's not optimized for recognizing black text on a white background. For example, I took a screenshot of this comment and ran it through the demo and got this:
1: https://github.com/goatslacker/pokemon-go-iv-calculator/blob...
Thanks for putting this out.
We'll see how it goes.
This looks great, and I'd really love to but
> Uncaught ReferenceError: progress is not defined
EDIT: works now!
Edit: this affected every browser because it was a typo. Fixed!
My current HRCloud2 project could benefit greatly if I ever get around to it. Currently I make the php interpreter jump through hoops and move stuff all over the place to OCR images and docs. This could save a ton of time and shift the processing to the client instead of my server.
You can even host the external language files yourself as described in the readme: https://github.com/naptha/tesseract.js#local-installation
> Dropan Enghsh Wage on (Ms page to OCR m
Should be
> Drop an English image on this page to OCR it!
It should work better if you feed it a screenshot of the black text at the top of the demo page though (Tesseract.js is a pure Javascript port etc...).
Well it's pure JS in that it's been running the C tesseract through emscripten. So in a way it's pure JS just as much as the original lib is pure assembly when compiled ;-)
I threw together a quick proof-of-concept in Go for exposing tesseract via a web API:
For languages that don't employ much whitespace, this would be nice.
They've added https://github.com/naptha/tesseract.js/blob/c26cae7ee956c399...