Tesseract.js wraps an Emscripten port of the Tesseract OCR Engine
github.com
github.com
I have little expertise in ML, but from my limited understanding, OCR is the bread and butter of the field. I've read exactly one "Intro to ML" article and it was about recognising digits. And yet, we have an abundance of high quality proprietary OCRs that can recognise printed or even hand-written text and the single open source one is having trouble with perfectly formatted text with a readable font.
Could anyone with more expertise shine some light on this current state of affairs?
I had the opposite experience.
My partner was doing a project for the Army Core of Engineers and they only provided information via some system called ProjNet that, best I can tell, exported PDFs of Web Pages in pure vector format so they were unsearchable. Of course they needed to search 10000 pages of documents to answer questions for the ACoE.
I was able to feed the PDFs into Tesseract and produce 1:1 text document per page of PDF and then marry it back up to the PDF so they could search the PDFs. It worked astonishingly well and took about a half an hour using the cringiest of shell scripts.
I did something similar with SDGE's published rate tables to convert their screenshots of XLS files back into tablur data. It didn't work as well but still got the job done.
Tesseract.js – A Javascript port of the Tesseract OCR engine - https://news.ycombinator.com/item?id=28105850 - Aug 2021 (37 comments)
Tesseract OCR - https://news.ycombinator.com/item?id=27876383 - July 2021 (65 comments)
Tesseract Teaser - https://news.ycombinator.com/item?id=26400168 - March 2021 (7 comments)
Tesseract.js: Pure JavaScript OCR for 100 Languages - https://news.ycombinator.com/item?id=21843713 - Dec 2019 (77 comments)
A guide to OCR with Tesseract, OpenCV and Python - https://news.ycombinator.com/item?id=21843342 - Dec 2019 (12 comments)
Using Tesseract OCR with Python - https://news.ycombinator.com/item?id=14741124 - July 2017 (47 comments)
Show HN: Tesseract.js – Pure JavaScript OCR for 60 Languages - https://news.ycombinator.com/item?id=12694004 - Oct 2016 (97 comments)
EasyOCR is definitely interesting and something that's worked well for us at a prototyping level.
Thanks for sharing these -- it's maybe just my very bad searching skills but I had been trying to set some stuff up with Tesseract and had come to the conclusion that I just couldn't use it for document photos and would either need to abandon that effort and buy a faster scanner, or hook into some proprietary service like Google/Apple.
Both of these look really promising, so now I'm excited again about the potential of setting up a fast Open Source way to digitize my documents.
Sure, I could use someone else's subtitle file from the Internet, but that's not as fun than doing it yourself.
The actual code that does the OCR is wraped and included via this package [0] which just wraps the original Tesseract in C++ [1] using wasm. Shameful title.
For my own purposes, the priority for me when reading "pure" is that the core runtime I'm using (the JS runtime) is the only runtime dependency - I'm not depending on external binaries and execution environments like an FFI implementation would.
It also opens up codebases to browser compat, where FFI would typically not be available.
WASM isn't an interface or a wrapper, it's a language/format. Having trouble understanding what you mean by this, unless you're arguing that the WASM VM itself is the FFI?
WASM "embeds" modules within the JS runtime, in a similar way that traditional FFI "embeds" native bindings compiled separately & externally. It's still quite different insofar as the VM is a part of the runtime, but there are vague parallels.
For me though, the practical problems related to runtime env that one encounters with traditional ffi bindings calling dynamically linked native libraries are rarely present with a WASM library, as the support within the runtime is explicit (the only real exception here is architecture, which is always an issue regardless).
I think the implications of that code are different, but yeah, I see your point and I think it's fairly reasonable.
> A foreign function interface (FFI) is a mechanism by which a program written in one programming language can call routines or make use of services written in another.
I use it in an Electron project and a documents that takes about 1.5 sec per page with the Tesseract CLI, I can get down to about 15 sec with Tesseract.js with parallelization.
Literally third sentence of the description:
> It works in the browser using webpack or plain script tags with a CDN and on the server with Node.js.
It may miss a few features (some which I needed I had to code in).
Is there any single-binary static OCR tool comparable to these two?
I can see where you're coming from, but I've never used or heard anyone in the web world use "pure" to mean only "written entirely in Javascript without transpilation or other tools."
If it hits the parts of "pure JS" that most people care about:
- it's running entirely in Javascript.
- it has no native dependencies.
- it can run entirely clientside.
- it can be embedded in a normal web page.
then I think most people will be fine with using "pure" to describe it.
----
I wouldn't even have that many quibbles with their phrasing even if they were compiling to WASM. Sure, at that point it wouldn't be running as pure javascript, but it would still hit 3 of the 4 points above.
(But as already mentioned, JS is still supported today, using wasm2js.)
I'll concede though that if they're targeting WASM it's not technically pure Javascript, but it still feels a bit to me like splitting hairs since it's always been explained to me that WASM and the Javascript runtime under the hood have a lot of overlap.
https://github.com/WebAssembly/binaryen#wasm2js
The emcc flag -sWASM=0 disables the wasm final output and emits JS instead.
The only place I've seen "pure JS" used in the JS world is differentiating e.g. a Postgres client implementation that does or doesn't depend on the specific version & configuration of libpq you have on your current system. That's about runtime deps, not about source language.
That’s exactly what “pure” means - “pure rust”, “pure go”, etc. IMO you can’t say the heart of all the work is a c++ lib and call it a “pure JS” anything.
More accurately / correctly / usefully would be calling it “JS wrapper over a c++ library cross compiled to JS”. Or maybe “All JS at runtime” or some other qualifier.
Want to see how to cross-compile a non-trivial c++ lib? Check out here!
Want to see a great JS wrapper library where we had to make cross-language & cross memory-management API decisions? Check this out!
But - want to read some awesome high performance image processing algorithms in JS? Not this.
I'm not sure where the line here is supposed to be drawn, but I don't personally think looking at asm.js code is significantly harder than looking at something like compiled Typescript, and certainly it's not any harder than looking at minified Javascript. If I'm going to be debugging code, they're both going to be annoying to look at. We're quibbling over definitions so I'm not going to say that you're wrong, "pure" can mean whatever you want it to mean. I'm just saying that most JS devs I know consider (for example) JSX code to still be pure Javascript when it's compiled.
It seems a little odd to me to look at something where every single line of code is Javascript, being run entirely in a Javascript interpreter, and say that isn't actually completely real Javascript, but if the programming circles you frequent are different and think about this differently, :shrug: more power to you.
> But - want to read some awesome high performance image processing algorithms in JS? Not this.
I also wouldn't necessarily assume that every Open Source program written in only JS without compilation is going to be well suited for reading or learning from. But again, sort of splitting straws here -- my only concern is that you're probably going to be disappointed a lot if you equate "pure Javascript" with "readable".
That being said perhaps a poll is needed to find out what most people think.
I mean, if someone compiles a Markdown document and sticks the result on their website, do you say, "this isn't HTML"? People are free to use words however they want I guess, but I don't understand the perspective where the way a project was written suddenly means that the compiled result isn't Javascript -- it feels like it's taking the word "pure" in a metaphysical direction that I just don't really grok.