OCR It – pull text out of un-copyable documents for your LLM
github.com
github.com
I might be a bit behind, all of this is from early this year for the most part, but for something like "I have 3000 movie posters and I want to get the titles with like 90% accuracy" it is good (much better than Tesseract), and it'll do that in like an hour.
EDIT: I guess one thing is Tesseract will kind of give gibberish back when it fails. The main issue with the LLMs are that instead they take a stab at it (like for a movie poster it'll give part of a quote, or a actor name) back. Makes knowing when it fails a little harder. As long as you have some way to verify when it is likely failing they are very good though.
Its much lower and the error rate is much higher. It works, but for our usecase, "clean rendered text" (extract from kindle for example) tesseract is much better.
For the input text rendered on screen, Tesseract did better on both accuracy and speed. We got about 0.1% character error vs 1–2% for RapidOCR, and Tesseract was roughly 2.5x faster. Blur was the biggest difference: 0.4% vs 14%.
The big problem is that this is synthetic rendered text, which is basically the easy case and also the only kind of input this extension captures. I wouldn't assume the same results for scanned documents.
I haven't tested EasyOCR yet.
Which models are EasyOCR and RapidOCR using?
Jury is still out on which is more trustworthy handling any personal data, Microsoft or Google. Neither.
“Pin a region once. Hit a hotkey on every page. Get the whole book as text.”
Much better than the old definition of “region lock”, nice.HN isn’t a fan of the generated readmes though, though vibed software (thoroughly used) can be all good.
Yes, but the problem with these vibe-coded crap is that they are pretty much always less than a week old, which means it wasn't even used before the “author” submitted it here.
(The author didn't even bother writing their comment themselves by the way: https://news.ycombinator.com/item?id=49415857)
OK yeah seemed tedious so figured that must’ve not been the only way [probably if you handwrite you could clear that up beforehand]
& the comment here https://news.ycombinator.com/item?id=49415857 actually violated the guideline as noted by another here https://news.ycombinator.com/item?id=49417725
Note since I last posted: looks like someone vouched for the comment posted by the account created at the same time as OP’s post, so no longer dead
It has a similar workflow for taking screenshots and then immediately annotating or editing them, without having to open a separate image editor. And: it provides also an local OCR feature (which is why I comment this here), you can extract text from a screenshot with on-screen OCR using Tesseract with the small button beside the "Crop Image" one.
Combined with the syntax-highlighting feature for screenshots of code snippets, the OCR is surprisingly useful in combination if you e.g. quickly discuss some code in a chat when copy is blocked for whatever reason (e.g. somone sent you a screenshot in the first place).
[1] https://gradia.alexandervanhee.be/
[2] https://flathub.org/en/apps/be.alexandervanhee.gradia
Edit: fixed wrong link index numbers
Until it's approved you guys can download the ready to use releases:
Download ocr-it-firefox-0.3.0.zip from https://github.com/thiagotigaz/ocr-it/releases/tag/v0.3.0
If you'd rather build from source, the steps are in the README: https://github.com/thiagotigaz/ocr-it#install
Local: https://github.com/datalab-to/chandra Hosted: https://www.datalab.to
Another decent option is GLM OCR. It's slightly less accurate but faster and cheaper.
Local: https://github.com/zai-org/GLM-OCR Hosted: https://docs.z.ai/guides/vlm/glm-ocr
Other models such as PaddleOCR, dots.ocr and DeepSeek OCR performed significantly worse.
Its rather interesting if it's correcting a mistake or picking a correct alternative.
How accurate do you think it is overall?
If you have a coding agent available, ask it to try transcribing a few of the PDFs.
We are still waiting for the firefox extension to be approved, i will post it here whenever we hear something. In the meanwhile it can be tested with the zip file here https://github.com/thiagotigaz/ocr-it/releases or by building manually.
OCR It is a Chrome extension for that gap. You drag out a capture region once — the text block of the reader, say. After that, one hotkey per page screenshots that exact rectangle, OCRs it, and appends the result to a running transcript. Or start an auto-run and it captures, turns the page, and repeats until the document ends. Then Copy all, or Download .txt, and you have a file to paste into Claude or drop into an agent's context.
Everything runs locally. Tesseract's wasm build and the language data (~10 MB) are committed into the extension, so there are no network requests at all, no API key, and no host permissions at install — single captures ride on activeTab. The irony of an AI-adjacent tool that never talks to a server was not lost on me, but the pages you're capturing are often exactly the ones you don't want to ship to a third party.
Three things turned out more interesting than expected:
- MV3 service workers have no DOM and no Worker, so cropping and OCR live in an offscreen document.
- The next-page control is stored as a point, not a CSS selector. A point survives DOM re-renders and reaches into cross-origin iframes and shadow roots, which nothing the top frame can express does. Routing it was the fiddly part: window.screenX inside an iframe reports the browser window, not the frame, so frames locate themselves by walking same-origin ancestors, and across an origin boundary the parent hands the offset down by postMessage.
- The auto-run waits for each page's OCR before turning. That's what makes end-of-document detection work; a timer-based loop sails past the last page and fills your transcript with copies of it.
Limitations: Chrome's own PDF viewer can't be auto-advanced (it's a plugin no extension can inject into, though capturing from it works fine); the region is a fixed rectangle on screen, so resizing or zooming mid-run breaks it; and accuracy tracks the source — crisp rendered text reads at 93-95% confidence, scans need cleanup before they're worth feeding to anything.
Tests drive a real headless Chrome over CDP, which had its own surprises: Chrome 137+ ignores --load-extension, and headless can't show the optional-permission prompt, so the suite installs a copy with the grant baked in plus a real toolbar click via Extensions.triggerAction to prove the ungranted path still works.
MIT, no build step: https://github.com/thiagotigaz/ocr-it
Don't post generated text or AI-edited text. HN is for conversation between humans.
https://news.ycombinator.com/newsguidelines.html