Running OCR against PDFs and images directly in the browser
simonwillison.net
simonwillison.net
Not saying that the article was being misleading about this, just saying that the LLM part is basically doing some standard interfacing and HTML/CSS/JS around that core engine, which wasn’t immediately obvious to me when scanning the screenshots.
The point here is more about highlighting that browsers can do this stuff, and it doesn't take much to wire it all together into a useful interface.
It matters in this case because for tesseract, the exact model is incredibly important. For example, v4 is pretty bad (but what is available on most linux distros when ran serverside) whereas v5 is decent. So I would have had a more accurate interest in this post if it was a bit more upfront that "Tesseract.js lets you run OCR against PDFs fairly quickly now, largely because of better processors we as devs have, not because of any real software change in the last 2-3 years".
I felt this before for his NLP content too - but clearly it works because he's such a great explainer and one for teasing content later that you do read it! I must say I've never been left confused by Simons work.
What else should I have done?
Was an earlier headline or subtitle here on HN what was misleading, but then that was changed to not be misleading?
The PDF parsing is done by the excellent PDFplumber Python library [2], the web app is built with Elixir's Phoenix framework and it is all hosted on Fly.io.
[1] https://data-tools.fly.dev [2] https://github.com/jsvine/pdfplumber
I built this because most of the basic layout detection libraries have terrible performance on anything non trivial. Deep learning is really the long term solution here.
https://notes.joeldare.com/handwritten-text-recognition
Tesseract was one of the tools I tested, although I used the CLI instead of the WASM version.
1. Pre-process PDF images to detect letters better?
2. Use LLMs to spell/grammar check and perhaps even auto-complete missing pieces?
3. Employ rich text to capture style (ex: lexical.dev)?
Unsure if it is feasible to bundle it all up for web.
See also: https://github.com/RajSolai/TextSnatcher / https://github.com/VikParuchuri/surya
Not sure how difficult it is to run it in the browser, though.
I would really want human review. Remember that copier that changed digits because it was being clever with compression?
A PDF to image processor is being built and should be out in a few weeks
No docs, but happy to help anyone wanting to use either. Email is kord @ the company I'm working on.
It would miss some cells from a table, or does not recognise all the numbers when they have commas.
There are a ton of potential tools out there like Tabula and AWS Textract table mode but none of them have felt like the perfect solution.
I've been trying Gemini Pro 1.5 and Claude 3 Opus and they looked like they worked... but in both cases I spotted them getting confused and copying in numbers form the wrong rows.
I think the best I've tried is the camera import mode in iOS Excel! Just wish there was an API for calling that one programmatically.
Seems like it can handle tables.
I just grabbed a two-column academic PDF, which performed as well as you would expect. If I was returned a json list of text + coordinates, I could do some dirty munging (eg footer is anything below this y index, column 1 is between these x ranges, column 2 is between these other x ranges) to self-assemble it a bit better.
const {data: {text}} = await worker.recognize(imageUrl);https://mdn.github.io/dom-examples/web-speech-api/speak-easy...
On this page: https://en.wikipedia.org/wiki/Typeface it takes almost ~10 seconds for the text in the first image to become selectable after page load.
With local images / PDFs in Preview it's really quick though
You can also use MacOS's OCR capability to create a shortcut that allows you to copy and paste the text out of any region on the screen -- for example, a stack trace someone is showing you in a screen share.
I use it for ingest of image and pdf type files for my own website chatting tool: tinydesk.ai!
I run the backend on an express js server so all js as well.
Smaller docs I do on the client side, but larger ones (>1.5mb) I've found take forever so those process in the backend.
https://github.com/simonw/s3-ocr
You would need to upload them all to S3 first though, which is a bit of a pain just to run OCR (that's Textract's fault).
You could try stitching together a bunch of scripts to run the CLI version of Tesseract locally.
I always wanted to make a chrome extension for one thing or another, but all the learning involved around the boilerplate always drained the motivation. But with GPT I built the initial POC in an hour and then polished and published it on store even. Recently I compiled some bash and cmd helper scripts, I don't know either of these enough (do know some bash) and don't have it in me to learn them. Specially the windows batch scripts. Using LLM it was matter of an hour to write a script for my need as either a windows batch script or even bash script.
Oh I even used GPT it to write 2-3 AutoHotKey scripts. LLMs are amazing. If you know what you are looking for, you can direct them to your advantage.
Very exciting to see that people are using LLMs similarly to build things they want and how they want.
> The LLM part is almost irrelevant to the final result to be honest
Oh I think I see. No, I'm not being paid to promote LLMs.
The point of my blog post was two-fold: first, to introduce the OCR tool I built. And second, to provide yet another documented example of how I use LLMs in my daily development work.
The tool doesn't use LLMs itself, they were just a useful speed-up in building it.
It's part of a series of posts, see also: https://simonwillison.net/tags/aiassistedprogramming/
> No, I'm not being paid to promote LLMs.
This is good enough for me, thanks.
The files become big but it just works.
If an alternative is quicker / lighter I will use it but it must just works.