OCR at Edge on Cloudflare Constellation
notes.willhackett.com
notes.willhackett.com
I know other companies have struggled with demand, so maybe they're doing it on an invite basis.
[1]: https://blog.cloudflare.com/workers-ai/ [2]: https://www.cloudflare.com/nvidia-workers/
I've ran face detection on cpu and it takes 1 / secondv and it's probably a pretty intensive "action".
Plenty of use-cases that don't require a gpu
And plenty that require one too though.
That would mean support in every dc and a easily to expand capacity.
They can't just make it available in one region and then see how it goes.
I like how cloudflare works, but this use-case for them ( on the edge) seems more difficult to plan. It's not just the tech, it's the infrastructure in this case.
Just my 2 cents
- Takes 1.5 seconds to run, so there goes the "edge" benefit.
- And if it takes that long, it would cost more than having it running on a proper instance/vps.
- if you have infrequent access patterns you don't pay for the time it isn't used; and
- if you have huge bursts (say your AI project gets on the homepage of hacker news) capacity automatically scales with demand.
In return each compute-second costs more than a normal vps, so there is some threshold where doing it yourself is more cost effective.
The benefit of "edge" here is imho that it works together with other cloudflare "edge" stuff. Not hugely useful for OCR, but imagine your blog on Cloudflare Pages has a contact form (with a CF Pages Function) and you want an AI spam filter; or bot detection that is updated each time a page is visited, or in the auth function you have offloaded to a CF Worker; or if you want to enhance your email-triggered worker with AI.
- Tesseract5 *demolished EasyOCR on paragraph detection, getting that 100% on the 10 pages I checked. EasyOCR missed most of the paragraph breaks.
- Tesseract got most of the punctuation correct, EasyOCR only got apostrophes and two double-quotes (out of 14) correct. Every single period, comma, exclamation mark, and hyphen was missing or wrong, as were most of the double-quotes. Some question marks were recognized, but with garbage after them.
- In general EasyOCR seems to just add in square closing brackets ("]") where none are
* What language are you trying to OCR? And only language or also things like math symbols? * Do you have a GPU or not? * Are you trying to OCR handwriting or typed words?
I explored OCRing English documents from the 1960s that were primarily typed, though some handwriting. I tried out PaddleOCR, TrOCR, Tesseract, EasyOCR, and kerasOCR for FOSS, and then Google, Amazon, and Microsoft for paid.
To be clear, the paid solutions beat the FOSs ones handsdown, no question. However for FOSS I found that TrOCR was the best for both typed and handwritten, however for typed, it was closely followed by tesseract, but for handwriting TrOCR was by far the best with all the others basically being worthless. However, TrOCR took ~200x longer even on GPU than Tesseract on CPU (Tesseract if fastttt, even more if you parallalerize it). Tesseract isn't the best, but it's the best all around, it's the one the Internet Archive uses.
Need to write up a blog on this. And the docTR looks interesting, I'm going to check that out.
I’m obsessed. Very keen to try PubSub next.