How to do OCR on a Mac using the CLI or just Python
blog.greg.technology
blog.greg.technology
I used a combination of RHetTbull's vision.py (for the actual implementation) [1] + ocrmac (for experimentation) [2] and was pleasantly surprised by the performance on my i7 6700k hackintosh.
I wouldn't call myself a programmer but I can generally troubleshoot anything if given enough time, but it did cost time.
[1]: https://gist.github.com/RhetTbull/1c34fc07c95733642cffcd1ac5...
Could you run a farm of macOS machines and turn this into an API for profit? Would that be legal?
I have never gotten truly garbled output from Apple’s, whereas Tesseract will frequently produce random Unicode characters from text.
Apple’s also handles things like overlapping text or changing font sizes and typefaces far better than any open-source OCR I’ve used.
https://findthatmeme.com/blog/2023/01/08/image-stacks-and-ip...
AWS textract provides sample python code to extract tables into csv which works great.
Have you compared it with Textract?
If you look at RAG frameworks as one example they'll typically use/support a variety of implementations. Tesseract is almost always supported but it's rarely ideal with projects like Unstructured[0] and DocTR[1] being preferred. By leveraging more-or-less SOTA vision models[2][3] they embarrass Tesseract.
I haven't compared them to the Apple Vision framework but they're absolutely better than Tesseract and potentially even Apple Vision.
There are also various approaches to use these in conjunction but that gets involved.
[0] - https://github.com/Unstructured-IO/unstructured-inference
[1] - https://github.com/mindee/doctr
[2] - https://github.com/mindee/doctr#models-architectures
[3] - https://github.com/Unstructured-IO/unstructured-inference#mo...
[0]https://docs.aws.amazon.com/textract/latest/dg/how-it-works-...
https://github.com/JaidedAI/EasyOCR#whats-coming-next
Happy to see OCR is advancing lately, but I really need HWR.
I am looking for something this polished and reliable for handwriting, does anyone have any pointers? I want to integrate it in a workflow with my eink tablet I take notes on. A few years ago, I tried various models, but they performed poorly (around 80% accuracy) on my handwriting, which I can read almost 90% of the time.
How well it works on your handwriting is for you to test, but if you, having all kinds of contextual information, can’t read it well, I guess it won’t, either.
docTR comes out as strongest open solution.
[1] https://learn.microsoft.com/en-us/windows/powertoys/
[2] https://learn.microsoft.com/en-us/windows/powertoys/text-ext...
I am impressed how it handles handwriting and crappy screen grabs.
Will we ever have programming languages that are primarily designed to take input from whiteboard grabs? (ie where not only handwriting, but also placement, connectivity, and maybe shape are meaningful?)
Also if it’s text of a URL/domain or a QR code (eg in a photo of a poster, or in a video) you can hold-press/hold-click to open the link directly from the image.
What are the advantages over native macOS shortcuts these days?
https://learn.microsoft.com/en-us/windows/powertoys/text-ext...
PyXA uses the Vision framework to extract text from one or more images at a time. It's only a small part of the package, so it might be overkill for a one-off operation, but it's an option.
ImageAnalyzer is newer and much better
I bet this shortcut from OP is also using the older API under the hood
OCRTHISFILE="ocr-test.jpg"
shortcuts run ocr-text -i "${OCRTHISFILE}"
pbpaste > ${OCRTHISFILE}.txt
or to view output and place in file:
OCRTHISFILE="ocr-test.jpg"
shortcuts run ocr-text -i "${OCRTHISFILE}"
pbpaste | tee ${OCRTHISFILE}.txt
0: https://support.apple.com/guide/iphone/lift-a-subject-from-t...
1: https://developer.apple.com/videos/play/wwdc2023/10176/
EDIT: Try replacing the "Extract text" action with "Remove background". When running the shortcut, use "-o" to specify output image filename.
shortcuts run remove-background -i ~/Downloads/portrait-beard.avif -o beard.jpgI’ve always had good results from the Preview.app. I wonder how this engine compares for number of errors in a difficult source versus Free alternatives.
> we already have a common, portable data format for social media. It's screenshots of tweets
I know that I'd definitely use it!
However, when creating a PDF from images using Preview and exporting using ‘Embed Text’ option to OCR, I have noticed the text is worse than if you OCR the exact same images using the shortcut above or using a script. Presumably Preview is using the Vision framework’s less accurate fast path when preparing the PDF. https://developer.apple.com/documentation/vision/recognizing...
ocr-text "$1" && pbpaste
shortcuts run ocr-text -i new-haven-pizza.jpg | catI went back into my shortcut and Shortcuts added a pseudo-action "Stop and output <copy to clipboard>; if there's nowhere to output: <Do Nothing>", and I would think that "Do Nothing" would mean don't create a file, but I guess Quick Actions has some kind of special meaning given that all the other ones seem to be intransitive actions, implying that the user wants a file as the output.
Error: The operation couldn’t be completed. (WFBackgroundShortcutRunnerErrorDomain error 1.)
I get that some people will want to create it from scratch themselves or incorporate the actual meat of it into a larger shortcut... but not sharing one that does what the article says, because of a bug 2 years ago, is a bit of a weird take.
why can't shortcuts be exported as ... shortcut files?
it's not ideal to have people recreate the shortcut step by step (which is what I ended up describing in my post) but... I couldn't find a better way..! :-)
if you'd be able to recreate the shortcut and share it, and post the link here (and/or email it to me), I'd love to place that in the blog article! thank you
I'll try it again on macOS when I'm back at my desk.
Edit: also works on macOS Sonoma (https://www.icloud.com/shortcuts/6216aa9072144846adcaae69a5a...) - this one has all input sources selected, the iOS created one has only images/media/pdfs/files/rich text selected for input.
shortcuts run ocr-text -i <A PATH TO SOME IMAGE> | say -v Fred
`aichat -f tmp/test.png -- output only text in the image`
- Press CMD+SHIFT+4
- Draw square on screen where you want to extract the text from
- (Quickly) click on the preview image in the lower right corner
- Copy text from image
I scanned about 100 A4 documents in just a couple of minutes.