Show HN: I wrote a free Mac app to OCR any text on screen
github.com
github.com
This type of functionality should be integrated or readily available in MacOS.
I would love to have a way to do basic math operations, unit conversions, etc. without resorting to write them in spotlight. For example hovering over a price, it should convert it to a different currency. Or compare similar types of informations from different sources.
Your solution goes in this direction, thanks!
It should be pretty trivial for someone to hook this into https://insect.sh/.
We had this fully client-side in the browser all the way back in 2013.
Another alternative is Copyfish. It is cross-platform and uses cloud ocr:
The cloud == someone else's computer. Never forget that.
> The cloud == someone else's computer.
I think everyone around here knows this and can make their own decisions, if and for what data they use "the cloud".
If you take any screenshot to clipboard that includes some text, then paste it into Notes, it will silently name the resulting image file using the text pictured.
Not terribly useful but I did find it helpful once while taking screenshots for documentation.
And the docs says >By default, a text recognition request first locates all possible glyphs or characters in the input image, then analyzes each string.
Since the code doesn't specify any preferred languages, I think it would try to detect any languages supported by the framework.
From short googling, I found this thread [3]. Looks like the supported languages depends on the MacOS version, and it only support en, fr, it, de, es, pt, zh on Big Sur.
Not sure about the rotation though.
[1] https://developer.apple.com/documentation/vision/vnrecognize...
[2] https://github.com/schappim/macOCR/blob/master/ocr/main.swif...
It also has some other snipping modes, supports more than english, and has the option to, after the OCR, immediately show a popup window where you can fix what the OCR inevitably failed to properly recognize.
I recommend use alongside a clipboard manager.
[0] http://capture2text.sourceforge.net
P.S. I recommend changing/disabling its Win E shortcut as that conflicts with Windows built-in shortcut for file explorer
The Mac program uses VisionKit which does handle these cases better (not as well as google cloud vision from my outdated experience but way better than tesseract)
I tweeted a GIF on how it looks like https://twitter.com/cheeaun/status/1395973544983425025.
I’m using (and I’d like to recommend) https://www.keyboardmaestro.com/ for this.
It requires self-written macro, however it can do much more than that, including parsing & formatting OCRed text. For one job I went directly Image->OCR->File so I could copy OCRs into text for some non-elegant hardcodes ;-)
https://gist.github.com/jerieljan/f86843388c4ee9ecbc44d687a3...
“TextSniper is an easy-to-use desktop Mac OCR app that can extract and recognize any non-searchable and non-editable text on your Mac's screen. As an extra feature, it can turn OCR text into speech. It is a super convenient alternative to complicated optical character recognition tools.”
https://apps.apple.com/us/app/textsniper-ocr-simplified/id15...
“Meet lightning-fast text recognition on Mac. TextSniper is an app that can extract text from a selected portion of your screen. Forget taking notes — get TextSniper to capture and save what’s important.”
https://setapp.com/apps/textsniper
Consolidating into this note, OwlOCR is mentioned elsewhere in this post:
“Capture any text on your Mac's screen. Digitize images and PDFs to searchable PDFs using OCR right on your Mac.”
“OwlOCR allows grabbing a part of the screen and having any text in that area be instantaneously recognized and copied to clipboard. Additionally, the application supports recognizing text from PDF files, images and converting the contents to plain text. All conversion is done securely on-device - none of your images or files are sent to third-party services in the cloud.”
https://apps.apple.com/us/app/id1499181666
// Both process on device, OwlOCR mentions Apple's algo. Users of both in this thread are happy.
For example using tesseract on Linux and trying to OCR the "terminus" font I get better result by first resizing the screenshot to something bigger (and blurry) and even then it's far from perfect OCR'ing. When in the first place it's a pixel perfect font...
(and, yes, there are cases where OCR'ing screen fonts make sense)
The OCR software would then just need to be smart enough to recognize “things that look like glyphs”, and put bounding boxes around them; and everything from there could be implemented in logic, rather than a model. (Just apply the same transforms to the thing in the bounding box, and then search the fingerprint DB.)
/usr/local/bin/ocr | say
I've been successfully using Mathpix Snip [1] to do general OCR for quite some time.
It's not as well communicated as its initial purpose of applying OCR to LaTeX equations, but it currently supports much more than that, such as mixed text/math and tables.
On a personal note, I'm actually surprised it wasn't mentioned thus far in this thread.
See more discussion here on HN in [2], [3].
Also, does anyone know of a similar project for Linux?
tmp=/tmp/out
maim -s -u | tesseract - "$tmp"
# Remove empty lines
sed -ir '/^\s*$/d' "$tmp".txt
copyq add "$(cat "$tmp".txt)"
rm "$tmp".txt
rm "$tmp".txtr
tesseract insists on adding on txt extension and what I assume is some intermediary file txtr, making it awkward to use with mktemp. Probably explained in the manual which I skipped.But like others have said tesseract is not very reliable, at least with default settings -- it's common for it to add extra spaces or various single quotes, or omit spaces.
Abutting that "r" option against the "i" option is likely why you ended up with a file named .txtr and therefore implies that it did not actually hear the "-r" you intended
I've had the best luck picking an actual backup suffix such as "-i.bak" or "-i~" to keep BSD sed and GNU sed on the same page, although I've also seen scripts that go as far as "--version" sniffing and changing the actual invocation as "${SED_I} -E" type stuff
Woe unto those who write scripts as "sed -i -e /whatever/" since for half(?) of their users they'll end up with "somefile-e"
dyld: Symbol not found: _OBJC_CLASS_$_VNRecognizeTextRequest
Any chance to get it to work on poor macOS 10.14?
Any idea on how this compares to tessaract (or other local OCR).
I currently have an Alfred workflow that invokes tessaract and it works decently well but the accuracy could be better.
By far the best for text was Google Vision and then Azure. Whilst Google Cloud and Azure both also do handwriting recognition, Azure did better at this.
The cloud platforms performed better than pure on device with Apple’s vision API outperforming Tesseract.
Do you have the source for those workflows/would you be willing to share them?
I'll keep an eye on this one too though!
/* Begin PBXBuildFile section / 0425D1C16E9B7E34F8EBCCFB229F6BCF / Pods-ocr-umbrella.h in Headers / = {isa = PBXBuildFile; fileRef = E52F12A9CD9DA185DB6C7CFAF9971233 / Pods-ocr-umbrella.h /; settings = {ATTRIBUTES = (Project, ); }; }; 69F017594F16B64B4E70E96B863F38D1 / Pods-ocr-dummy.m in Sources / = {isa = PBXBuildFile; fileRef = 812D67335813B22DFC54237ACEB07CC8 / Pods-ocr-dummy.m /; }; 9D8F5FD727B32865EE80BA6ACDA12AF4 / ScreenCapture.swift in Sources / = {isa = PBXBuildFile; fileRef = CAE82544998B753F1708876308FF330D / ScreenCapture.swift /; }; AC8C4224C366FAD03EFFDC427D793373 / ScreenCapture-dummy.m in Sources / = {isa = PBXBuildFile; fileRef = 2BFFD24873C787E751AFC41D8C497ECB / ScreenCapture-dummy.m /; }; BE8E791706F107976678CAA1DE681FA6 / ScreenCapture-umbrella.h in Headers / = {isa = PBXBuildFile; fileRef = B9D6CB7E3F7CD4599F66F1F010D4CADD / ScreenCapture-umbrella.h /; settings = {ATTRIBUTES = (Project, ); }; }; CA9117D8B1C22828347BFE8326E2F7D2 / ScreenRecorder.swift in Sources / = {isa = PBXBuildFile; fileRef = FC34AC3B539E1EFA3B0D1E086E1BA1D9 / ScreenRecorder.swift /; }; / End PBXBuildFile section */
You can achieve the same thing using Tesseract on Linux, or even better quality using Google Vision.
[1] https://files.littlebird.com.au/Shared-Image-2021-05-22-17-1...
Maybe the issue your complaint unexpectedly tries to surface is that many awesome, highly useful projects like this one depend on code that isn't human-readable. I think this is a noteworthy point and should be discussed more often.
But then again, having the code in some form of source control /at all/ is far, far better than the alternative, which is depending on some instructions in a README.md or just hoping the user will know how to use XCode properly such that the real contribution of the project is used
Maybe your post also somewhat points out the fact that to newcomers or people looking at XCode code (auto-generated or otherwise) for the first time, it's /not/ obvious which files you should be looking at, and so we should give the parent poster some slack. Is this a problem that projects should worry about or take into consideration when auto-generated code starts to mix with non-generated code in source control?
N.B.: the parent post was talking about ./Pods/Pods.xcodeproj/project.pbxproj https://github.com/schappim/macOCR/blob/ca9a6379e07a8e1a5eaa...
The actual code here that isn't just Xcode boilerplate is in this very simple to read file: https://github.com/schappim/macOCR/blob/master/ocr/main.swif...
This application calls the native macOS libraries to do the OCR so I don't think you'd find anything useful here to do a Linux port - you can certainly use the idea and combine it with a linux compatible OCR library though.