> Resistant to Optical Character Recognition (OCR), most laypeople will need to print+rescan to OCR
If print+rescan (or equivalently, screen-grab+OCR) works, which it will, then it's hardly OCR-resistant!
The only thing this "blocks" is text extraction from the PDF with things like copy/paste or pdftotext/html/whatever conversion tools, which will "see" the codepoints used rather than the glyph images.
I'm 1000% sure there are gurus who could whip up a script to overcome this. But its kind of one of those things where you don't have to outrun the bear, you have to outrun your friend running next to you. It makes your sensitive documents just that much less likely to be scanned/found.
1) As many people pointed out, this doesn't prevent OCR, it just prevents copying strings (e.g. with crawlers). 2) Majority of OCR doesn't deal with PDFs produced from a text source but either from a) jpg-scans of documents b) pdfs produced from those jpg-scans. 3) The first thing I tried, was OCR with my iPhone and it obviously worked. As someone else said, there're solutions that let you batch process many documents.
Don't get me wrong, your stuff works for what you designed it to. However, it provides <false sense of security> by <falsely> claiming that it prevents OCR; which in turn, can lead to more harm[1].
[1] - e.g., it may convince people to share stuff that they wouldn't otherwise.
Retrieving text from images is literally the definition of OCR.
This seems like quite the oversight to me...
If you’re a divorce attorney who used this to convert documents in response to a discovery request, and the opposing side had a valid reason for needing the unobfuscated text, then you’re probably going to end up having a nice conversation with the judge about acceptable formats.
Sending compressed TIFFs would probably be just as good. A bit larger file sizes, but it would be just as effective as stopping automated scraping of text. Also, less likely to piss off a judge. Any opposing firm that would be sophisticated enough to automate scrapping the text from a normal PDF would be able to OCR these files just as easily.
Or maybe you have a second site that sells the decoder, so you get to sell to both sides. Not a bad business model, if you can work it.
... And most attorneys simply print documents. Once the PDF is on paper, OCR-ing it back into text is just one scanner away.
You don't need any scripts, just Acrobat itself (or any comparable PDF viewer) can do this. Export the PDF to images, make a new PDF out of the images, scan the text, done.
Example (took your example and did just that with it, now everything can be copied & pasted as normal text): https://filebin.net/qse2e0oaqkl1hjof/ocred.pdf
In general, if it LOOKS like text, SOMETHING can OCR it. That's the whole point of OCR. If you want to try to block OCR, you need something like CAPTCHAs, and that's getting less and less effective every day. In fact many are already more easily solved by computers than humans.
There is no form of OCR this is resistant to, simply the change the description to be accurate and remove references to being OCR resistant as this is false.
Security through obscurity is stupid:
gs -sDEVICE=tiffscaled24 -dNOPAUSE -dBATCH -dSAFER \
-sOutputFile=filename.tiff \
filename.pdf
tesseract filename.tiff filename.txt
All you need is ghostscript and tesseract. Both are an apt-get away.Any image you open in Safari, Preview, etc. (official Apple programs) will be OCR'd automatically, allowing you to extract the text with copy+paste. I think it works with PDFs, but I haven't tested it.
https://support.apple.com/guide/preview/interact-with-text-i...