Pdfsandwich
tobias-elze.de
tobias-elze.de
Once you have text in the PDF, you can use any sort of text analysis tools. You can use tools to convert it to plain text and grep through, or anything else you want.
That being said, it's not perfect, but still pretty awesome. Sometimes the spacing was off or it would confuse symbols like 1, I, or l. But these are minor and usually only on poorly scanned PDFs.
I have e.g. a directory with all weekly lecture slides for one lecture, and can directly find where (both file and page) we learned something related to photosynthesis via `rga photoshynthesis`.
Then, they put a warning on their site saying that they explicitly consider it to be misconduct search the book (e.g., by making the book searchable with OCR):
> The NYLC/NYLE Course Materials are locked in a non-searchable format in accordance with the Board’s misconduct rule prohibiting candidates from electronically searching the Course Materials when taking the NYLE. If a candidate, because of a disability, uses a screen reader to access written material, please contact the Board office by phone (518-453-5990), mail, or fax.
Kind of silly, sure! They don't want you searching the book during the exam, but they're fine with you going through it. And, following passing the bar exam, there is a character and fitness process, so students are fairly terrified of doing anything unethical, particularly things the bar explicitly says is unethical. So, it's basically an honor system, but with a big stick (although I haven't heard of any actual enforcement). If you OCR it, brag about it to your friends, and your friends really hate you, I guess they could report it.
[1]https://www.nybarexam.org/Content/NewYorkCourseMaterials.pdf
During covid, I stopped having my accountant come to my home office to do this work. After a great deal of experimentation, I resorted to purchasing an expensive Epson (approximately $600) document scanner that reliably does full duplex scans very quickly. The documents are stamped with a serial number (e.g. 2021-207) and stored in a single directory on my system with a file name that corresponds to the serial number (i.e. 2012-207.pdf).
The Epson hardware is great, the software is just barely adequate. It processes one document at a time, and the Epson application's OCR running on a M1 Mac mini can't keep up with the scanner and slows the whole process down somewhat. I would like to batch process the scanned documents in the background to convert the pdfs generated by the scanner into "searchable" pdfs. (Pdf files produced by scanners have an image layer but no text layer underneath it. Optional OCR done post-scan then adds the text layer.) I've tried a number of other OCR applications, one of the best is Adobe Acrobat Pro; it has the Adobe Pro kind of price, unfortunately, and does a million things I don't need.
Back to my filing system. I keep each years physical documents in one file drawer sorted by serial number. Because the documents are stamped with a serial number before being scanned, I can always find the physical document easily if I am looking at the pdf. Furthermore, because my pdf's are searchable I can quickly locate a bill or a tax document by a relevant name or even a particular amount (like, where did this $1,808.17 discrepancy come from).
Is this perfect, no far from it. Many little irritations afflict the actual process. Many statements have large amounts of small barely readable disclosures and footnotes, sometimes in faint small fonts. This is largely useless, slows down the OCR, and increases the file sizes. Barcodes, QRCodes, and DataMatrix codes often appear on the first page of these documents or even every page of the documents. It would be great if these were somehow scanned and used to tag the documents. The Epson software insists on embedding spaces in the generated file names, doesn't allow me to use auto generated ISO dates in file names, that makes working with the files from the command line less than ideal. (File names like "2021-Aug-07 122.pdf" are user friendly but not some friendly for scripts or sort commands.) I use several different configurations for the scanner, the software supports it, but I have to pay attention to pick landscape and double-sided when needed.
Thank you HN for the many suggested solutions to the OCR issues. It gives my hope that I'll be able to wire together something better than I've got now.
That said, pdfsandwich's 'one thing well' approach does have an appeal. I will definitely try it out, thanks for posting. Something in its "logo" reminded me of the OpenBSD fish. :)
It would allow you to take a physically scanned document and create a PDF with selectable text you could copy+paste, search over, etc.
the text will be added to each page invisibly "behind" the images.
ImageMagick to convert the pages to images
Tesseract-ocr by Google to transcribe the text in the images, which puts it’s output into singular pdf files
Pdfunite to stitch together the pdfs back into a whole file
I’m sure I’m missing a few, iirc it can call a tool that straightens the pages as well.
EDIT: Messed around and remembered the stuff:
where a.pdf is a 2 page PDF:
>convert a.pdf a.png
makes a-0.png and a-1.png
OCR's each image:
>for x in {0..1} ; do tesseract a-$x.png a_ocr-$x PDF ; done ;
combines them into 1 PDF:
>pdfunite a_ocr-{0..1}.pdf a_ocr_combined.pdf
The input is a scanned PDF. The output is the same PDF with the recognized text on top, in a transparent font.
Copy and paste now works because when you click the PDF you are selecting the transparent text.
But yeah, agree with your overall point, PDF is currently the easiest, most well-supported way of sending stuff that doesn't fudge around with the design. What you send someone is almost guaranteed to be what they see, unless they use some weird PDF reader.
How often do people embed fonts in the html via a font-face + base64 ttf or woff data-uri?
Not to mention need to save file in original Adobe Reader ("do you really want to overwrite this file?') every time you add a comment.
It converts documents to just images, then converts those back to PDF, all in a sandbox.
I work in a soewhat big corp. and PDF is viewed as the basis for the digitization of existing documents, be it for archiving or for the exchange of data with business partners. Also, we still have a fax machine in the company.
But, this is not just the opinion in my company. The public service also envisions sending and receiving digital documents using PDF as the future. E.g.: at the end of my digital tax return everything is summarized in a pdf (great), which then has to be signed and sent by post (lol).
many OCR software can scan this but not "understand" it to use it. ABBYY has something like this for scanning invoices so is there something for the foss folks?
The project I had in mind was similar to this one but I can't remember the name currently: https://github.com/tabulapdf/tabula
However, if you're looking for a ML-based, invoice-specific project looks like the other comment to your reply might be more useful.
> Popular examples
> Steak
I love it
Currently pdf is not supported, but I've considered adding that in.
https://kebekus.gitlab.io/scantools/image2pdf/
Also worth a look, works great and also supports PDF/A and jbig for a nice compression of b/w scans
The source code seems to be on sourceforge.net. I site once important, but now when I see it I either think "the project is most-likely dead" or "can this project be legit? Am I getting malware here?"
This is such a sad thing for me. When I was a kid, Sourceforge.net is where I would always go first to look for software, because I knew that it was a reputable host and that open-source projects participated in a culture of greater respect for users than most freeware projects demonstrate.
Not saying anything from 2018 isn't valuable... I'm currently working on PDF scraping tools, and my lord, the stuff we're left to work with is abysmal... this could be state of the art.
Filezilla, for example, initially participated in SourceForge's adware program under the old owner, but after the owner change, actually provided a clean version on Sourceforge (while the version on the website was and still is adware-bundled).
there is something called scantailor if you are scanning books yourself. that gives you more contrl over the orientation and margins and ocr and contrast controls. that said, pdfsandwich gives you a miniscule file size which i could not achieve otherwise.