What's so hard about PDF text extraction?
filingdb.com
filingdb.com
(( if your answer is no because this is part of your secret sauce, I totally understand ))
The real question is how did such a torturous (once proprietary) format manage to become the only ubiquitous standard for paginated, self-contained, formatted text suitable for printing. In principle HTML could actually do the job, but arguably it's not the ideal tool either, and I've not seen an example of an HTML document that just prints as expected across multiple pages either.
Odd that this got posted almost the same time as https://news.ycombinator.com/item?id=33145498 btw...
It's a hard enough domain area that warrants a startup to solve the pain point.
If you want an editable document, use a word document instead.
HTML does have page-break-before and page-break-after css properties, but css could do with a "scale element to fit page", "only this element on a page" and "center on page" css properties. Even the print specific css properties that do exist currently are not well supported by chrome. Google really needs to fix that situation IMHO.
* sure, docx may be a form of OpenXML but it's still largely based around MS Word's proprietary behaviour - very few other tools can reliably display all docx files well.
From what I remember the PDF standards groups are currently investigating adding something just for temper detection.
And many certificate authorities sell such certificates but for high prices.
Typing in the name or signing with a touch screen is actually considered a valid and binding, official signature in most countries. It is the lowest form of a signature equivalent to a written signature.
Agreed on cumbersome, I have a PNG file with transparency containing my signature and use Xournal++ (https://xournalpp.github.io/) for placing it on the PDFs. That works fine but still takes more time than I want to spend. But it's faster than typing it out every time.
Anyone wanting to extract text from a pdf would do better to hand it to a pool of typists for them to retype into a new document.
Or better yet, edit the original document that the pdf was generated from.
My most effective (so far) extractive program is 'pdftotext' and using the '-layout' option:
pdftotext -layout my_file.pdf