What's so hard about PDF text extraction?
filingdb.com
filingdb.com
Edit: To add a little more color, given that none of us was (or at least certainly I wasn't) an expert on the PDF format, we had so far treated the bug like a bug of probably at-most moderate complexity (just have to read up on PDF and figure out what the base unit is or whatever). After discovering what this article talks about, it became evident that any solution we cobbled together in the time we had left would really just be signing up for an endless stream of it-doesn't-work-quite-right bugs. So, a feature that would become a bug emitter. I remember in particular considering one of the main use cases: scientific articles that are usually in two columns, AND also used justified text. A lot of times the spaces between words could be as large as the spaces between columns, so the statistical "grouping" of characters to try to identify the "macro rectangle" shape could get tricky without severely special-casing for this. All this being said, as the story should make clear, I put about one day of thought into this before the decision was made to avoid it for 1.0, so far all I know there are actually really good solutions to this. Even writing this now I am starting to think of fun ways to deal with this, but at the time, it was one of a huge list of things that needed to get done and had been underestimated in complexity.
The funny thing is that creating a universal algorithm to convert PDFs and/or HTML to plaintext is probably comparable in difficulty to building level 5 self-driving cars, and would accrue at least as much profit to any company that can solve it. But there are hundreds of billions of dollars going into self-driving cars, and like zero dollars going into this problem.
The basic issue imho is that NLP algorithms are very inaccurate even with perfect input. E.g. even with perfect input, they're maybe only 75% accurate. And even an a text-processing algorithm that's like 99.9% accurate will yield input to your NLP algorithms that's like 50% accurate, so any results will be mostly unusable.
- Better search engine results
- Identifying experts within a company
- Better machine translation
- Finding accounting fraud
- Automating legal processes
For context, the reason why Facebook is the most successful social network is that they're able to turn behavioral residue into content. If you can get better at taking garbage data and repackaging it into something useful, it stands to reason that there are lots of other companies the size of Facebook that can be created.
There's almost no question in my mind that most new data will endure in some form, by virtue of being digital from day 1.
The endgame for such a company, imho, is to become the "source entity" of information management (in abstracted form), whose two major products are one to express this information in the digital space, and the other in the analog/physical space. You may imagine variations of both (e.g. AR/VR for the former).
Kinda like language in the brain is "abstract" (A) (concept = pattern of neurons firing) and then speech "translates" into a given language, like English (B) or French (C) (different sets of neurons). So from A you easily go to either B or C or D... We've observed that Deep Learning actually does that for translation (there's a "new" "hidden" language in the neural net that expresses all human languages in a generic form of sorts, i.e. "A" in the above example).
The similarities of the ontology of language, and the ontology of information in a system (e.g. business) are remarkable — and what you want is really this fundamental object A, this abstract form which then generates all possible expressions of it (among which a little subset of ~1,000 gives you human languages, a mere 300 active iirc; and you might extend that into any formal language fitting the domain, like engineering math/physics, programming code, measurements/KPI, etc.
It's a daunting task for sure but doable because the space is highly finite (nothing like behavior for instance; and you make it finite through formalization, provided your first goal is to translate e.g. business knowledge, not Shakespeare). It's also a one-off thing because then you may just iterate (refine) or fork, if the basis is sound enough.
I know it all sounds sci-fi but having looked at the problem from many angles, I've seen the PoC for every step (notably linguistics software before neural nets was really interesting, producing topological graphs in n dimensions of concepts e.g. by association). I'm pretty sure that's the future paradigm of "information encoding" and subsequent decoding, expression.
It's just really big, like telling people in the 1950's that because of this IBM thing, eventually everybody will have to get up to speed like it's 1990 already. But some people "knew", as in seeing the "possible" and even "likely". These were the ones who went on to make those techs and products.
PDF is de-facto standard for any invoicing, POs, quotes, etc.
If you solve the problem you can effectively programmatically deal with invoicing/payments/ large parts of ordering/dispensing. It's a no brainer to add it on to almost any financial/procurement software that deals with inter business stuff.
Any small-medium physical business can probably half their financial department if you can dependably solve this issue.
Small-mediums should be looking to consolidate buying through a few good suppliers and working with them directly to automate process, or adopting interchange formats.
Problem for some small-business is the cost (process changes, licencing etc) of adopting interchange formats and working with large vendors is prohibitive at their scale e.g. the airline BSP system.
I agree that solving the problem generally i.e. replacing an accounts payable staff person capable of processing arbitrary invoice documents will be comparable to self-driving in difficulty.
If a company deals with a lot of a single type of PDF, then the approach could be economical. I am actually involved in a project looking at doing this with AWS Textract.
Building machines that understand formats that are understood by humans is exactly what we should be doing. People should read, write, and process information in a format that is comfortable and optimized to them. Machines should bend to us, we should not bend to them.
If businesses only dealt with machine readable formats, everyone's computer would still be using the command line.
And there's real condescension in your post:
> Small-mediums should be looking to consolidate buying through a few good suppliers and working with them directly to automate process
You're saying that businesses need to change their business to accommodate data formats, but it should be the other way around.
The proliferation of computers in business over the last 50 years is precisely because businesses can save money/expand capacity by adapting the business processes to the capabilities of the computers.
Over that time, computers have become more friendly to humans, but businesses have adapted and humans been trained to use what computers can do.
As far as I understand there are at least two standards (I know of in Germany): XRechnung and ZUGFeRD/Factur-X (which is PDF A/3 with embedded XML).
You can go further. Invoices often contain block sections of text with important terms of the invoice, such as shipping time information, insurance, warranties, etc. To build something that works universally, you also need very good natural language processing.
This should be obvious, but the answer is because OCR engines are not terribly accurate. If you have a native PDF, you're far better off parsing the PDF then converting to an image and OCRing. But if OCR ever becomes perfect, then sure.
While Abbyy is likely the best, it's also incredibly expensive. Roughly on the order of $0.01/page or maybe at best a tenth of that in high volume.
For comparison, I run a bunch of OCR servers using the open source tesseract library. The machine-time on one of the major cloud providers works out to roughly $0.01 for 100-1000 pages.
Changing their process might be more expensive than paying a lot of money for them to carry on as is for a few more years while getting the benefit of modern eyes on their content.
Edit: concrete example would be government publications like budget narrative documents.
Legal technology. Pretty much everything a lawyer submits to a court is in PDF, or is physically mailed and then scanned in as PDF. If you want to build any technology that understands the law, you have to understand PDFs.
PDFs are incredibly flexible. Text can be specified in a bunch of ways. Glyphs can be defined to the nth degree. Text sometimes isn’t text at all. There’s no layout engine and everything is absolutely positioned. Fonts in PDF’s are insane because they’re often subset so they only include the required glyphs and the characters are remapped back to 1, 2, 3 etc instead of usual ascii codes.
I've actually seen obfuscation used in a PDF where they load in a custom font that changes the character mapping, so the text you get out of the PDF is gibberish, but the fonts displayed on rendering are correct (a simple character substitution cipher).
The important thing to remember whenever you think something should be simple, is that someone somewhere has a business need for it to be more complicated, so you'll likely have to deal with that complication at some point.
If you care to check us out: https://siftrics.com/
Feel free to email me at siftrics@siftrics.com with any questions. We can setup a phone call, zoom meeting, or google hangouts too, if you’d like.
I've often thought about creating products like these but as a one-man operation I am daunted by the "getting customers" part of the endeavour. How do you get a product like this into the hands of people who make the decisions in a business? (For anyone, not just OP). PPC AdWords campaigns? Cold-calling? Networking your ass off? Pay someone? Basically, how does one solve the "discoverability problem"?
In the meantime, if you have any questions, feel free to send me an email at siftrics@siftrics.com. I’d love to hop on the phone or do a Zoom meeting or a Google Hangouts.
The page is clear and easy to understand, looks good. Well done.
I like to say that anyone with a good old ThinkPad and an internet connection can mint fortunes and build empires :-)
I am especially thinking about Japanese. Our company could probably find good uses of such service if it had Japanese support.
If you have any questions or need help trying it out, please email me at siftrics@siftrics.com. We can hop on the phone, too, if you'd like.
Of course, there is absolutely no semantics, just display.
I would believe that. It was a pretty poor obfuscation method as they go, if it was intended for that.
> Believe it or not, you can often rebuild the cmaps using information in the pdf to fix the mapping and make the extraction work again.
Oh, I did. That's the flip side of my second paragraph above. When there's a business need to work around complications or obfuscations, that will also happen. :)
Can't stress this enough. The next time you open a multi-column PDF in adobe reader and it selects a set of lines or a paragraph in the way you would expect, know that there is a huge amount of technology going on behind the scenes trying to figure out the start and end of each line and paragraph.
At least :)
Can you explain a bit more about why this is so valuable? I don't know anything about this industry.
There's a lot more than zero dollars going into this... it's just that the end result is universally something that's "good enough for this one use-case for this one company" and that's as far as it gets.
Converting PDFs to HTML well is a very hard problem, but hard by itself to create a very big company. When processing PDFs or documents generally, the value is not in the format, it's in the substantive content.
The real money is not going from PDF to HTML, but from HTML (or any doc format) into structured knowledge. There are plenty of companies trying to do this (including mine! www.docketalarm.com), and I agree it has the potential to be as big as self-driving cars. However, technology to understand human language and ideas is not nearly as well developed as technology to understand images, video, and radar (what self-driving care rely on).
The problem is much more difficult to solve than building safer-than-human self-driving cars. If you can build a machine that truly understands text, you have built a general AI.
Why wouldn't you just zoom with the center point being where the tap occurred?
- Use https://github.com/flexpaper/pdf2json to convert the PDF in an array of (x, y, text) tuples
- Use a good text parsing library. Regexes are probably not enough for your use case. In case you are not aware of the limitations of regexes you may want to learn about Chomsky hierarchy of formal languages.
Here is the section of our Dockerfile that builds pdf2json for those of you that might need it:
# Download and install pdf2json ARG PDF2JSON_VERSION=0.70 RUN mkdir -p $HOME/pdf2json-$PDF2JSON_VERSION \ && cd $HOME/pdf2json-$PDF2JSON_VERSION \ && wget -q https://github.com/flexpaper/pdf2json/releases/download/$PDF... \ && tar xzf pdf2json-$PDF2JSON_VERSION.tar.gz \ && ./configure > /dev/null 2>&1 \ && make > /dev/null 2>&1 \ && make install > /dev/null \ && rm -Rf $HOME/pdf2json-$PDF2JSON_VERSION \ && cd
Regexes have limitations but I was able them to leverage them sufficiently for PDFs from a single source.
I parsed over 1 million PDFs that had a fairly complex layout using Apache PDFBox and wrote about it here: https://www.robinhowlett.com/blog/2019/11/29/parsing-structu...
https://github.com/AXATechLab/pdf2json
Bounding box also can be off with pdf2json. Pdf.js do a better job but have a tendency to no handling some ligature/glyph well, transforming word like finish to "f nish" sometime (eating the i in this case). pdfminer (python) is the best solution yet but a thousand time slower....
[0] https://www.thoroughbreddailynews.com/getting-from-cease-and...
Most programming languages offer a regex engine capable of matching non-regular languages. I agree though, if you are actually trying to _parse_ text then a regex is not the right tool. It just depends on your use case.
I have used this in the past to extract tables, but it doesn't help much in cases where you need font size information.
The PDF standard is a mess, and the number of 'tricks' I've seen done is astonishing.
Example: to add shade or border effect to text, most PDF generators simple add the text twice with a subtle offset and different colors. Result: your SaaS service returns every sentence twice.
Off course there were workarounds, but at some point it became unmaintanable.
The demand for automating text extraction is still very high — or at least it feels like it when you’re working around the clock to cater to 3 of your customers, only to wake up to 10 more the next day. We’re small but growing extremely quickly.
It’s definitely harder to get government business because the sales process is so long and compliance is so stringent. That said, we are GDPR compliant.
I've bookmarked your site for future research... but the aviation part has me curious!
Also, how do you manage things when one of those banks decides to change the layout/format?
I managed to find your email address from your GitHub profile. Going to send you an old fashioned email.
pdftotext -layout -nopgbrk -eol unix -f $firstpage -l $lastpage -y 58 -x 0 -H 741 -W 596 "$FILE"
After that, it was a matter of writing and tweaking custom text parsers (in python or java) until the output was acceptable, generally an XML file consumed by the build (mainly to generate code).A frequent need was to parse tables describing fields (name, id, description, possible values etc.). Unfortunately, sometimes tables spanned several pages and the column width was different on every page, which made column splitting difficult. So I annotated page jumps with markers (e.g. some 'X' characters indicating where to cut).
As someone else said, this is like black magic, but kind of fun :)
Edit: grammar
I've come across most of the problems in this post but the most memorable thing was when we were asked to support Arabic, when suddenly all your previous assumptions are backwards!
See:
https://news.ycombinator.com/item?id=22156456
In the GNU Awk User's Guide:
https://www.gnu.org/software/gawk/manual/html_node/Multiple-...
Tracking column and field widths across page breaks is ... interesting, but more tractable.
Parsing statement PDFs from every bank is pretty hellish.
In the past we did purposely make it more difficult to parse our PDFs
The experience helped me to roll out an API, as https://extracttable.com, for developers.
OCR tricks? Assuming post processing dev stuff - may I know your OCR engine. We are supported with Kofax and openText along with cloud engines like GVision as a backup.
We need some metadata to rearrange and sort PDF pages for mailing and delivery (such as name, address, and start/end page for that customer).
Our general rule is you provide metadata in an external file to make it easy for us. Otherwise, we run pdftotext and hope there's a consistent formatting for the output (e.g. every first page has "Issue Date:", "Dear XYZ,", or something written on it).
If that doesn't work then we're re-negotiating. It is not too difficult usually to build a parser for one family of PDF files based on a common setup as you've said and you get to learn various tricks. It is very difficult though to write a general parser.
Personally, I found parsing postscript easier since usually it was presented linearly.
Good thing about this was as you have already outlined: It allowed for some flexibility in what was acceptable input data. For specific address formats or names we could accept multiple formats as long as they were consistent and in the proper position in the input file.
Regarding renegotiating: We didn't get that far. However, if a customer within our organization was enlisting our expertise and could not produce an acceptable input file, then we would go back to them and explain the format that we require in order to generate the necessary documents. Of course, creating our document through our data pipelines is obviously the better choice, but this was not an option in some cases at the time.
As far as doing the work of creating these documents in a tool like Planetpress is concerned, well, don't use Planetpress. You are better of doing it in your favorite language of choice's libraries tbh. Nothing worse than having to use proprietary code (Presstalk/Postscript.) that you have to learn and never be able to use anywhere.
The problem we have with a lot of client files is that they look fine but printers don't care about "look fine", they crash hard when they run out of virtual memory due to poor structure. And usually without a helpful error message, so that's more billable hours to diagnose. The most common culprit is workflows that develop single document PDFs then merge them resulting in thousands of similar and highly redundant subset fonts.
# Parser drift and maintenance hell
Let's say that you receive 100 invoices a month from a company over the course of 3 months. You look over a handful of examples, pick features that appear to be invariant, and determine your parsing approach. You build your parser. You're associating charges from tables with the sections their declared in, and possibly making some kind of classification to make sure everything is adding up right. It works for the example or two pdfs you were building against. It goes live.
You get the a call or bug report: it's not working. You try the new pdf they send you. It looks similar, but won't parse because it is--in fact--subtly different. It has a slightly different formatting of the phone-number on the cover page, but identical everywhere else. You change things to account for that. You retest your examples, they break. Ok, two different formats same month, same supplier. You fix it. Chekhov's Gun has been planted.
A month passes, it breaks. You inspect the offending pdf. Someone racked up enough charges they no longer fit on a page. You alter the parser to check the next page. Sometimes their name appears again, sometimes not, sometimes their next page is 300 pages away. It works again.
A few more months later, a sense of deja-vu starts to set it. Didn't I fix this already? You start tracking three pdfs across 3 months:
pdf 1 : a -> b -> c (Starts with format a, change to be same as pdf 2, then changes again)
pdf 2 : b -> b -> c (Starts with one format, stays the same, changes the same way as pdf 1)
pdf 3 : b -> a -> b (Starts same as pdf 2, changes to same as pdf 1 first month, same as pdf 3)
What's the common factor between these version changes? The return address is determining the version.
PDFs are slightly different from office to office, with templates drifting slightly each month in diverging directions. You have to start reevaluating parsing choices and splitting up parsers. It's difficult to account for incurring linear maintenance cost for each new supplier and amortize that over a sizeable period of time. My arch nemesis is an intern who got put to work fixing the invoices at one office of one foreign supplier.
# PDFs that aren't standards compliant
In this case, most pdf processing libraries will bail out. Pdf viewers on the other hand will silently ignore some corrupted or malformed data. I remember seeing one that would be consistently off by a single bit. Something like `\setfont !2` needed to have '!' swapped out for another syntactically valid character that would leave byte offsets for the pdf unchanged.
TLDR: If you can push back, push back. Take your data in any format other than PDF if there is any way that is possible.
I have used Google's fine OCR results to simulate a hacker.
- Download a youtube video that shows how to attack a server on the website hackthebox.eu
- Run ffmpeg to convert the video to images.
- Run a jpeg to pdf tool.
- Upload the pdf to google drive.
- Download the pdf from google drive.
- Grep for the command line identifiers "$" "#".
- Connect to hackthebox.eu vpn.
- Attack the same machine in the video.
ImageMagick. convert *.jpg out.pdf
By the way, why do you wait 10 minutes? Is there a signal that the PDF is done processing?
Or is there just some kind of voodoo magic that seems to happen that just takes 10 minutes to do?
Also, I am not aware of a signal when it is done.
Now I think about it, I don't know what you mean by "upload a pdf to google drive and download it 10 minutes later".
Uploading and downloading a file shouldn't change it at all, at bit level.
I'm not sure (I haven't thought about it a lot) that you could come up with a format that duplicates that function and is also easier to parse or edit.
It's not that the thing you're trying to do is stupid. It's probably entirely legitimate, and driven by a real need. It's just that the original designers of the thing you're trying to work on didn't give a damn about your ability to work on it.
QFT. PDF should really have been called “Print Description Format”. At heart it’s really just a long list of non-linear drawing instructions for plotting font glyphs; a sort of cut-down PostScript.
https://en.wikipedia.org/wiki/PostScript
(And, yes, I have done automated text extraction on raw PDF, via Python’s pdfminer. Even with library support, it is super nasty and brittle, and very document specific. Makes DOCX/XLSX parsing seem a walk in the park.)
What’s really annoying is that the PDF format is also extensible, which allows additional capabilities such as user-editable forms (XFDF) and Accessibility support.
https://www.adobe.com/accessibility/pdf/pdf-accessibility-ov...
Accessibility makes text content available as honest-to-goodness actual text, which is precisely what you want when doing text extraction. What’s good for disabled humans is good for machines too; who knew?
i.e. PDF format already offers the solution you seek. Yet you could probably count on the fingers of one hand the PDF generators that write Accessible PDF as standard.
(As for who’s to blame for that, I leave others to join up the dots.)
Currently, there is no viable alternative if you want the pros but not the cons
I remember OpenXPS being much easier to work with. That might be due to cultural rather than structural differences, mind - fewer applications generate OpenXPS, so there's fewer applications to generate them in their own special snowflake ways.
The problem with this is that from an average person perspective it doesn't have the pros. There is no built-in or first-party app that can open this format on Mac and Linux. More than 99% of the users only want to read or print it. It's hard to convince them to use an alternative format when it's way more difficult to do the only thing they want to do.
Once you have something as a PNG (or any other format you can get into a Bitmap), throwing it against something like System.Drawing in .NET(core) is trivial. Once you are in this domain, you can do literally anything you want with that PDF. Barcodes, images, sideways text, html, OpenGL-rendered scenes, etc. It's the least stressful way I can imagine dealing with PDFs. For final delivery, we recombine the images into a PDF that simply has these as scaled 1:1 to the document. No one can tell the difference between source and destination PDF unless they look at the file size on disk.
This approach is non-ideal if minimal document size is a concern and you can't deal with the PNG bloat compared to native PDF. It is also problematic if you would like to perform text extraction. We use this technique for documents that are ultimately printed, emailed to customers, or submitted to long-term storage systems (which currently get populated with scanned content anyways).
pdftk form.pdf multibackground additions.pdf output output.pdf
Sounds like you are in need of OCR if you want to be able to use arbitrary screen coords as a lookup constraint.
From my experience, it seems to grab text just fine, the tricky part is identifying & grabbing what you want, and ignoring what you don't want... (for reasons mentioned in the article)
https://github.com/itext/itext7-dotnet
https://itextpdf.com/en/resources/examples/itext-7/parsing-p...
Not even when they try to select and copy text?
Epubs in comparison are easy, as all it takes is a single tap or button press to continue. When there's no DRM on the file (thanks HB, Baen) I read in FBReader with custom fonts, colors, and text size. It doesn't hurt any that the epub files I get are usually smaller than the PDF version of the same book.
Personally, I think the fact that Calibre's format converter has so many device templates for PDF conversion says a lot.
[1] recently read an SEO post on okta's site. who can read that garbage?
[2] only GA ... which isn't a 3rd-party tracker.
Why not? It's not self-hosted and results are stored elsewhere.
some might add: for the purpose of resale of the data, but I don't think that's a requirement to be classified as 3rd party tracker. the mere act of correlation, no matter what you then do with the data, makes you a 3rd party tracker. in case you think that's just semantics, this is important for GDPR and the new california law.
you can turn on the "doubleclick" option, which does do said correlation and tracks you. but that's up to the site to decide. GA doesn't do it by default.
As noted in the article, it is extremely difficult to figure out the original text given only a "normal" PDF, so you end up using a lot of heuristics that sometimes guess correctly. There's no guarantee that you'll be able to extract the "original text" when you start with an arbitrary PDF without embedded data. So if you're extracting text, neither way guarantees that you'll get "original text" that exactly matches the displayed PDF if an attacker created the PDF.
That said, there's more you can do if you have an embedded OpenDocument file. For example, you could OCR the displayed PDF, and then show the differences with the embedded file. In some cases you could even regenerate the displayed PDF & do a comparison. There are lots of advantages when you have the embedded data.
When you're done filling the form, the PDF runs form validity checks and generates a 2D barcode [1] -- which stores your all field entry data -- on the first page. This 2D barcode can then be digitally extracted on the receiving end with either a 2D barcode scanner or a computer algorithm. No loss of fidelity.
Looks like Acrobat supports generation of QR, PDF417 and Data Matrix 2D barcodes.[2]
[1] https://www.canada.ca/en/revenue-agency/services/tax/busines...
[2] https://helpx.adobe.com/acrobat/using/pdf-barcode-form-field...
The Canadian tax agency offers free storage for whatever receipts you mail them? Sounds nifty. Does the IRS (or any other tax agency) do this?
Is there a special "data" section of the PDF that includes this? Can you point me to any documentation regarding this? It sounds quite good TBH.
Using white-on-white dark-hat SEO techniques for keyword boosting? Check. Custom fonts with random glyphs? Check. I didn't see custom encodings (yet).
We try to keep HTML semantic, but google has been interpreting pages to a much higher level in order to spot issues such as these. If you ever tried to work on a scraper, you know how it's very hard to get far nowdays without using a full-blown browser as a backend.
What worries me is that it's going to get massively worse. Despite me hating HTML/web interfaces, one big advantage for me is that everything which looks like text is normally selectable, as opposed to a standard native widget which isn't. It's just much more "usable", as a user, because everything you see can be manipulated.
We've seen already asm.js-based dynamic text layout inspired by tex with canvas rendering that has no selectable content and/or suffers from all the OP issues! Now, make it fast and popular with WASM...
"yay"
Though absolute-positioning of all text elements via CSS at some arbitrary level (I've seen it by paragraph), such that source order has no relationship to display order, is quite close.
I started reading the ISO spec for postscript used in modern PDFs. You can read it yourself here: https://www.adobe.com/content/dam/acom/en/devnet/pdf/pdfs/PD...
What actually needs to be done to extract text correctly is to be able to parse the postscript, have a way of figuring out how the raw text.. or the curves that draw the text.. are displayed (whether they are or not and in relation to each other) using information that the postscript gives you.
Edit: More than anything I think understanding deeply the class of PDFs you want to extract data from is the most important part. Trying to generalize it is where the real difficulty comes from.. as in most things.
I didnot get time to enhance it further but planning to containerize the whole application. See if you find it useful in its current form.
Not a fan of the potential vendor lock in though, so it's only really suitable for those in an already AWS environment not worried about them harvesting your data.
* Polish ebooks, which usually use Watermarks instead of DRM, sometimes hide their watermarks in a weird way the screen reader doesn't detect. Imagine hearing "This copy belongs to address at example dot com: one one three a f six nine c c" at the end of every page. Of course the hex string is usually much longer, about 32 chars long or so.
* Some tests I had to take included automatically generated alt texts for their images. The alt text contained full paths to the JPG files on the designer's hard drive. For example, there was one exercise where we were supposed to identify a building. Normally, it would be completely inaccessible, but the alt was something like "C:\Documents and Settings\Aneczka\Confidential\tests 20xx\history\colosseum.jpg".
* My German textbook had a few conversations between authors or editors in random places. They weren't visible, but my screen reader still could read them. I guess they used the PDF or Indesign project files themselves as a dirty workaround for the lack of a chat / notetaking app, kind of like programmers sometimes do with comments. They probably thought they were the only ones that will ever read them. They were mostly right, as the file was meant for printing, and I was probably the only one who managed to get an electronic copy.
* Some big companies, mostly carriers, sometimes give you contract templates. They let you familiarize yourself with the terms before you decide to sign, in which case they ask you for all the necessary personal info and give you a real contract. Sometimes though, they're quite lazy, and the template contracts are actually real contracts. The personal data of people that they were meant for is all there, just visually covered, usually by making it white on white, or by putting a rectangle object that covers them. Of course, for a screen reader, this makes no difference, and the data are still there.
Similar issues happen on websites, mostly with cookie banners, which are supposed to cover the whole site and make it impossible to use before closing. However, for a screen reader, they sometimes appear at the very beginning or end of the page, and interacting with the site is possible without even realizing they're there.
With that in consideration, and the existing resources are little help especially on skewed, blurry, handwritten and 2 different table structure in the input, I ended up creating an API service to extract tabular data from Images and PDFs - hosted as https://extracttable.com . We cared it to be robust, average extraction time on images is under 5 seconds. On top of maintaining accuracy, A bad extraction is eligible for credit usage refund, which literally not any service offer it.
i Invite HN users to give it a try and feel free to email saradhi@extracttable.com for extra API credits for the trail.
You're "handwritten" example looks a bit "too decent" as well. I can see how that works. You first look for the edges of the table, and then you evaluate the symbol in each cell as something that matches unicode.
So, how well does this cope with increasing degradation? i.e. pencil written notes that bleed outside cell borders, curve around borders, etc.? Stamps and symbols (watermarks) across tables?
"The Worst Image" is a close match to that, except it is a print.
Regarding increasing degradation - as stated above, the OCR engine is not proprietary - we confined ourselves to detect the structure at this moment, and started with the most common problems.
Feel free to reachme at manuel at jazzido dot com
Am I reading the repos correctly? It looks like Extractable copied Tabula (MIT) to its own repo rather than forking it, removed the attribution, and then tried to re-license it as Apache 2.0. If so, that would be pretty fucked up.
Still, I would have loved at least a heads up from the team that sells Tabula Pro. I know they're not required to do so, but hey, they're kinda piggybacking on Tabula's "reputation".
(IANAL)
What do you recommend us to do, to not make you feel we made a dick move.
TIA
Did you ask permission of the original author to use a derived name?
Did you discuss your plan to commercialize the original author's work with the author? Before starting out?
Since starting a commercial project, how much money have you given to the original author?
1) No. 2) No. 3) Zero.
"commercialize the original author's work with the author" - No, but let me highlight this, any extraction with tabula-py is not commercialized - you can look into the wrapper too :) or even compare the results with tabula-py vs tabulaPro.
Copying the TabulaPro description here, "TabulaPro is a layer on tabula-py library to extract tables from Scan PDFs and Images." - we respect every effort of the contributors & author, never intended to plagiarize.
I understand the misinterpretation here is that we are charging for the open-sourced library because of the name. We already informed author in the email about unpublishing the library, this morning, I just deleted the project and came here to mention it is deleted :)
It may be that you didn't realize that people would see your appropriation as wrong, although I have a hard time believing that as well given that the author tried to contact you and was ignored. As they say, "The wicked flee when no man pursueth."
So what I see here is somebody knowingly doing something dodgy and then panicking when getting caught. If you'd really like to make amends, I'd start with some serious introspection on what you actually did, and an honest conversation with the original author that hopefully includes a proper [1] apology.
[1] Meaning it includes an explicit recognition of your error and the harms done, a clear expression of regret, and a sincere offer to make amends. E.g., https://greatergood.berkeley.edu/article/item/the_three_part...
I would like to get the author's comment on "tried to contact you and was ignored", as I was the one who emailed yesterday.
And meanwhile if you say on HN that HTML should be used instead of PDF for papers, people will jump on you insisting that they need PDF for precise formatting of their papers—which mostly barely differ from Markdown by having two columns and formulas. What exactly they need ‘precise formatting’ for, and why it can't be solved with MathML and image fallback, they can't say.
People feeling the urge to defend PDF might want to pick up at this point in the discussion: https://news.ycombinator.com/item?id=21454636
We need the opposite, we need a format that stays the same size, same proportions and is vectorized so you can zoom to any size - however, the relationship of space between elements remains constant.
PDF is an amazing format IMO. Think of it like Docker - the designer knows exactly how its going to appear on the user's device.
If you have opposing thoughts, please elaborate further instead of simply saying "nothing to do with it". HN works by explaining and arguing about issues to get to the bottom of something.
Sorta like this then, more rigidly conveying the designer's intended layout: https://bureau.rocks/projects/book-typography-en/ ?
(Though these ‘books’ are almost the opposite of what I'm advocating for, in terms of formatting.)
They key difference is it contains no ambiguity (such as all fonts must be embedded).
This conjecture would have some practical relevance if I had access to the same papers in other formats, preferably HTML. Yet I'm saddened time and again to find that I don't.
In fact, producing HTML or PDF from the same source was exactly my proposed route before I was told that apparently Tex is only good for printing or PDFs. I hope that this is false, but not in a position to argue currently.
It is worrying if places that are “libraries” of knowledge aren’t taking the opportunity to keep searchable/parseable data, but it’s no worse than a library of books.
That's not my complaint in the first place. The problem is that while we progressed beyond books on the device side in terms of even just the viewport, we seemingly can't move past the letter-sized paged format. The format may be a bit better than books—what with it being easily distributed and with occasionally copyable text—but not enough so.
I'm not even touching the topic of info extraction here, since it's pretty hard on its own and despite it also being better with HTML.
PDF is simply better than HTML when it comes to preserving layout as the author intended.
First, there's no need to stretch my argument to the point of it being ridiculous. I don't have to reach for a watch to suffer from PDF. Even a tablet is enough: I don't see many 14" tablets flying off the shelves. I also know for sure that the vast majority of ubiquitous communicator devices, aka smartphones, are about the same size as mine, so everyone with those is guaranteed to have the same shitty experience with papers on their communicators and will have to sedentary-lifestyle their ass off in front of a big display for barely any reason.
Secondly:
> Likewise, when layout makes the difference between reader understanding or confusion, it's hard to trust automatic reflowing on unknown screen sizes. PDF is simply better than HTML when it comes to preserving layout as the author intended.
As I wrote right there above, still no explanation of why the layout makes that difference and why preserving it is so important when most papers are just walls of text + images + some formulas. Somehow I'm able to read those very things off HTML pages just fine.
I think it's the artist in them.
This is understandable in some instances:
a. Picasso's or Monet's works probably wouldn't be as good if you just roll them up into a ball. Sure, the component parts are still there (it's just paper/canvas and paint after all!) but the result isn't what they intended.
b. A car that has hit a tree is made up of the composite parts but isn't quite as appealing (or useful) as the car before hitting the tree.
c. A wedding cake doesn't look as good if the ingredients are just thrown over the wedding party's table. The ingredients are there, but it just isn't the same...
Presentation is sometimes important.
You're right that most of the relevant semantics would fit into Markdown. So store the markdown! There are problems with PDF but HTML is the worst of all worlds.
If you're talking about images and whatnot falling off, that's a problem of delivery and not the format.
Markdown translates to HTML one-to-one, it's in the basic features of Markdown. For some reason I have to repeat time and again: use a subset of HTML for papers, not ‘glamor magazine’ formatting. The use of HTML doesn't oblige you to go wild with its features.
I am indeed, and of other tags that are no longer supported. Old sites are often impossible to render with the correct layout. Resources refuse to load because of mixed-content policy or because they're simply gone - which is a problem with the format because the format is not built for providing the whole page as a single artifact. And while the oldest generation of sites embraced the reflowing of HTML, the CSS2-era sites did not, so it's not at all clear that they will be usable on different-resolution screens in the future.
> Markdown translates to HTML one-to-one, it's in the basic features of Markdown. For some reason I have to repeat time and again: use a subset of HTML for papers, not ‘glamor magazine’ formatting. The use of HTML doesn't oblige you to go wild with its features.
This is one of those things that sounds easy but is impossible in practice. Unless you can clearly define the line between which features should be used and which should not, you'll end up with all of the features of HTML being used, and all of the problems that result.
You'll notice that I said in the top-level comment that my beef is with PDF papers (i.e. scientific and tech). I don't care about magazines and such, since they obviously have different requirements. So let's transfer your argument to current papers publishing:
“Since PDF can format text and graphics in arbitrary ways, you'll end up with papers that look like glamor and design magazines and laid out like Principia Discordia and Dada posters. You'll have embedded audio, video and 3D objects since PDF supports those, and since it can embed Flash apps, you'll have e.g. ‘RSS Reader, calculator, and online maps’ as suggested by Adobe, and probably also games. PDF also has Javascript and interactive input forms, so papers will be dynamic and interactive and function as clients to web servers.”
You can decide for yourself whether this corresponds to reality, and if the hijinks of CSS2-era websites are relevant.
What is it with people, one after another, jumping to the same argument of ‘if authors have HTML, they will immediately go bonkers’? If really looks like some Freudian transfer of innate tendencies. We have Epub, for chrissake, which is zipped HTML—what, have Epub books gone full Dada while I wasn't looking? Most trouble with Epub that I've had is inconvenience with preformatted code.
> Old sites are often impossible to render with the correct layout. Resources refuse to load because of mixed-content policy or because they're simply gone - which is a problem with the format because the format is not built for providing the whole page as a single artifact.
Yes, as I mentioned under the link provided in the top-level comment, the non-use of a packaged-HTML delivery is precisely my beef here. The entire idea of using HTML for papers implies employing a package format, since papers are usually stored locally. It's a chicken-and-egg problem. It's solved by the industry picking one of the dozen available package formats and some version of HTML for the content. Which would still mean that HTML is used for formatting. HTML could be embedded in PDF for all I care, if I can sanely read the damn thing on my phone.
Those things don't happen in PDFs in the wild, or at least not to any great extent. It's not that technical paper authors have shown some special restraint and limited themselves to a subset of what the rest of the PDF world does. Technical papers look much like any other PDF and do much the same thing that any other PDF does; if they were authored in HTML, we should expect them to look much like any other HTML page and do much the same thing that any other HTML page does. Based on my experience of HTML pages, that would be a massive regression.
> Yes, as I mentioned under the link provided in the top-level comment, the non-use of a packaged-HTML delivery is precisely my beef here. The entire idea of using HTML for papers implies employing a package format, since papers are usually stored locally. It's a chicken-and-egg problem. It's solved by the industry picking one of the dozen available package formats and some version of HTML for the content. Which would still mean that HTML is used for formatting. HTML could be embedded in PDF for all I care, if I can sanely read the damn thing on my phone.
The details matter; you can't just handwave the idea of a sensible set of restrictions and a good packaging format, because it turns out those concepts mean very different things to different people. If you want to talk about, say, Epub, then we can potentially have a productive conversation about how practical it is to format papers adequately in the Epub subset of CSS and how useful Epub reflowing is versus how often a document doesn't render correctly in a given Epub reader. If all you can say about your proposal is "a subset of HTML" then of course people will assume you're proposing to use the same kind of HTML found on the web, because that's the most common example of what "a subset of HTML" looks like.
This makes zero sense to me. You're saying that technical papers look the same as Principia Discordia or glamor/design magazines or advertising booklets, including those that just archive printed media. That technical papers include web-form functionality just like some PDFs do—advertising or whatnot, I'm not sure. If that's the reality for you then truly I would abhor living in it—guess I'm relatively lucky here in my world.
However, if you point me to where such papers hang out, I would at least finally learn what mysterious ‘complex formatting’ people want in papers and which can only be satisfied by PDF.
It looks promising for these kinds of daunting tasks
https://dl.acm.org/doi/pdf/10.1145/3183713.3183729
And additional follow-up work on extracting data from PDF datasheets is here:
https://dl.acm.org/doi/pdf/10.1145/3316482.3326344
One thing to point out about our library is that while we do take PDF as input and use it to calculate visual features, we also rely on an HTML representation of the PDF for structural cues. In our pipeline this is typically done by using Adobe Acrobat to generate an HTML representation for each input PDF.
[1]: https://github.com/HazyResearch/fonduer/tree/master/src/fond...
Most of our Xerox printers spoke Postscript natively, these days more printers can use PDF. We generally used a tool to convert PCL to PS to suit our workflow if that was the only option for the file, because being able to manipulate the file (reordering and applying barcodes or minor text modifications) was important. Likewise for AFP and other formats. PCL jobs were rare so I never worked on them personally.
If someone could make a service that lets you upload a PDF that contains a form, and then let users fill out that form and e-sign it and collect the results, and then print them out all at once, it would be great.
It's not a billion dollar idea but there are a lot of little companies that would save a lot of time using it.
Again, not sexy, but it is so stupid I have to fill out a direct deposit form by hand and turn it into my company, who checks it, then hands it off to the payroll vendor, who has to check it, just to enter the damn data into a form on their end.
In theory it sounds like it should be straightforward but it hinges so much on how well the document is structured underneath the surface.
Being that these tools were primarily designed for non-technical users first the priority is in the visual and printed outcome and not the underlying structure.
One document can look much the same as another in form—uses black borders to outline fields, similar or same field names, etc, but may be structured entirely differently and that can be a madhouse of frustrating problems.
It can be complex enough to write a solution for one specific document source. Writing a universal tool that could take in any form like that would probably be a pretty decent moneymaker.
My first intuition, though, would be it may be more successful (though no less simple) to develop a model that can read from the visual of the document rather than parsing it successfully.
Open to learning something here, though!
There is a wide world outside of consumers of SaaS products for every little niche problem.
Sometimes they are baked in processes that still use PDF's to share information, sometimes they're old forms of any kind, sometimes even old scanned docs that are still in use but shared digitally. A lot of the businesses that carry on that way are of the mind that "if it's not broke, don't fix it" which is quite rational for their problem areas and existing knowledge base. They might be a potential market at some point for a new solution, but good luck selling them on a web-based subscription SaaS solution when a simple form has been serving their needs for 30+ years.
OP's problem of the PDF being the go-between to digital endpoints is more common than you might think.
The universality I was referring to was the wide range of possibilities for how a given form might be laid out. And old documents contain a lot of noise when they've been added to or manipulated. Look inside an old PDF form from some small-medium sized business sometime. Now imagine 1000 variations of that form one standard problem. Then multiple that by the number of potential problem areas the forms are managing.
Also like OP said—it's not sexy, but it's very real and having an intelligent PDF form reader and consumer would be a time-saver for those businesses who aren't geared to completely alter their workflow.
The tool could do anything with the extracted data. If it allowed you to connect to any of your in house services (like payroll or accounting) either with a quick config/API or a custom patch, or Google Drive, or whatever without complications like online-required and web accounts especially. No whole solution like that exists to my knowledge. At least nothing accessible to the wider market.
Thanks again, I just want to make sure I understand.
OTOH, a PDF form works exactly they way you’d like. Maybe there’s a small market in helping convert one to the other for collecting input from old paper-ish forms.
It lets me type in to forms - or draw text over them if necessary. Then I paste in a scan of my signature. Then save as a PDF an email across.
I've been doing this for years. Job applications, mortgages, medical questionnaires. No one has every queried it.
If you're hand delivering a printed PDF, it's just going to be copy-typed by a human into a computer. No need to make it too fancy.
* https://www.hellosign.com/products/helloworks
* JotForm (https://www.jotform.com/help/433-How-to-Add-an-E-Signature-t...)
(I know about all these because I'm working on a PDF generation service for developers called DocSpring [1]. I'm also working on e-signature support [2], but that's still under development, and still won't be a perfect fit for your use-case.)
Sorry, this is a bit off-topic regarding PDF extraction, but it distracted me greatly while reading...
I'm pretty sure the intention was A B C D (cut then wash). Not sure why the author would not use alphabet order for the recipe...
[edit] Sorry, I made it read to a colleague and he mentioned the A B C D annotations were probably not in the original document. This was not clear at all for me while reading, and if they are not included it's indeed hard to find the correct paragraph order.
And of course, even if the letters were there in the original document, it would be clear to a human that they're incorrect because it doesn't make sense to wash vegetables after cutting.
One workaround I've found is that sometimes it helps to "print to PDF" the original PDF using Preview on Mac. This doesn't fix all the problems, but it does sometimes fix issues with the input PDF — even though both files appear identical to the human eye.
Are there any other workarounds or "PDF cleaners" out there? It would be awesome if there were a web-based service where you could get a PDF de-gunkified, for lack of a better term.
I should probably write a blog post on this.
However... this didn't work in some cases, mainly with formatted text but sometimes with PDFs that looked like they were compiled in some nonstandard way. As a result I ended up chucking the XML structure entirely and recompiling the text from character-level coordinates. Formatted text was also an issue, with slightly offset y coordinates from regular characters on the same line.
I'm not sure I could take this experience and say that extracting _all text_ would be straightforward. Hopefully for most documents the XML is nicely structured, but I imagine there are many more opportunities for inconsistencies in how the PDF is generated when thinking about diagrams, tables etc. rather than just abstracts.
Considered writing up a blog post about my experiences with the above but imagined that it was far too niche. Code's here [1] if it's of interest.
[1] https://gist.github.com/GuyAglionby/4b55d00803710f2e2e9877fd...
pdfless ()
{
pdftotext -layout "$1" - |
sed 's/\f/\n\n ----------------- ----------------- <page> ----------------- ----------------- \n\n\n/g' |
${PAGER:-less -S}
}
The key is the "-layout" argument, which preserves original layout of the document. This ... may not be what you want visually, but makes backing out the original text somewhat easier.Of course, requesting the LaTeX sources would be preferred.
I walked away from the product over a decade ago since it always seemed like it’d be trivial for adobe to implement the feature in reader. Though every couple of years there’s a redaction scandal and I keep wondering how lucrative the product could have been with some marketing.
That's the most interesting point in the article.
Reminds me of how a friend managed to fix bugs in an assembly source file written in the original programmer's very own undocumented special language implemented in the assembler's macro language. He disassembled the resulting object file, fixed the problems, and checked in the disassembly as the new source code.
We provide support for retrieving words as well as a bunch of different algorithms for document layout analysis [1]. But like the other commenters here mention, it's an extremely difficult problem which doesn't have an easy or general solution.
I was trying to build a custom library on top of the open-source library that did a bit more processing, multi-column analysis, statistical analysis of whitespace size, etc. But building something that works for the general case is difficult enough to be functionally impossible.
Despite that I think the PDF format is well suited to what it is for and there are very few "implementation mistakes" in the spec itself (no up-front length for inline image data is the main one, plus accessibility obviously). It's ultimately become too successful and as a result developers are stuck handling cases where it's being used for entirely the wrong purpose but I can't see a way to another format gaining purchase for the correct purpose (perhaps it's like JavaScript in that way, it has huge adoption because it was first, not because it does all jobs well).
Perhaps a content-first format which also handles presentation well could gain a foothold if it came with a shim for PDF viewers and software to use but I dread to think how much effort that would be.
[0]:https://github.com/UglyToad/PdfPig
[1]:https://github.com/UglyToad/PdfPig/wiki/Document-Layout-Anal...
I use it quite successfully to turn my bank statements into text, which can then be further processed.
Also, muPDF for Windows: https://www.mupdf.com/downloads/archive/mupdf-1.16.0-windows... unzip in a folder and run mupdf.exe
Now all I need is a good front end search system for my document archive.
Does the scanning system you use not do OCR?
I use a Mac and Spotlight does a good job of indexing the files. I think alternatives for other OSes might be something like Apache Solr?
So that's where Ghostscript comes in. On a schedule I have a script that picks up new PDFs in the share, runs them through Ghostscript to create a multipage TIFF, that TIFF is then given to Tesseract (as it can't handle PDFs natively) which does the OCR and outputs a nice PDF with searchable text. All very simple.
The scanning of the pages is very fast, but the scanner takes an age sending the PDFs over the network - it's ethernet port is only 100mbit/s but to be honest I just think the CPU inside the scanner is slow. It also doesn't have enough internal buffer which means you can't scan the next document until the previous one has completed being sent to the share.
If I hooked the scanner up to USB, then the PC could run the Brother software which does use OCR - but it's not automatic, all it does is display the PDF inside Paperport once the scan is complete. For bulk scanning, it's not workable.
Regarding indexing - I've started looking at Solr, and it might suit my needs. I was hoping for a visual type search system, where you could see thumbnails of the PDFs in the results.
---
When in doubt, use plain text. It's a million times better in every way that counts.
I wish my bank statements and such could be downloaded as plain text files, instead of massive PDF files that embed another copy of a bunch of typefaces in each file.
It's not.
It used to be long ago, but now it has full programmability with JavaScript.
Chrome displayed them fine, Preview on Mac did not.
Trying to communicate this to them was like talking to a tree, or an alien, or a room of catatonic individuals.
Thankfully I think they've fixed it now.
I don't envy working on ingesting even more diverse PDFs.
We've taken no intentional action to change the way the back button works - in fact, I too hate it when websites do that.
Can you PM me with some details about what you're seeing? I'm having issues reporducing it with my particular setup.
My browser is the latest (v73.0.1) Firefox on the latest build of Windows10. I confirmed the issue with all addons disbled so it is not an addon issue. I think I know what may be responsbile. When initially I load the page the back button works as intended for about a second. After that delay the page seems to load some resources from static.parastorage.com and www.mymobileapp.online. Once those resources are finished loading the back button does not navigate back to the HN article on the first press. Have to press once more. So I presume a script from one of those domains is responsible. Hope this helps!
There are sites that explicitly mess with the back button but I haven't seen one in ages.
I've done lots of work in this space, including computer vision and ML approaches, and Tabula[1] which was the gold standard for extraction.
PDF Plumber is better on just about every example I've tried.
The CEO Simon Mahony basically told me to piss off when I told him I thought the site was misleading since there was no product or service to directly purchase. They make custom developed software that you must pay their consultants to integrate. I would not do business with such a company that acts so unprofessionally even if they have a decent team.
It still hits the issues mentioned in this article (surprise spaces appearing in middle of words, etc)
The workflow was:
- Extract the page images as TIFF, and store the page ranges so I could map the page ranges back to the individual articles afterward.
- Concatenate a range of images one big file, with an upper limit of (IIRC) about 4000 pages. FR would start to generate weird errors when I made the files any bigger than this.
- Run OCR over the giant 4000 page file.
- Export the result as one big PDF with OCR text layer under the scanned pages.
- Split the PDF back into individual PDF files corresponding to articles, using the data I saved in step 1.
- Optimize the individual PDF article files for compact storage, using the Multivalent [1] optimizer.
I did this with a combination of FineReader -- the only paid software -- Python, Multivalent, AutoHotKey, and PDFtk.
I was living on a grad student stipend at the time so I optimized for spending the least amount of cash possible, at the cost of writing my own automation to replace the batch processing found in more expensive editions of FineReader.
The most time consuming part was dealing with weird one-off errors thrown by FR's OCR engine. I had to resolve them all manually. They were too varied and infrequent to be worth automating away.
I tried Acrobat's own OCR too before I resorted to FineReader, but it was pretty terrible. At the time it also appeared to make the PDF files significantly larger, which was weird since a text layer shouldn't take much additional storage.
Sometimes, publishers make their PDF e-books from printed source in which images are “optimized” to low quality JPEGs, but next to them non-display Photoshop data streams with pristine megapixel illustrations are kept. If you catch big PDF files, check their insides, it's one line of `mupdf extract`.
It’s still a pretty manual process, but it does the most difficult part good enough.
Way too busy to write more, but I’ll be back to read comments later tonight.
Wonderfully amusing code comment, about Adobe PSD format: https://fallenpegasus.livejournal.com/854615.html
That's not to say it isn't great; I could well believe it. But I'm just not shocked that PSD and PostScript have ended up being a bit of a mess over the decades. I doubt I could have done any better.
The fact that text might be oriented different wasn't covered in the article. IIRC Preview on Mac might search there (not near my mac ATM to check)
How you process PDF depends greatly on the scale at which you're working with documents. For large-volume, high-speed processing, automation is necessary. Where you're translating a more stable corpus, human input may be tractable. The ability to look at source PDF, OCR, and an edited text version to correct for errors seems a part of that workflow.
Often it's possible to get close or approximate transcription using standard tools. I've found the Poppler library's "pdftotext" remarkably good with many PDFs, so long as there's some text within them: https://poppler.freedesktop.org
There's a general concept I've been working toward of a minimum sufficient document complexity, which follows a rough (though not strict) hierarchy. It's remarkable how much online content is little more than paragraph-separated text, with no further structure. Even images are not strictly informational, but rather window-dressing.
Typically, additional elements added are hyperlinks, images, text emphasis (italic and bold, often only the first), sections, lists, blockquotes, super- and sub-script, in roughly that order.
(A study looking at the prevalence of specific semantic HTML elements within a corpus would be ... interesting.)
Then there are the elements NOT natively supported in HTML: equations, endnotes/footnotes, tables of contents, etc.
It seems to me there should be an analogue to Komolgrov complexity as concerns layout of textual documents. That is: there is a minimum necessary and sufficient level of markup (perhaps: number, type, and relationship of elements) necessary to lay out a specific work.
I've tagged out novel-length books in Markdown with little more than the occasional italic and chapter marks.
Documents which use more markup than is required are overspecified. This is the underlying problem with a great deal of layout, and the ability to reduce texts to their minimum complexity would be useful. It's a nontrivial problem, though large swathes of it should be reasonably achievable.
Another approach would be for information-exchange formats to actually be, you know, information exchange formats rather than PDF.
(Though the latter is often, though not always, well-suited to reading.)
It's critical that the training data is good quality and much of the engineering effort should go into good annotation interfaces. We built an end-to-end system for all this at evolution.ai. Please email me if interested in an off-the-shelf solution. martin@evolution.ai.
Which brings me to a question: what alternatives do we have to PDF?
PDF needs to die. Djvu is good.
> Why not OCR all the time? > Running OCR on a PDF scan usually takes at least an order of magnitude longer than extracting the text directly from the PDF.
... so? Google can OCR video and translate it in something that feels like real-time; what PDF processing are they doing that is so performance bound?
> Difficulties with non-standard characters and glyphs OCR algorithms have a hard time dealing with novel characters, such as smiley faces, stars/circles/squares (used in bullet point lists), superscripts, complex mathematical symbols etc.
Sure, but more than the random shit you find in PDFs anyway?
> Extracting text from images offers no such hints
Finding an algorithm that approximates how a human approaches a page layout doesn’t feel like it would be all that hard.
Obviously it’s very easy to stand on the sidelines and throw stones, but parsing PDFs using anything other than OCR + some machine learning models to work out what the type of a piece of text feels like pretending we are still constrained by the processing costs of 5 years ago
1) Do you have trillion or so dollars at your beck and call? If not, you're not Google.
> Finding an algorithm that approximates how a human...
2) ...is generally nigh impossible even for someone with Google's resources (e.g. Waymo, although when it comes to reading, it's somewhat usable). Also, look at 1)
Unless by approximate you mean toddler level. In that case:
3) The approximation is probably useless
"In CS, it can be hard to explain the difference between the easy and the virtually impossible."
You italicized the word 'solved' to emphasize that under technical scrutiny it is not true in response to a claim about being pedantic.
100% accuracy on identifying birds is, epistemologically, impossible.