Pdf2htmlEX – Convert PDF to HTML without losing text or format
github.com
github.com
It is definitely the best solution I've found so far. The outputted HTML / CSS / images look almost identical to the source PDF. That being said, there are a few issues still:
* One Gigantic (600kb) CSS file from a single PDF
* Hundreds of individual fonts
* HTML semantics are non-existent
These are all relatively easy to fix, I believe. I have found my own solutions to most of the issues in post-processing.
Kudos to you, coolwanglu. Also, I'd like to get in touch with you about lending a hand to fix some of the issues I've encountered.
Thanks for a cool piece of software!
2nd & 3rd are in the future plan, as I'm still working on accuracy and speed. And #115(https://github.com/coolwanglu/pdf2htmlEX/issues/115) is about the 2nd issue.
About the first one, I've not got an elegant solution yet, maybe a CSS file per page?
Please file new issues at GitHub if you think it's necessary :)
These are all relatively easy to fix, I believe. "
How? For example, how would you identify <span>'s (or whatever this converter uses) to identify headers, and page headers/footers, or a ToC, or a preface? IMO this is an AI-hard problem, for which even the 'simple' approximation (statistics) is very hard due to the wide variety in inputs (a corpus trained for multi-column journal articles will most likely not work at all for books, although I haven't tried and would love to be proven wrong).
Use case: a working (i.e., preserving semantics) pdf-to-epub converter. This would, imho, be a killer product / service.
wkhtmltopdf [0] is probably the most popular, but it's also ridiculously buggy.
The PDF's it outputs are full vector not just rasters, it the same engine used in Chrome to view PDF's and print web pages from my understanding.
Then use ps2pdf from ghostscript.
You can automate this with a small amount of work.
The idea is that now the document becomes more controllable and accessible, say you can put Google Analytics in your resume written in LaTeX; or maybe an social reading service, where you can comment, annotate and share.
Unlike PDF viewers, web browers are never optimized for this kind of messy inputs. The next version of pdf2htmlEX will be focused on optimizations, e.g. smaller size of background images, hopefully that would help.
I truly wish there was at least one ground that hadn't been touched by "social" crap.
Still like old Google Reader with its OLD social features.
It is no surprise iOS handles rendering PDF's so quickly and so well and without the need for an third party app, it always has from the release of the first iPhone. This is also why print to PDF is built in on OSX.
PDF the file format adds many, many things to that (forms, encryption, DRM, notes, a JavaScript engine, reflow information, etc)
I guess I don't really see much practical purpose for it -- most browsers these days seem perfectly fine opening PDF files natively, after all. But it's a very cool technological demonstration.
Maybe this could be some kind of bridge tool for generating sites with fancy typographical layout? You could use Adobe Illustrator etc. to do fancy column work, drop caps, hyphenation, all that jazz -- and then "render" into HTML. It would certainly be as anti-"responsive" as you can get, but it would certainly have the ability to generate more advanced typography much faster than you can produce with HTML/CSS by hand.
Convert to HTML -> Edit -> Print back to PDF (if needed)
Say you have a resume written in LaTeX and you want to insert Google Analytics inside?
It would be definitely interesting in that way, but in that case it may not be worth it to rewrite everything in JS.
Features about recognition would be planned in the future, usually PDF viewers do not recognize too many things, do they? :)
This means that you can create one of your own.