Quite strange this scepticism here about good old djvu. Better fetch a book from the Internet Archive in both formats and check for yourselves why it's a thing.
50 years from now, you are not going to have any trouble finding a .PDF viewer, but DjVu is a different question entirely.
Frankly I don't expect having to read PDF files 50 years from now, but if I were thinking about preservation, the timeproof archiving medium still is acid-free paper –better yet, vellum– and probably India ink. As for the need of any technology, beyond stone tools and fire nobody really needs anything, it's just convenient.
I think you may be confusing tools with file formats. There are any number of ways to do this with .PDFs, none of which I've personally had to explore.
Frankly I don't expect having to read PDF files 50 years from now, but if I were thinking about preservation, the timeproof archiving medium still is acid-free paper –better yet, vellum– and probably India ink. As for the need of any technology, beyond stone tools and fire nobody really needs anything, it's just convenient.
Meanwhile, back on this planet...
Exactly, in this planet we still can read, for instance, books printed by Gutenberg, yet try to convert some Word files rotting in a low-density macintosh floppy from 25 years ago. Paper beats everything else because it's a passive medium that lasts centuries and comes with all the hardware you need. A computer file needs a whole ecosystem and being copied around often.
Will we even use files in 2068? I don't know. Maybe in 2047 PDFs will be out of mainstream support because 99% of customers will be using the next better format or workflow.
Meanwhile, a .PDF from the same era would still be perfectly readable. There's a lesson there.
Ordinary users just couldn't make pdfs without paying for additional publishing software until some ten years ago.
The lesson is that some formats become de facto standards, then people forget how that happened (and that it could happen again), and ignore better technical solutions for specific jobs.
That is a very bold statement. 50 years is a long time even in human years. It's an eternity in technology years.
Specifically, what will kill off the .PDF format while leaving DjVu untouched?
- Moving to HTML and signed HTML file.
- JPEG > PDF on mobile so people send the first over the later.
But going out on a limb with a format doesn't instill confidence for long term. It's great there's one "perfect" implementation that seems to just work, but perceived lack of an ecosystem is still worrying. Not that looking through the source of say libjpeg instills confidence either! I was also considering JP2 for nicer-looking artifacts, but couldn't bring myself to go that route.
I eventually just increased my size budget (because disappearing a ream into 5GB is still damn useful!), and decided to just store lossless FLIF at 300dpi. I statically compiled the flif binary, do a test decode to make sure the bits round trip, and store checksums of the decoder binary and raw raster alongside.
I then stuff this into a zip container (aka .cbz), along with some really poor jpg thumbnails so present evince can view it. I still need to write the transcoder to .djvu - when I'm no longer buried under piles of paper!
Something a little more up to date and more focused on today's technologies would be really nice.
It’s a very specific and ( hopefully) narrow need and is the only reason i’m starting this project in the first place.
Do you know a way to crop a PDF? Basically, if I have a PDF of a single letter-size page with a 6x9 label on it, and I'd like to crop it down to that 6x9 label only, is there a way to do this? I've been looking into this with, for instance, the poppler library, and haven't found anything.
You could parse the pdf for that bit of text then simply recreate a new one with the correct size. But it depends a lot of how that label is printed ( using one block for the whole label or one per letter). Basically it's a lot of work or not depending on the number of cases you need to deal with.
I know python has a python miner library, you may want to try that
My issue with JPG conversion is the lossiness factor introduced there; I'm trying to avoid that.
There should only be one case: a single label on a letter-size page surrounded by a lot of whitespace. I'll check out that miner library; thanks!!