DjVu, an open PDF alternative
en.wikipedia.org
en.wikipedia.org
Actually, now that I think about it, the slide started all the way back in 2010 when I couldn't find a decent DjVu viewer for the iPad.
Also, I don't always scan myself. Sometimes I'll get a PDF generated by someone else, which contains scanned images. Then I'll often convert it to DJVU for speed and space savings.
Just be careful that it doesn't silently change the text, because JBIG2 compressors can be lossy that way:
I presume it only works with NTFS? Surprised the page doesn't say.
[1] https://www.petri.com/compress-files-with-compact-exe
[2] https://stackoverflow.com/questions/7928840/how-to-use-compa...
I presume that's for simplicity?
[1] https://serverfault.com/questions/617648/transparent-compres...
I'm not particularly knowledgeable about filesystems, or indeed of deep Linux matters in general, but the idea of layering ZFS and ext4 strikes me as the bad kind of 'cute' solution. If you want compression, ext4 surely isn't the place to start.
I like the tool though. Couldn't get compact.exe to work, it reported compressing 0 of the files in the directory. CompactGUI worked fine though. Maybe it invokes compact.exe once per file, rather than across the directory.
One interesting trick I've tried, however, is to first remove all internal compression from PDF streams, and then apply LZMA to the resulting PDF. This can be a good way to compress PDFs losslessly.
And this out-compresses PDF's internal compression?
Could PDF be updated in future to support superior compression algorithms?
A less ambitious alternative might be to use something like zopfli. It is totally backwards compatible so it can be used today to compress PDFs better without sacrificing compatibility.
But going out on a limb with a format doesn't instill confidence for long term. It's great there's one "perfect" implementation that seems to just work, but perceived lack of an ecosystem is still worrying. Not that looking through the source of say libjpeg instills confidence either! I was also considering JP2 for nicer-looking artifacts, but couldn't bring myself to go that route.
I eventually just increased my size budget (because disappearing a ream into 5GB is still damn useful!), and decided to just store lossless FLIF at 300dpi. I statically compiled the flif binary, do a test decode to make sure the bits round trip, and store checksums of the decoder binary and raw raster alongside.
I then stuff this into a zip container (aka .cbz), along with some really poor jpg thumbnails so present evince can view it. I still need to write the transcoder to .djvu - when I'm no longer buried under piles of paper!
Quite strange this scepticism here about good old djvu. Better fetch a book from the Internet Archive in both formats and check for yourselves why it's a thing.
50 years from now, you are not going to have any trouble finding a .PDF viewer, but DjVu is a different question entirely.
Frankly I don't expect having to read PDF files 50 years from now, but if I were thinking about preservation, the timeproof archiving medium still is acid-free paper –better yet, vellum– and probably India ink. As for the need of any technology, beyond stone tools and fire nobody really needs anything, it's just convenient.
I think you may be confusing tools with file formats. There are any number of ways to do this with .PDFs, none of which I've personally had to explore.
Frankly I don't expect having to read PDF files 50 years from now, but if I were thinking about preservation, the timeproof archiving medium still is acid-free paper –better yet, vellum– and probably India ink. As for the need of any technology, beyond stone tools and fire nobody really needs anything, it's just convenient.
Meanwhile, back on this planet...
Exactly, in this planet we still can read, for instance, books printed by Gutenberg, yet try to convert some Word files rotting in a low-density macintosh floppy from 25 years ago. Paper beats everything else because it's a passive medium that lasts centuries and comes with all the hardware you need. A computer file needs a whole ecosystem and being copied around often.
Will we even use files in 2068? I don't know. Maybe in 2047 PDFs will be out of mainstream support because 99% of customers will be using the next better format or workflow.
Meanwhile, a .PDF from the same era would still be perfectly readable. There's a lesson there.
Ordinary users just couldn't make pdfs without paying for additional publishing software until some ten years ago.
The lesson is that some formats become de facto standards, then people forget how that happened (and that it could happen again), and ignore better technical solutions for specific jobs.
That is a very bold statement. 50 years is a long time even in human years. It's an eternity in technology years.
Specifically, what will kill off the .PDF format while leaving DjVu untouched?
- Moving to HTML and signed HTML file.
- JPEG > PDF on mobile so people send the first over the later.
Something a little more up to date and more focused on today's technologies would be really nice.
It’s a very specific and ( hopefully) narrow need and is the only reason i’m starting this project in the first place.
Do you know a way to crop a PDF? Basically, if I have a PDF of a single letter-size page with a 6x9 label on it, and I'd like to crop it down to that 6x9 label only, is there a way to do this? I've been looking into this with, for instance, the poppler library, and haven't found anything.
You could parse the pdf for that bit of text then simply recreate a new one with the correct size. But it depends a lot of how that label is printed ( using one block for the whole label or one per letter). Basically it's a lot of work or not depending on the number of cases you need to deal with.
I know python has a python miner library, you may want to try that
My issue with JPG conversion is the lossiness factor introduced there; I'm trying to avoid that.
There should only be one case: a single label on a letter-size page surrounded by a lot of whitespace. I'll check out that miner library; thanks!!
Particularly because Xerox photocopiers (yes, self-contained photocopiers) had exactly the same kinds of issues due to image compression engine glitches (https://news.ycombinator.com/item?id=6156238, https://news.ycombinator.com/item?id=9584172)
JBIG2 was based upon JB2.
Do you have some concrete examples of corrupted texts? I'd very much like them to use as test cases in a document processing pipeline.
How would you know? Maybe if a document had an unusual misspelling. But how would you know if DjVu munged a 3 into an 8?
It seems also that the problem happens more often with cyrillic glyphs which have less descending/accending elements that typical latin ones.
I totally understand how this could theoretically happen but it all hinges on it actually happening in the documents that I'm working with and to have a test case where it happens for sure would make all the difference in determining whether or not the documents I'm working with are at risk or not and if so if I can detect which ones are at risk (that would be half the battle won).
It would have to be a case where a human would see a 3 or an 8 (or a similar transposition with other glyphs) resulting in a document that is corrupted in such a way that afterwards the human would see the (invalid) alternative.
Even a single instance of this happening would be very relevant.
To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity. The compression algorithm could not only mistake '3' as '8' with small gaps left, it would replace this dirty '3' with the image of reference '8' image, so that human reader think that he sees a scan, i.e. an image of page scanned, while in fact it sees 'edited' image, not corresponding to actual page.
I take your word, there is a technical reason behind this, not that I don't believe you.
> To clarify, it's not about OCR errors, where advanced OCR engine could use dictionary, language autorecognition etc. It's about dvju-compressed text scans without OCR layer. The compression method relies only on glyphs similarity.
Yes, I totally get that. It's about DjVu's compressor replacing the image of one character with the image of another.
You would obtain more credibility by actually providing such an example.
By comparing it with the original, surely.
DjVU: https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
PDF: https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
I don't have an easy way to find all of the errors I observed but typical ones are on page 58 at the end of the third line, 'Germanic' with a little bit of bleed between the serifs at the bottom of 'n' gets transformed into 'Germauic', and on page 420 at the end of line 7, 'compass' with a small dot in the middle of the 'o' becomes 'cempass'.
FWIW I've transcribed a bunch of other DjVu files produced by the Internet Archive and not noticed anything like this.
If I can reliably flag problem cases then at least I will know which files are going to have a lot of manual work.
There aren't a lot of viewers out there for DjVu and the encoding side is patent encumbered, so I'm not interested in the format.
You can get pretty close with JBIG2+Jpeg2k in a PDF file, I believe archive.org does this, but I don't know of an open source encoder that does it and sometimes PDF viewers don't decode jbig2/jpg2k efficiently.
DJV View http://djv.sourceforge.net/
WinDjView https://windjview.sourceforge.io/
Other than that, DjVu is a great format.
For those who wonder how it's better than PDF for scanned texts - DjVu uses different compression for background and actual text, thus saving tons of space.
But worst case scenario you could run something like Pale Moon or WaterFox only for accessing such documents.
Also I know how slow institutional entities can be to adapt, but they will catch up eventually.
Okular(provided by KDE), for example, is blazingly fast. Or at least, I haven't noticed any slowness.
Also PDF is an open format.
Now that it is, there's probably no reason to use DjVu.
btw if you want quickly some random large scanned PDFs for testing, archive.org is a goldmine. For example here is a scanned 1200 electronics handbook: https://archive.org/details/NationalSemiconductorLinearAppli...
They also do DjVu, so you can see the size difference there. For that particular book, its 58M for the PDF vs 31M for DjVu. So basically half the size.