JBIG2 Undetectable Data Corruption: Destroying Our Past, One Character at a Time
circuitousroot.com
circuitousroot.com
Immediately this talk from David Kriesel comes to mind. :)
It's also telling that the example trotted out is always a decades-old photocopier implementation. Unless someone does some comparisons of recent scans and finds a bunch of altered characters, I just don't buy this concern. MRC+JBig2 is really great and the size savings have not been trivial for me personally.
(I hesitate to add that the alternative would be, I assume, DCT/jpeg page images, which introduce lots of noisy artifacts of their own... destroying our history one DCT-block at a time!)
There's apparently 22 pages of fluff at the beggining, so you should add that the page numbers. Also "the third line" was meant to be the 3rd line from the bottom.
Direct links (to be opened in DjView):
https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
https://ia800204.us.archive.org/28/items/cu31924026442156/cu...
So, I have to conclude that the original conversions were done with very aggressive settings (or very bad jbig2 implementations). If that's they way the internet archive did all its encodings, then I admit there must be several problem documents, and they should re-do the conversions from the original JP2k files they seem to have kept. Not sure if google books made similar mistakes.
The most obvious way to compress pages is to losslessly encode them as a single 1 bpp image (see Figure 2). However, we get much better compression by using symbol encoding and accepting some loss of image data. Although our compression is lossy it is not clear how much information is actually being lost - the letter forms on a page are obviously supposed to be uniform in shape, the variation comes from printing errors and lack of resolution in image capture.
Probably only a small fraction of errors are JBIG2 substitutions though.Jpeg artifacts on text are ugly but they tend to stand out. You will see artifacts long before the text becomes unreadable or ambiguous, and that is a good thing.
I don't belive that's correct; it's rather the other way round.
https://github.com/barak/djvulibre/blob/release.3.5.28/libdj...
> JB2 has strong similarities with the forthcoming JBIG2 standard developed by the "ISO/IEC JTC1 SC29 Working Group 1" which is responsible for both the JPEG and JBIG standards. This is hardly surprising since JB2 was our own proposal for the JBIG2 standard and remained the only proposal for years. The full JBIG2 standard however is significantly more complex and slighlty less efficient than JB2 because it addresses a broader range of applications.
The other comments here have linked to the previous articles about this, which do give far more detailed information about the problem. JBIG2 in lossless mode won't do this.
But if it happened once, who's gonna guarantee this won't happen again? When it happens with numbers, worst case it can have fatal consequences.
The problem is the existence and use of a codec for a purpose, that harms that purpose. It doesn't matter what it's name is.
Are scans being produced with this codec? And is it not true that not only is there data loss or corruption like with jpg, but that unlike jpg artifacts you can sometimes not know that the data loss or corruption occurred?
That does make it not merely a problem of "I wish they scanned this in higher quality" but a far worse problem of "this document looks good so I trust it" when in fact it was corrupt.
"A new research report demonstrates the connection between the rising amount of mass produced housing featuring a dedicated sex dungeon and biased content in the training set used for the compression algorithm in a popular line of architectural plotters"
Most probably its raison d'être is for non-critical stuff, such as literature, where character flips may be easily detectable during encoding and after the fact (eg. with a spell-checker) and won't possibly subvert the meaning of a whole literary work anyway.
That really can’t be hard to do, since there are lots of common books in GB, there are even OCRs available so you could even do the error checking automatically for thousands of books.
If you can’t show a single case of JBIG data corruption in GOogle Books, you have absolutely no justification in calling it such a serious problem!