Xz format inadequate for long-term archiving
nongnu.org
nongnu.org
1. The article itself recommends that "a tool is supposed to do one thing and do it well". Well, xz is for compression, nothing else. If you want to check integrity, use a hashing tool such as sha256sum. If you want to recover from errors, generate parity files using par2 or zfec.
2. The article claims xz was chosen by prominent open source projects due to hype. No, it was chosen because of its favorable size/speed tradeoffs.
3. It's not clear what the author means by "long-term archiving". Archives professionals (as in, people who are actually employed by memory institutions) will tell you of many factors (all unmentioned here) that have a bearing on whether a file format (particularly a compressed one) is suitable for use in data preservation.
4. The section on trailing data is particularly bizarre. By claiming that xz is "telling you what you can't do with your files", it seems that the author considers it perfectly reasonable to append arbitrary data to the end of a file and expect it to continue functioning. Is my normality detector off today or is this just a wacky thing to want to do?
5. The article is written by the author of lzip (which it compares favorably to xz), but does not disclose this. Admittedly it is hosted on the lzip site, but overall the article comes across as an opinionated hit-piece designed to sow doubt in a competitor.
Making a backup with tar is done by typing something like that on bash:
> tar -c - dir1 dir2 dir3 > /dev/tape
That will (hopefully, I doubt I got the tar switches right) backup those dirs into the tape (that will actually have a weird name, not '/dev/tape').
Now, in practice Linux doesn't always know the size of a tape you inserted. But this is not the issue, if you accept the seeks needed for that, you'd better write at the beginning anyway.
Kids these days don't appreciate having random addressable storage for archive/backup data!
No: that's why the LZMA compression format was chosen, and why the .lzma container format temporarily became popular as people started using lzma-tools. The move to xz, which results in slightly larger files for seemingly arbitrary reasons and which technically provides one benefit (seekability) almost everyone defeats (by compressing a tar file) while using the same relying on the same underlying compression algorithm is less-well motivated.
My impression was that the whole 7zip->p7zip->lzma_alone->xz journey was concluded before any "xz hype" began.
[0] https://lists.debian.org/debian-devel-announce/2011/08/msg00...
When I burn data (including xz archives) on to DVD for archival storage, I use dvdisaster[2] for the same purpose.
I've tested both by damaging archives and scratching DVDs, and these tools work great for recovery. The amount of redundancy (with a tradeoff for space) is also tuneable for both.
[1] - https://github.com/BlackIkeEagle/par2cmdline
[2] - http://dvdisaster.net/
What's the oldest disc you've tested?
What do you think about PAR3?
I just tried a DVD from 2013 (Verbatim DVD-R, 16x, 4.7 GB) and it read fine, without any errors, and dvdisaster found that the checksums on it matched.
I usually try to use at least 20% redundancy, but will settle for less if the data's not very important. If the data's really important I'll max out the amount of redundancy (up to dvdisaster's limit, which I don't remember off-hand). Sometimes I'll even burn an extra DVD with the same data, both with dvdisaster error correction on them.
As for par3, I only heard about it for the first time today in this thread. So I have no opinions on it except to say that if it's an improvement over par2, I'm all for it. Backwards compatibility with par2 also would be nice.
Shame about the status of the dvdisaster project - as of 2015 the Mac OS X and Windows ports are discontinued.
With compressed files you usually need it to be perfect to recovery any data, but you can use fixtar [0] to recover some data from a corrupted tar archive.
I used BluRay as the media cost is about the same as DVDs per GB here (at least 25GB discs). Also there are HTL discs available, which supposidely have a greater longevity than DVDs.
[0] -http://riaschissl.bestsolution.at/2015/03/repair-corrupt-tar...
However what's totally awesome about xz as a container format is seekable random access to compressed files. I use it to store infrequently used disk images, which I can boot up without even decompressing them. (https://rwmj.wordpress.com/2013/06/24/xz-plugin-for-nbdkit/)
A compressed container format that allows quick access to individual files is very useful and actually improves data security - corruption in one part of the data will be far less likely to ruin the entire collection of files, whereas corruption in a compressed tar file may lead to the loss of everything (I know recovery tools exist, but they cannot prevent one file in the container being dependent upon data from another file.)
It does seem like xz is somewhat overengineered, but I don't think that's a characteristic unique to it; any other algorithm with similar compression performance will yield similar behaviour on corrupted data, since the whole point and why compression works is to remove redundancy. I say use something like Reed-Solomon on the compressed data, thus introducing a little redundancy, if you really want error correction.
>> Just one bit flip in the msb of any byte causes the remaining records to be read incorrectly. It also causes the size of the index to be calculated incorrectly, losing the position of the CRC32 and the stream footer.
This sounds severe enough to me.
What happens if bits in the block header are corrupted? If it can't find the start of the next block, the same thing will happen.
Also, breaking up the data into blocks will decrease compression, since each block starts with a fresh state. It is ultimately a tradeoff between compression ratio and error resistance.
This has been addressed in the article. "Bzip2 is affected by this defect to a lesser extent; it contains two unprotected length fields in each block header. Gzip may be considered free from this defect because its only top-level unprotected length field (XLEN) can be validated using the LEN fields in the extra subfields. Lzip is free from this defect."
> Also, breaking up the data into blocks will decrease compression
This has been tested very thoroughly. Larger block sizes give rapidly diminishing marginal returns (man bzip2). Now, the largest you can go with bzip2 is 900KB.
I've tried all the usual implementations that are available on Windows recently, and they were all unusable. Like in "try to give it a hundred files, each in a single-megabyte range, and it will crash hard".
I'd agree about the tooling problems with it, it's "acceptable" on a linux/unix command line but I've never seen anything elsewhere that even looked halfway usable.
You seem to confuse format and compression algorithm. From what the article says, the format seems bad. Once any data is damaged, all the rest of reading the file (as opposed to decompressing the data from file) goes out of the window. No way to re-synchronize reading with data blocks after corruption.
The argument makes more sense for the casual user who isn't paying for a professional backup service. But how many of those are using xz? They're more likely to be using zip files; and they'll be in a better place for it.
The zip (concatenation of compressed streams) vs tar.gz (compression of concatenated streams) distinction is useful in an archival context, if one is genuinely worried about bitrot and the risk of needing to recover partial data where the container is damaged. A concatenation of compressed streams will have worse compression especially for lots of small files, but it is far easier to recover from an error.
xz wasn't chosen for its archival integrity AFAIK, it was chosen for its size. Package distribution ought to spend as much resource as practical to reduce bytes over the wire. Size is why bzip2 "won" over gz despite being many times slower and requiring much more RAM. lrzip would be my first choice for maximum compression, but its downstream resource requirements are far too high vs any of the others, so from my experience xz would be the answer anyway.
No: that's why the LZMA compression algorithm is being used. The switch from the original LZMA container format to the "improved" xz container format had nothing to do with size, as the old container generated slightly smaller files in addition to not having this combination of negative properties.
But I dogress. When you create archival packups you use must use error correcting systems. Par is a good place to start. https://multipar.eu/ https://github.com/Parchive/par2cmdline
If it offered none, and explicitly stated that, it would be one thing, but it offers some that functions quite poorly in practice, which is rather another.
I think there is scope in xz to embed some FEC in the archive as a backwards compatible extension. Not such if the decoding scheme could be changed to reduce error cascades etc.
Debian source packages are cryptographically signed. Debian archives are also cryptographically signed. There's no error-correction, but you don't want that. You just want to verify integrity every time you make a copy, and re-copy if it didn't copy right. And all the tools do that.
I remember this discussion on debian-devel, and there were a lot of people pointing out that the author's guesses at Debian's threat model had little to do with Debian's actual threat model, at which point the author became frustrated.
So it was the right choice for the Xz format to use variable-length integer encoding. Not all integer fields will need to represent values up to 2^63, but it doesn't matter.
Individuals may care about space more because the storage or medium is expensive for them. Even then, bulk data like photos and video doesn't compress well, so this argument is still pointless.
Is that a Turing-complete bytecode? If so, I envision some interesting applications with regard to procedurally-generated content...
But of course for real world use cases one would choose a general compression algorithm or write a specific one for a certain type of data you want to handle (because it will perform better than a general algorithm).
So, what does mince do? https://github.com/Kingsford-Group/mince/blob/master/README.... says:
"Mince is a technique for encoding collections of short reads."
That doesn't tell me much.
Also, http://bioinformatics.oxfordjournals.org/content/early/2015/...:
"We present a novel technique to boost the compression of sequencing that is based on the concept of bucketing similar reads so that they appear nearby in the file"
I'm still not sure I understand that fully, but guesstimate that it shuffles file content in such a way that it compresses so much better that it more than offsets the space needed to store the information on how to deshuffle the decompressed data. Is that correct?
Or does it require you to specify the parts and discards the reshuffling information, assuming you aren't interested in the order?
It sounds like it's meant results of DNA sequencing processes that produce sequences for random small pieces of the molecule.
I'm a developer and I was LITERALLY in the process of backing up my iPhoto library (200GB) using Xz just now.
welp
So rzip is probably excellent for archival situations (not counting any flaws the algorithm may have) but not so great for things like compressed tarballs (which are meant to be appendable).
$ file COPYING.*
COPYING.lz: lzip compressed data, version: 1
COPYING.xz: XZ compressed data
To me, this indicates a lack of data in the magic database used by file, more than something inherently bad with xz.This is why it's usually paired with gzip or xz.
It boggles my mind why reflow is not one of the most important and most highly regarded features of any mobile browser.
on opera mobile: are you truly comfortable to use a 'default on man-in-the-middle-ware' from china? (it proxies everything and compresses it, to save data bandwidth)
i stopped using it in August because of that. might've been too paranoid though.
Personally I am less concerned about the Boogiemen from China/Russia than about the pervasive ruin of privacy by US companies.
>it is your browser's duty to render it nicely
I think browsers should have the ultimate say on how to display websites, according to the user's preferences (think of text size and colors). Sadly, with websites gradually transforming into "web apps", users are losing control.
[1]: https://addons.mozilla.org/En-us/firefox/addon/fit-text-to-w...