Just use a file system that does proper integrity checking/resilvering. Also use TLS to transfer data over the network.
Just use a file system that does proper integrity checking/resilvering. Also use TLS to transfer data over the network.
Storage with integrity checking would be the solution to bitrot, but TFA also seems concerned with "how do you unarchive/recover a random file you found?" which seems a somewhat valid concern.
And xz does have support for integrity checking, so it seems reasonable to have a discussion on whether that is a good support, rather than on whether it should be there at all.
Archives need not only a way to check their integrity but also error correction, which xz does not have.
However, you can easily combine xz with par2, which does provide error correction.
Self-referentially incorruptible data is the standard to beat here, moving the concern to a different layer doesn't increase the efficiency or integrity of the data itself.
It is arguably less efficient, as you now rely on some lower layer of protection in addition to whatever is built into the standard itself.
It is less flexible - a properly protected archive format could be scrawled on to the side of a hill, or more reasonably onto an archive medium (BD-disk), and should be able to survive any file-system change, upgrade, etc. Self-repairing hard drives with multiple redundancies are nice, but not cheap, and not wide-spread.
It also does nothing for actually protecting the data - I don't care how advanced the lower level storage format is, if you overwrite data e.g. with random zeros (conceivable due a badly behaving program with too much memory access, e.g. a misbehaving virus, or a bad program or kernal driver someone ran with root access, also conceivable due to EM interference or solar radiation causing a program to misbehave), the file system will dutifully overwrite the correct data with incorrect data, including updating whatever relevant checks it has for the file. The only way around this is to maintain a historical archive of all writes ever made, and that should be evidently both absurdly impractical (how do you maintain the integrity of this archive? With a self-referentially incorruptible data archive perhaps?) and expensive.
Compared to a single file, which can be backed up, agnostic to the filesystem/hardware/transport/major world-ending events, which can be simply read/recovered, far into the future. There's a pretty clear winner here.
I respectfully disagree. By putting it in the layer below, there is the ability to do repairs.
For example, consider storing XZ files on a Ceph storage cluster. Ceph supports Reed-Solomon coding. This means that if data corruption occurs, Ceph is capable of automatically repairing the data corruption by recomputing the original file and writing it back to disk once more.
Even if XZ were able to recover from some forms of data corruption, is it realistic that such repairs propagate back to the underlying data store? Likely not.
If you can't read the data in question though, you cannot do the repairs, it doesn't matter if you do Reed-Solomon coding or not. You are thinking about coding for the underlying hardware, which is what the data corruption you are talking about is designed to fix - it does not solve the problem for writes coming from above.
To do that, you actually have to decode the data in question and perform a reed-solomon encoding on the actual file inside of the archive, and this only gets worse e.g. if you have nested archives.
If the data is self-referentially repairable, however, it doens't matter if the file gets overwritten with e.g. a cat gif, the format will work around that. The filesystem on the other hand will have written the cat gif to the file and updated the Reed Solomon encoding for your file, assuming (incorrectly) the file writes were valid.
I suppose you could mandate that for any file to be written to your filesystem it must first be completely decompressed, and then store some encoding information alongside the archive, but this would be inefficient to the extreme, since merely copying a file onto the system would mean you have to decompress the file and then checksum it.
At any rate, even if you did decompress the file in question, you have failed to separate the layers like you want to, since now you have mandated the XZ and LZMA algorithms also be baked directly into the filesystem itself.
Better not to needlessly couple the filesystem to some compression algorithm, let the compression system handle its own error correction.
The point about being able to use different media like bluray drives is a valid point but since xz doesn't do any correction it doesn't really matter, it has to be done out-of-band anyway.
The simple cost effective thing is not to engineer a complex redundancy system above and below to try to adhere to some misguided "separation of concerns", its to use the simplest, most effective solution which presents itself.
When you try to separate things which should not be separated in a software (or other) system, you get high coupling, low cohesion. Not everything should be attempted to be "de-coupled".