4Q: the final archive format
github.com
github.com
And no, storing it on a ZFS or BTRFS volume with error recovery enabled does not count. (They don't use FEC, they use 1960's triplicate storage. Hugely wasteful of space, does nothing to protect against transmission errors and can still be corrupted by two of the exact wrong bits being damaged.)
Storing it on a media that does use FEC also does not count. I want a per-file tunable FEC knob, not one vendor-determined setting. And as history has shown, it needs to be done through FOSS code and not trade secret firmware.
Supposed to. If you do any backups to DVD/BD and still want to get your data back in 5 or 10 years, though, you'd be well advised to do some sort of FEC - burn multiple copies of each disc, generate a bunch of PAR2, whatever.
(You might want to do that for backups on hard drives too. Yeah, maybe the hard drive firmware is supposedly taking care of any errors below the block level and you're not too worried about bitflips, but that just means you'll lose entire blocks and files when you lose something.)
Also read up on RAID levels. RAID5/6 or ZRAID7 uses error correction not duplication.
nc -l 7000 | tar -xf -
tar -czf - * | nc 1.2.3.4 7000
Files appear one by one on the remote end.Personally I'd like a tool that lets me extract files from TB size archives easily without decompressing everything. Only virtual disk images seem to have that random access functionality. There are ways of using gzip/bzip/xz that use sync points, so you could produce a compatible archive that allowed decompressing metadata bits, though it would suck for many small files.
I'm not too familiar with Coffeescript, but it doesn't seem like a good choice of language to write an archiver. There's no actual draft file format spec I can see, either? But from a first pass, I have the following comments:
Crypto: Encrypted blocks: AES-256-CBC, random IV, with no MAC (!!!). You need to look at that again: that could be a Problem. Hashed blocks: SHA-2-512. Maybe OK (how's length encoded? Look out for extension attacks). That crypto is 14 years old and missing a vital bit: not "modern". Modern choices would include CHACHA20_POLY1305 (faster, more secure, seekable if you do it right); hashes like BLAKE2b (as the new RAR already does); signing things with Ed25519. Look into that kind of thing. You need a crypto overhaul. The keybase.io integration is a nice thought for a UX - but is an online service in invite beta really ready for being baked into an archive format?
Packing: LZMA2 is pretty good: 7z and xz already use that. For a fast algorithm, Snappy is not as good as LZ4, I understand? Neither is the last word in compression. Text/HTML/source code packs much better with a PPM-type model, like PPMd (7z has that, too, as had RAR, but removed it recently), but you need to weigh up the decompression memory usage. ZPAQ's context model mixing can pack tighter, but that's much more intensive and while I like extensibility, I don't like the ZPAQ archive format having essentially executable bytecode.
Other missing features that other archivers have: Volume splitting? Erasure coding or some other FEC? Can you do deltas? (e.g. binary software updates)
You've got some pleasant UX ideas for a command-line archiver (compared to some other command-line archivers!), but sorry, I don't think you're ready for 1.0.
Now we're going to transition to USB-C.
Just checked, if you look at the rapidSSL CA cert, it uses SHA-1.
One advantage this could have over the above, is if you can open any file at random, as with the above scheme, you might have to linearly decrypt and decompress the entire archive up to that file.
(A counter point: while each file could be compress and encrypted, there's nothing in the tar format that explicitly says so, meaning that each file would have to be probed to determine if it was compressed or encrypted)
4Q uses CBC and the crypto lib doesn't seem to support random access, unless you manually divide your file into separately encrypted streams.