Tar is an ill-specified format (2015)
invisible-island.net
invisible-island.net
I simply tried to follow the tar(5) man page[2], and got a reference test set from another website posted previously on HN[3].
Along the way I discovered that NetBSD pax apparently cannot handle the PAX format[3] and my parser inadvertently uncovered that git-archive was storing the checksums wrong, but nobody noticed because other tar implementations were more lax about it[4].
As the article describes (as does the man page), tar is actually a really simple format, but there are just so many variants to choose from.
Turns out, if you strive for maximum compatibility, it's easiest to stick to what GNU tar does and favor GNU extensions over PAX header fields. If you think about it, in many ways the GNU project IMO ended up doing "embrace, extend, extinguish" with Unix.
[1] https://github.com/AgentD/squashfs-tools-ng/tree/master/lib/...
[2] https://www.freebsd.org/cgi/man.cgi?query=tar&sektion=5
[3] https://mgorny.pl/articles/portability-of-tar-features.html
In fact, one of the biggest questions I've still yet to answer is whether Docker or OCI have ever clarified just which TAR format must be followed? As far as I can tell, it's a mix of "whatever library X supports"? https://www.cyphar.com/blog/post/20190121-ociv2-images-i-tar for example.
I'm not sure I want a more complicated format, but I do wish somebody would pick a new file extension for their next use of tar so we could clearly differentiate one tar format from another the next time someone builds on top of simple tar files for their binary distribution, etc.
Off the top of my head, you need to specify text encoding, line endings, a quoting mechanism in case a field contains the separator, and an escaping mechanism in case a field contains multiple lines or the quote character.
If you have those specified and everyone follows the specification, there shouldn't be any problems, right?
Woe be unto you if you have to process input data where those are not used consistently, of course.
Second, this: http://georgemauer.net/2017/10/07/csv-injection.html
And what I wrote before is basically all you need for a specification. Yes, it sucks that there is no fixed one that everyone uses.
Well, yes, but I feel that that's the point at which CSV definitely earns its "ill-specified format" hat.
Given that there is no way to extract the file specification from the file itself, you can only follow the specification if someone tells you the specification beforehand. You can even use the "wrong" specification and don't notice it because your specification overlaps with the "true" specification in the types of data you're typically seeing (looking at you, optional quotes).
You forgot the localized decimal separator symbol.
it never ends
Arguably right, practically useless.
All file formats have limits for how much they define types for the payload. Most have very limited types. JSON for example knows objects, arrays, strings, numbers, booleans and null. It's not that big a difference to have a format that has only strings.
You also need to specify the separator itself, which despite the name isn't always a comma, but can also be a semicolon or (less commonly) a tab.
https://www.iana.org/assignments/media-types/text/tab-separa...
They are there, they are convenient, and they don't require any visible character sacrifices.
'\0'
Depending on the source this might be:
- quote back-slash zero quote
- zero
- back-slash zero
- a byte with the numeric value of 0
- quote zero quote
- quote back-slash zero quote
- quote byte-with-value-0 quote
- malformed and should be ignored
Worse the same application might mix different of this quoting mechanisms for different columns or even in the same column. Normally it shouldn't but it's defintely a thing you can find in programs which don't use a proper marschalling library/code, and you know. CSV is easy so you surly don't need to add a library dependency just to use it. (sorry sarcasm).
Or with other words "because CSV is easy" it's somteimes not properly specified and sometimes no "proper" marschalling/serialization library/module is used resulting in ad-hoc fixes for quoting where needed which potentially don't follow any specification either.
Salespeople enter things like Company Foo, the One (which used to be Bar). The worst part is that often these are legal names, which means you do need to store them.
When I worked for a FAANG this was one of my major annoyances as there was no canonical number for a customer, so everyone did string matching which broke whenever a company changed their name (like for instance, if they'd just gone public).
Just say no to CSV's with commas in quotes.
It's almost like they really wanted to avoid this issue back in the 70s so it is built right into ASCII and unicode.
I wonder how you embed US/RS/GS/FS within a field though. Prefix with ESC?
As many as there are localized datetime formats :)
there shouldn't be any problems, right?
That would be reallys awesome, yes. With the aid of the other sibling comments you already figured by now it's not as simple, unfortunately.
- Values in most columns are escaped, but a few are not. (You have to count commas from both ends of the line and then log a warning and make a guess if there are too many commas.)
- ASCII whitespace is url-encoded (why?) but no other characters are.
- Columns without double quotes can only contain digits.
- Double commas are used as the field separator because they didn't want to implement quoting. (Some of the fields came from user input on a web page, and I prayed for them to hit a double comma and learn the error of their ways, but it never happened. I think their front end developers were sanitizing for them.)
- And my favorite, no quoting was supported at all, double quotes were left unquoted, commas were stripped from values when writing, and when reading, fields were compared against a list of mappings like "Acme Inc." -> "Acme, Inc." to restore the commas.
The "extend" phase from the Microsoft playbook is always something that competitors cannot duplicate. Either for legal reasons (patents) or technical reasons (having the "spec" depend on all of the Windows ecosystem).
I don't think that's what GNU did.
POSIX aside, Unix-like systems that are still around nowadays and not directly use GNU user space, have copied many of the GNU extensions out of necessity. People have become accustomed to those extensions and written programs & scripts depending on them, so you now need to implement many of the GNU extensions simply to stay compatible.
Even if not consciously done and without malicious intents, this is IMO still fairly similar to how Microsoft played it: add extensions that are convenient, so people use them, eventually rely on them and become a de facto standard.
Any C compiler vendor (even commercial ones) does language extensions, mob => pushing the language forward.
Microsoft does C compiler with language extensions, mob => get your pitchforks.
On purpose and on accident: when Debian switched its /bin/sh from Bash to dash all sorts of things broke (for us) because GNUism were leaking into what should have been a standard language/feature set.
We had to tell our users: change your code or change your shebang.
That reads like a lack of interest to update POSIX to catch up with the de-facto sstandard
I mean, I doubt that GNU would try to block the standardization process to avoid including their extensions.
GNU licenses are very much embrace-extend-extinguish licenses.
What makes open source software "open source" is the source code, which is subject to copyright law. Ideas, protocols, and interfaces are not subject to copyright; only their expressions are subject to copyright.
The Embrace, Extend, Extinguish approach makes things incompatible with the way others do them to cement hold on a large user base that prevents those users from changing their minds later. Sometimes that takes the form of forcing people to reimplement everything in a cleanroom setting, because they can never keep up that way; that's where the GPL excels within the open source world for EEE. Other times, it's by diverging from norms in ways that others cannot follow without breaking their own ecosystems, which is where the GNU project has excelled with its plethora of subtle incompatibilities.
> Tools that correctly read ZIP archives must scan for the end of central directory record signature, and then, as appropriate, the other, indicated, central directory records. They must not scan for entries from the top of the ZIP file, because (as previously mentioned in this section) only the central directory specifies where a file chunk starts and that it has not been deleted. Scanning could lead to false positives, as the format does not forbid other data to be between chunks, nor file data streams from containing such signatures
You might think that ZIP files would have a header, or at least a fixed-width footer, but they don't. Instead you're supposed to scan backward through the file looking for a magic number indicating you've found the central directory. It's amazing to me that this is the format that won the compression wars and stuck around to this day.
The zip spec says anything can come in before the "zip data". Info-Zip gives this example
unzip unz552x3.exe unzipsfx.exe // extract the DOS SFX stub
cat unzipsfx.exe yourzip.zip > yourDOSzip.exe // create the SFX archive
zip -A yourDOSzip.exe // fix up internal offsets
http://infozip.sourceforge.net/FAQ.html#unixSFXIn other words, append any data you want, then fix the pointers in the central directory.
There are tons of various zip utilities that will mess up depending on what you put in that self extractor including MacOS's built in finder based zip support.
PKWare, the keepers of the documentation for the ZIP format refuse to specify that you must read the end of central directory. They do this by claiming you can stream a zip file. But you can't stream a zip file if you have to read the central directory because data you may have already used, "the definition of streaming", maybe not be listed in the central directory in which case you shouldn't have used it.
In other words imagine each file in the zip is a compressed script and you are going to execute them. Imagine there are 2 scripts, the first one is "rm -rf /" and the second "ls /". The central directory only points to the 2nd script but if you stream you'll read the first script an execute it.
The ZIP spec claims you can stream but clearly you can't.
ZIP files do have per file headers, which duplicate information in the directory, such as the file name. And malicious files have used this quirk to circumvent safeguards. See, e.g., WinRAR filename spoofing exploit: https://www.rapid7.com/db/modules/exploit/windows/fileformat...
Any time metadata is duplicated in a system (specifically, an interface or protocol), alarm bells go off for me. Simply duplicating metadata can inadvertently create an explosion of failure cases and implementation complexity. Conversely, avoiding duplication of data is an example of failure-proof design--if it's not duplicated, you don't have to worry about inconsistencies and reconciliation rules.
I can't agree with this.
The reason you deal with inconsistencies and reconciliation rules is because, despite your efforts to the contrary, the world will give you inconsistent or incomplete copies of your data from time to time.
A sure-fire way to ensure that your data is lost is to only keep one copy of it. At some place in your storage stack, you want some level of redundancy. The general idea for an archive format is that it is somewhat self-synchronizing, and somewhat possible to repair. All media has a non-zero error rate. And then there are all of the various reasons (software, human error, disk full, etc.) that a file might get accidentally truncated (which destroys the central directory in a zip file).
You may not personally have to deal with the consequences of damaged media or similar types of failures, but for those that do, it's nice to have a data format where a stray error or two doesn't render the entire archive useless. If you are building on top of "reliable" storage systems (like cloud storage providers), you're just pushing the problem to lower levels in the stack... but there is less context at lower levels in the stack, which means recovery can't be very sophisticated.
If you are writing a package manager or backup program you might make a different choice, but how often do you find yourself writing a package manager?
You might be very nervous to learn that most filesystems duplicate some metadata in one way or another. FAT, NTFS MFT, ext4 superblocks and group descriptors come to my mind. In some level you ought to have duplicates; it's a matter of where to place them. ZIP's choice made sense at the time of smaller and slower disks. (I do think that the modern standard ZIP should make some fields scrubbed to prevent compatibility problems, though.)
The point was you could create a multi-disk zip (zip came out when we had 1.4meg floppy disks). If you zip up 10meg across 7 disks and you want to update the README of your 7 disks, by putting the central directory at the end you just ask the user to insert disk 7, read the central directory, append the new README, write a new central directory. With the directory at the start you'd have to re-write all 7 disks. With it at the end you only have to re-write disk 7 and only a small portion of the file.
The only bad part of that particular feature in the zip implementation is the scanning part. The variable sized underspecified "Zip comment" that you have to scan through should have either come before the central directory OR there should have been a length or pointer to directory after it so no scanning needed.
startxref
12345
%%EOF
where 12345 is the file offset of the TOC.Usually when copying a game...
Yes. ZIP has no header, it has a footer.
What gets really messy is that the footer contains a variable length comment field, with the length stored before the field, so good look finding where the footer starts.
And here we are just talking a about a correctly generated file. This also opens up the door to all kinds of intentional abuse.
And of course googling zip file steganography reveals this is not a new idea: https://aroralakshay2014.medium.com/hide-it-lest-find-it-zip...
The header at the end actually helped it a lot. It allowed the self-unpacking archives to exist.
I still call him gnu as he of course calls me gumby.
ustar tar files are compatible with almost all tar archivers, and are in fact well-specified.
Im sure if you're deeper down the stack this ends up being one of those nuanced peas under a mattress, but Im guessing that 99% of people who use tar do not run into issues using tar.
http://tiamat.name/blogposts/fast-appending-files-to-tar-arc...
Um, yes. A tar file is just a blob of 512 byte aligned files with 512 byte headers in front of them + an optional 1k null-bytes signifying the end. Technically all there is to appending at the end is to slap on another header and a file blob?
EDIT: Ok, on second thought, you do need a linear scan. There could be ancillary data after the 1k termination blob. Even if not an intentional polyglot nonsense, but if you actually use tar on a main frame tape drive as originally intended, simply seeking to the end doesn't work.
That's when you realize how many cases you actually have to handle.
For reasons known only to themselves, Docker originally used different implementations for the docker command-line tool and the docker-compose tool. This means that cache entries produced by two tools that you got from the same place don't match up, and you can't use one tool to cache a layer and have the other pick it up.
At least, not by default - if you know about it, there's a switch you can flip to have one use the other's implementation. But they had to add that in, and do that extra work, because tar isn't well-specified.
It helps that we're only ever appending gzipped tar fragments, and we have a tracking database format to know where the offsets in the uncompressed stream are:
https://fastmail.blog/2015/12/21/building-a-backup-system-fo...
But yeah - much worse if you're parsing arbitrary tar files than if you create everything yourself and just have to make it readable by gnu tar and that's about it.
Personally I wish sqlar didn't have this limitation since SQLite is awesome for tooling.
Most of the PDP era DEC machines are designed favoring an octal encoding. The PDP-8 was a 12 bit machine (12=4 * 3). They also had few 18 (=6 * 3) and 36 bit (=12 * 3) machines.
The PDP-11 being a 16 bit machine with 8 bit granular memory access was an odd man out, but still had many aspects designed around splitting words into groups of 3 bits, e.g. if you look at the instruction set encoding. Or just take a look at the front panel[1].
Because of the DEC machines that early Unix was developed on (and the people involved in early Unix development being very familiar with them), octal encoding crept into many places where it stayed until today, including the C programing language and some Unix specific file formats.
Using octal, stored as plain text ASCII makes it very easy for a human to debug, for whom reading octal is second hand nature.
cat chapter*.mpg >movie.mpg
Are there any file archive formats that support this?It turns out that GNU tar supports extraction of concatenated tarballs, but it requires the --ignore-zeros (-i) option:
> Normally, tar stops reading when it encounters a block of zeros between file entries (which usually indicates the end of the archive). ‘--ignore-zeros’ (‘-i’) allows tar to completely read an archive which contains a block of zeros before the end (i.e., a damaged archive, or one that was created by concatenating several archives together).
-- https://www.gnu.org/software/tar/manual/html_node/Ignore-Zer...
(As a bonus, this also works for concatenated .tar.gz files because gzip supports concatenation.)
Regarding other formats that support concatenation, the Linux kernel's initramfs format is interesting because it is based on CPIO but is explicitly defined to be the concatenation of CPIO archives:
https://www.kernel.org/doc/html/latest/driver-api/early-user...
CPIO itself doesn't support concatenation. There is a neat hack to extract such concatenated archives:
while cpio -i; do :; done < archive.cpio
(source: https://unix.stackexchange.com/a/266090; also mentioned there is the tool skipcpio, which is part of dracut: https://github.com/dracutdevs/dracut/blob/master/skipcpio/sk...)This hack would also work for extracting concatenating tarballs (without GNU tar's --ignore-zeros option). One annoyance is that it shows a warning when it gets to the end of the archive.
You sometimes get transports streams named ".mpg", which complicates everything.
Tar is tar. Let it be.
tar -czvf if you’re gzzy.
tar+gzip beats zip for most uses. I don’t know about whether 7z beats tar+gz, but I don’t think 7z is as ubiquitous, though that may not be true now.
But I don't know how it got to the wares/hacker scene from there. I assume they will also have dropped the 7 bit thing as it would reduce efficiency by 1/8th on modern systems.
2.00 Beta 1 1999-01-02 - Original beta version.
I'm pretty sure its popularity increased because it was one of the few options on Windows that was completely free to install with no nag screens, ever. It had a decent right-click menu; excellent usability for its time. And you wouldn't be tempted to use an older copy such as whatever shipped with Windows XP to extract archives. Plus it offered compatibility with all other formats including its own, and eventually apps like Firefox publicly shipped with 7zip Self-Extractors and such. And if you wanted to test whether 7zip really was more efficient, you could quickly try compressing files using different options - in my personal testing, 7zip often won back then. Not sure if results have changed recently given there are a few new compression techniques these days.
GNU tar has -t and --list, while NetBSD, OpenBSD, and FreeBSD all have -t. None of them extract the archive to list the contents as far as I am aware.
As a user, I can use tar and an outside compression tool together in one command to see the file list.
tar -Jtf foo.tar.bzip
tar -xtf example.tar.gz
tar -I /usr/bin/mydecompressor -tf somethingelse.tar.myc
So unless there's some huge, huge archive out there that takes forever to decompress and read it's not really an issue for the user.So why are we still using it if ~99% of Linux users and administrators have never seen a digital tape drive? I've built quite a number of Linux, BSD and Windows servers yet the only tape with digital data I have seen in my entire life was for a Spectrum. Do we really need our files stored in a way compatible with tape storage?
> As a user, I can use tar and an outside compression tool together in one command to see the file list.
Which would decompress the whole archive, the seek through it. Ridiculous waste with no benefit.
Tar is the ultimate idiosyncrasy of of Linux.
I would question this assertion. Some of us have been doing this a while, you know.
> Do we really need
Probably not.
> no benefit
Benefits include being able to use the same dictionary across the whole archive, being able to leverage multiple improvements in multiple outside compression programs, backwards compatibility going back decades, and the ability to use a tool that's standard not just on Linux but BSD, macOS, Solaris, and any POSIX compliant system.
> ultimate idiosyncrasy of of Linux
Are you certain that's not systemd? Or the way ptys are handled? Or its sound and video systems? Or the /proc file system? Or devd? Maybe the multiple families of different incompatible package managers? I mean tar isn't originally from Linux or even GNU. It sure beat shar files.
If you want your files inside your archive compressed before being added to the file, you can do that too with tar, but other tools do it for you. You're absolutely free to use those other tools.
Yes, because that's the only famous oddity which is actually annoying (sound and video systems just work and do everything a user might want of them) and has no real reason to exist on a today computer.
> If you want your files inside your archive compressed before being added to the file
No, I want to see a list of files inside without decompressing anything and without reading the whole archive.
Being able to extract a particular file could be a very nice bonus - using the same dictionary for all of them doesn't require making this impossible.
Your mom was ill-specified.
However, could you please stop creating accounts for every few comments you post? We ban accounts that do that. This is in the site guidelines: https://news.ycombinator.com/newsguidelines.html.
You needn't use your real name, of course, but for HN to be a community, users need some identity for other users to relate to. Otherwise we may as well have no usernames and no community, and that would be a different kind of forum. https://hn.algolia.com/?sort=byDate&dateRange=all&type=comme...
...it’s not at all obvious to me...