The .zip file specification is flawed
github.com
github.com
I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal...
> Be liberal in what you accept, and conservative in what you send
There are several analyses of this maxim and whether it's a good choice for designing robust, secure systems. This particular Internet Draft doesn't agree [2].
[1] https://tools.ietf.org/html/rfc1122#section-1.2.2 [2] https://tools.ietf.org/html/draft-thomson-postel-was-wrong-0...
For M2M (machine to machine): Be strict in what you accept, and strict in what you send.
That's my take on it now.
For instance, a user will copy/paste URLs to his browser, there will probably be a space too much before or after. It's okay to clean it (and a better experience than a "site not found").
When a machine sends something "weird". Well, it's not possible to know if it's really wrong and it can't be corrected in any meaningful way. Just fail and throw an error so the developer can fix it.
I'm a fan of the strict solution as far as that goes, but there's a reason XHTML failed. Somehow asking people who write web sites to do it right or else it doesn't work is a big deal.
Well, when the job is to accept a hugely complex flexible poorly defined input format from 30 years ago that is written by millions of people who have no clue what they are doing, fixing common errors and getting anything to render is part of the spec ^^
And... oh wait!
I said "Be liberal in what you accept from the user" when it comes to user input. Web pages are user input, so yeah, browsers are no exception to the rules, they're actually a perfect example! :D
Obviously, your field of work dictates a lot here. And, if you are accepting something that has severe consequences on acting out, then yes, be more strict. However, the general principal holds. In general.
What you call a human mistake should be called a bug. The only acceptable way to handle a bug is to fix the code.
Not to be dismissive but the way you think is a classic beginner mistake. It's not the responsibility of other software to guess what bugs you'll put in yours. ^^
> Obviously, your field of work dictates a lot here.
large distributed systems, financial exchanges, trading systems, national government projects, aerospace, and even web stuff at times.
Some fields have low standards. That doesn't mean that the good practises don't apply, it just mean that people don't apply them ;)
You also have an odd misconception. It isn't having lower standards to be liberal in what you accept. It is actually harder. Much harder. And you should do it.
Easy example from finance. You shouldn't just take one currency. You should take in as many currencies as you can understand.
Does this mean to be magical? No. But you should ask, "how many different ways could this be given to me?" And you should instrument this with a marker for "unexpected."
This is a business feature request.
The "Being liberal about what you accept" is a technical guideline for protocol/format design and input processing/sanitation. Don't apply that rule to feature requests, it's not meant for that :(
It is fun to be pedantic with kids about "may I" versus "can I?" However, both work and if something can be understood, it should be understood.
So, should you build in all currencies? Not necessarily. However, you should definitely consider tagged inputs immediately, with the ability to extend the tags at some point. I'm not talking about big bang feature development, but ultimately, useful tools will have a plethora of ways they take in inputs. Because that is useful. And no design constraint should preclude this growth.
> "if you can understand it, you should act like you do"
For an unanticipated event, it's often less "if you can understand it", than "if you believe you can correctly guess the intention of an ambiguous request".HTML is a great example of that ambiguity. Take this example markup:
<h2 title="The <a href="https://en.wikipedia.org/wiki/Fall_of_the_Western_Roman_Empire">Fall of Rome</a>" class="roman_empire">The <a href="https://en.wikipedia.org/wiki/Fall_of_the_Western_Roman_Empire">Fall of Rome</a></h2>
This isn't too naive a markup error, and it's easy to see how a simple error in a CMS template could cause this output. So how should a parser treat it? It's non-conformant to the standard, but should it totally reject it? Recognise that the opening-bracket of the a tag in the title attribute is the start of the incorrect markup, and strip tags within that attribute? How should it appropriately recognise the misuse of double-quotes within a property attribute, that - according to the spec - should close the attribute?There is some merit in strict behaviour, rejecting non-conformant input. Especially when it comes to computer-to-computer API design, where rejecting non-conformant input may be better than attempting to interpret ambiguous input. Limitations can be an arrow to direct us back on the best path.
My argument here would be to do your best to not create ambiguity in what can be fed to you. The currency one is a great example. Never ever make something that takes in untagged integers as a value. It could mean too many things and you would have no method of knowing what was intended. But, once you start accepting strings, "$4" or "4 dollars", or "Four dollars" should all be on the table. (Though, yes, in this case probably best to stick with "4 USD" and complain about "dollars" being ambiguous without another qualifier.)
I'm not sure how this fits in with Postel's law, but your comment jogged my memory of it :)
The .zip spec is flawed, but…
I've never had an issue with extracting or using a .zip file
I'd suggest this is thanks to most software following (to some extent) the robustness principal: be conservative in what you do, be liberal in what you accept.Most of us will typically encounter fairly well packaged, conforming zip files. Occasionally we may come across something unanticipated - like this example, where HTML content is accidentally appended to the end of a zip file - and I suspect this is where we will find ambiguous behaviour: some package tools may crash, others might "extract" it as though it were content, others might ignore it.
It's this area of ambiguity that lends itself to vulnerabilities and attacks.
Re: physical engineering, I'd recommend a great book: "To Engineer Is Human". It talks about the evolution of engineering, which is a surprising amount of trial and error, with emphasis on the error.
IMHO it's a reflection of the software developers involved. The best tools, the ones we turn to time and time again, generally just work.
> Surprisingly, while this format is very common, it has never been formally documented. [1]
No. Excel is wrong when it comes to CSV.
Paste a Unicode string into Excel. e.g. Beijing in Simplified Chinese (北京市). Now Save As Windows CSV as beijing.csv. Close the file. Open beijing.csv. The cell now reads `___` (on Excel for Mac 2011 - maybe they deigned to fix it).
Excel just outputs bad data.
Once you add that, excel will open the file with utf8 encoding (if you use the utf8 byte order mark obviously). I haven't tried with other utf-* encodings.
Again don't know how to tell excel how to add that though :/ I've only had to deal with arabic in generated csv's.
All applications and operating systems should assume files WITHOUT a BOM are either ASCII, or the superset there of, UTF-8.
Please remember to say WHY you disagree if you do.
If you have a file without a BOM, you have to pick one.
As every 8 bit combination is an ASCII character of some kind, you can interpret every UTF-8 character as a combination of ASCII characters. And what you output will be different to what was input (unless you restrict yourself to single byte UTF-9 characters).
Without some other way of indicating the encoding format of a file, a BOM can be a tool to indicate "It's probably encoded using UTF-X".
However, the question "which?" can still apply. There are many encodings that are a superset of ASCII. UTF-8 is a superset of ASCII, but so are ISO-8859-X (for any "X"), Windows-1252, and many others.
When I've had problems in the past with this it's been around windows machines, which love their own encoding formats.
If it's merely ASCII, it doesn't matter. Nearly every charset contains all valid ASCII texts as a strict subset. UTF-7, UTF-16, UTF-32, and EBCDIC are the major counterexamples, and UTF-7 and EBCDIC aren't going to come up unless you're actually expecting them to. (Technically, ISO-2022 charsets can introduce non-ASCII characters without use of high bit characters, since they switch modes using the ASCII ESC character as part of the sequence. In practice, ESC isn't going to come up in most ASCII text and ISO-2022-JP (the only one still in major use) will frequently use 8-bit characters anyways).
The only useful purpose of a BOM is to distinguish between UTF-16LE and UTF-16BE, and even then it's discouraged in favor of actually maintaining labels (or not storing in UTF-16 in the first place). You can detect UTF-8 in practice without a BOM quite easily, and it's only Microsoft who feels obliged to need them.
I don't suppose I should get my hopes up too much that that option is going to be more prominent than the terrible default of saving as Mac OS Roman, right? (Whoever decided that Excel on OS X should export CSVs in an obsolete encoding for Mac OS Classic must have been trying to hurt Mac users.)
Unfortunately, to open the UTF-8 file again, it doesn't work ("北京市"). You have to make a new workbook and then import the csv in, specifying UTF-8 csv. It's pretty messed up but at least it's possible now.
http://www.slate.com/blogs/the_eye/2014/04/17/the_citicorp_t...
An undergraduate doing a class project on the building uncovered the flaw.
Amazing story.
Those dualisms can usually be resolved if you realize the vast and complex efforts that go into working around all the spec bugs - in this case, the various "find the magic number" heuristics.
A few things that I've hit in the wild that screwed up zips:
* Viruses/Worms that were rather primitive and start adding weird bytes in all kinds of places in all kinds of files
* Incomplete transmission over a modem. Depending on the protocol, you might even have most of the data, so you could read part of zip header, but not the actual archive or vice-versa. Normally, you knew right away the file was incomplete with a CRC check so not the worst problem, but the check itself was slow on old hardware.
* Weird things that added metadata and screwed with byte order, ends of file marks, etc. For instance there were some early attempts in the BBS scene to add metadata formats similar to what ID3 is to mp3s. Sometimes the writers would ruin the original file. I hit a few cases of strange attempts at steganography with software pirates trying to be "3l337" or whatever.
* Tools people tried to use to fix broken zips, that didn't quite fix them how they thought.
* Not really the fault of zip, but I've seen people rename a zip's extension to another archive format, causing the unarchiver to assume the format based on the extension. Never trust an extension if you can help it. (ex: rename a zip to rar, arc, lzz/lha, tar whatever)
* Floppy and hard disk repair programs. Sometimes these things would end up corrupting the bytes of files instead of fixing them when they tried to be more clever than moving things around. Sometimes moving things around also would result in things being ordered wrong for whatever reason in these programs. Some of the DOS Norton/Fastback/etc. ilk were especially frequent offenders.
In theory, the .zip spec is broken. In practice, it's the most reliable format for transferring a group of files. (And don't even get me started on JPEG, where iirc the file format wasn't even specified until after JPEG files had been popular for years.)
> I wonder if this same pattern can be applied to physical engineering disciplines- e.g. a structural engineer assessing a bridge and finding numerous faults in the design, despite traffic still using the bridge as normal...
I would be amazed if this weren't the case. I know I've encountered a few cases where mechanical engineers had cocked up the design but the resulting machine still managed to limp along and mostly perform its function.
For those who speak only English, yes.
The encoding of filenames have always been a mess for zip files.
Oh yes. Basically, EVERYTHING is flawed in INFINITE ways.
Or to put it differently, perfection doesn't exist. Luckily for us, the barrier for "being practical and useful" is a lot lower than perfection =)
if (commentLength !== expectedCommentLength) {
return callback(new Error("invalid comment length. expected: " + expectedCommentLength + ". found: " + commentLength));
}
Why not just determine this isn't the record, keep looking, and handle the error after the loop? if (commentLength !== expectedCommentLength) { continue }
[0] https://github.com/thejoshwolfe/yauzl/blob/master/index.js#L...I think it's the programming equivalent of not telling someone they have some food on their face - yeah, it's no skin off your back if you don't tell them, but who knows who they're going to run into later.
That said, a hybrid approach which fails by default but then offers a way to "force unzip" the file in a non-standards-compliant way might be the best of both worlds here, because the users get an immediate way to fix the problem and the emitters are likely to be notified that they are creating malformed zips.
[0] https://github.com/thejoshwolfe/yauzl/blob/master/index.js#L...
- A valid zip does not begin with, as you would normally expect, a magic number.
- A valid zip can contain arbitrary prepended data.
- A valid zip can contain sections with no identifier value.
- A valid zip can contain arbitrary data between sections.
- A valid zip is validated starting with a tail section located at the end.
- A valid zip can have a valid tail section that contains a valid tail section.
- A valid zip can contain arbitrary appended data.
I don't get it: Why design a format that's so hard to parse? Implementing a single-pass streaming parser is impossible. It should be a basic requirement for most file formats. /usr/bin/unzip cannot even extract from standard input. I'm sure the implementer didn't feel like receiving user complaints about exhausted memory.
To write the directory at the beginning they'd need to keep all the data in memory or do some extra postprocessing that would make it prohibitively slow on a 4.7MHz machine with a 20MD hard disk and a slow 360K floppy disk drive. So the directory goes at the end.
The files do not begin with a magic number because ZIP files can be embedded in other files - most commonly executable files for the self-extracting feature of PKZIP that placed the ZIP file right after the decompressor. This allowed the executable to be used both by itself and by PKZIP.
As for the data after the magic number, it was used for the comment which was most likely added at a later point and they decided to put it after the magic number so that it remains compatible with existing decompressors. In the 80s and early 90s people didn't had Internet to get the latest version and a lot of old versions of PKUNZIP were floating around for many years.
Things got further complicated when RARs made the scene because I remember that some of the early .RAR apps offered their own implementations of .zip support that would create multi-span .zips where the installer data was on the first disk, the .zip data was on the other disks, but the container spanned all the disks.
This meant that you could potentially have 3 different ways of dealing with the same data and there was no safe way to assume which method was used:
1. Installer on disk 1, multi-spanned .zip on disks 2-n 2. Multi-span .zip on disks 1-n 3. Multi-span .zip on disks 1-n but installer data only on disk 1
http://stackoverflow.com/a/12393597/60910
Therefore the straightforward way to parse a zip file is to proceed from the beginning and parse out each file sequentially. The End of Central Directory record is then only a redundant convenience to avoid sequentially scanning files e.g. in large zips for random access.
On second thought, not entirely redundant, as the central directory does contain the file permissions. But those can be parsed and set after file extraction, without increasing the overall memory complexity.
It is annoying that support for these non-standard zips has become de-facto standard (you don't want your tool to fail where others succeed)...
Here's a short-list off the top of my head of tech with specs and horrific offenders in the wild, including many popular implementors.
* Telnet
* VT-XXX, i.e. VT-100 or any term spec for that matter
* MIDI
* MKV, AVI, etc.
* GIF, JPEG
* A huge amount of tech in any browser
* SAUCE
* ANSI escape parsers/writers/sequences
* LZH/LHA
* Open Document Format
* XMP
* MAPI
* POP, IMAP
As I think of this, I realize I'll be typing all night. In some cases, the implementors all adopt things from the most popular implementors even if they are historical. In other cases, things just don't work or break in awesome ways. Zip bombs were mentioned, but there are plenty of awesome things.
I'll finish by saying I wonder if many authors of anything related to terminals even ever bother to read specs or care. Over the years I've started to wonder how some of these things even work in the wild, but the answer is most people don't notice what actually isn't working, while anyone who needs to care has to draw from what become unofficial new specs.
> D.1 The ZIP format has historically supported only the original IBM PC character encoding set, commonly referred to as IBM Code Page 437. This limits storing file name characters to only those within the original MS-DOS range of values and does not properly support file names in other character encoding, or languages.
Extra fields are generally used to implement other code pages. There is a direct way to do it with an extension to the GPBF. Problem is that some writers don't implement the proper flags.
There are noncompliant implementations that generate zip files with filenames in ISO 8859-1 and other 8-bit codepages, and conversely, there are decompressors that try to accommodate such malformed files.
Yes, the spec does not accommodate this. However there are conventions among popular ZIP creation tools that allow it.
If the zip spec were governed by an open body, then this is the kind of thing that would get quickly fixed. (basically standardize the common practice already in use). I don't know why PKWARE never fixed the spec. Maybe because there's no money in it for them, and the pain for the industry is really not that severe.
1. Begin linear search backwards from end of file 2. If magic constant found, push onto stack of possible locations 3. Once 32k has been reached stop searching 4. For each location on the stack, validate the rest of the header and that the remaining comment length matches the recorded length. 5. The first location to validate as a correct header is chosen
This should account for the appearance of the magic header in comments or data preceding the end of directory entry.
Have I missed something?
So I think what you missed is pedantry. :)
Step 4, specifically "the remaining comment length matches the recorded length" will fail in the presence of appended data, as there is no end of comment or end of record marker. This is exactly how the file failed in the article, it used that algorithm and found zero valid headers.
Download.php:
if (isset($POST['whatever']) write($file_contents);
print "download page contents here"
Including an ending tag (?>) followed by whitespace could have the same effect.
Tools that correctly read .ZIP archives must scan for the end of central directory record signature, and then, as appropriate, the other, indicated, central directory records. They must not scan for entries from the top of the ZIP file, because only the central directory specifies where a file chunk starts. Scanning could lead to false positives, as the format does not forbid other data to be between chunks, nor file data streams from containing such signatures.
Also, there are no tools I'm aware of that will handle that archive structure.
From the man page: "Dar can also use a sequential reading mode, in which dar acts like tar, just reading byte by byte the whole archive to know its contents and eventually extracting file at each step. In other words, the archive contents is located at both locations, all along the archive used for tar-like behavior suitable for sequential access media (tapes) and at the end for faster access, suitable for random access media (disks). ... Note also that the sequential reading mode let you extract data from a partially written archive (those that failed to complete due to a lack of disk space for example)."
> Consequently, to use the API, your program must be released under the GPL as well.
GPL is a bad thing for programs like this. It hampers usability. Especially at least the decoder should be BSD/MIT/etc.
A combined 'contain & compress' format is perhaps at odds with the unix philosophy of tools doing one thing well. Meanwhile, zip definitely exemplifies 'worse is better', in that despite the format's quirks and darker corners, it's good enough for most casual uses, and network effects and backwards compatibility help it win out.
If you just want to compress some stuff, zip is perfectly fine and widely supported. (or gzip or tar)
If you want fancy things with better compression, user rights/ACL (either linux or windows style), file dates, encryption, password protection. Well, I don't know what can do that.
An additional feature that sometimes is important as well is making itself stream friendly so you can read the metadata without the entire file (or easily before streaming) as it is streamed over some context or using some technology.
The landscape is much different now and so some of features you'd design into a new format didn't exist then, while others that were concerns are less concerns now.
As parent hints, the compatibility alone makes zip still very desirable and actual compression is good enough given size of disks today for most people. Obviously there are exceptions.
I personally also like ZPAQ for generating archives, though I find it a very poor fit for its suggested use as a backup tool (doesn't handle symlinks, owner, xattr, ACL &c.). As an added bonus the reference implementation is completely composed of public domain and MIT/X11 licensed code, which makes it friendly for linking pretty much anywhere.
https://github.com/century-arcade/src/tree/master/tools/mkiz...
(One person found it, but the whole thing was way too hard. I drastically overestimated people's operational cleverness.)
Maybe your target audience was just not quite right? People go gaga over ARGs and devote thousands of hours to 'em.
Maybe hint that your game has something to do with HalfLife 3...
Nah. It's insanely difficult to find this kind of stuff when you have no clue what to look for.
The SAUCE spec adds trailing bytes after the EOF mark for instance. I wrote a few SAUCE reader/writers embedded in various codebases (and will release a stand-alone good one some day soon).
Generally, a trick was that if you wanted some sort of metadata on a file, you could add it before the contents of the file (as in MKV allows ASCII at the start) or after the EOF marker. The latter made sure you didn't interfere with a lot of existing programs on various operating systems of the time, including but not limited to DOS, Windows, and CP/M. This trick still works for a lot of OSs or in other forms today, we just don't rely on markers for EOF as much so there are other methods, and people have started to bake metadata into specs.
As for the BBS scene, although SAUCE was used, for zips DIZ tended to be used more and was just packed into the archive and could be found using the internal TOC, or otherwise just uploaded separately.