Zip – How not to design a file format
games.greggman.com
games.greggman.com
Edit: For anyone seeing this much later in time, the website got flooded when it hit the Hacker News main page. I grabbed the archive mirror of it from earlier today.
Saving things to the Wayback Machine at the Internet Archive is a far superior solution than IPFS. IIRC, IPFS only keeps things as long as they're popular, which has famously resulted in things like most NFTs pointing to dead links.
It’s the best of both worlds - IPFS will still do it’s decentralized thing but with a pinning service even when individual IPFS nodes do garbage collection or otherwise expire an object locally they can always go fetch it from the persistent copy stored with the pining service.
I'd prefer something like archive.is for longer-term storage.
> This is undefined by the spec.
> There are 2 obvious ways.
> 1. Scan from the front, when you see an id for a record do the appropriate thing.
> 2. Scan from the back, find the end-of-central-directory-record and then use it to read through the central directory, only looking at things the central directory references.
I was recently bitten by this at work. I got a zip from someone and couldn't find inside the files that were supposed to be there. I asked a colleague, and they sent me a screenshot showing that the files were there, and that they didn't see the set of files that I saw. I listed the content of the zip using the "unzip -l" command. They used the engrampa GUI. At that point I looked at the hexdump of the file. What caught my eye was that I saw the zip magic number near the end of the zip, which was odd. The magic number was also present at the beginning of the file. At this point I suspected that someone used cat(1) to concatenate two zips together. I checked it with dd(1), extracting the sequence of bytes before the second occurrence of the zip magic number and the remainder into two separate files. And sure enough at that point both "unzip -l" and "engrampa" showed the same set of files, and both could show both zips correctly. Turns out engrampa was reading the file forwards, whereas unzip was reading the file backwards.
You mean like the GIFAR attack (https://en.wikipedia.org/wiki/GIFAR)?
This is also why traditionally zip are read from backwards.
> The polyglot file pocorgtfo10.pdf is valid as a PDF, as a ZIP file, and as an LSMV recording of a Tool Assisted Speedrun (TAS) that exploits Pokémon Red in a Super GameBoy on a Super NES. The result of the exploit is a chat room that plays the text of PoCkGTFO 10:3.
[0]: https://www.lexaloffle.com/pico-8.php
[1]: https://www.lexaloffle.com/bbs/?cat=7&carts_tab=1#sub=2&mode...
The most iconic /i/ kit spread this way being Dangerous Kitten (https://web.archive.org/web/20120322025258/http://partyvan.i...)
No need to wait. Microsoft does this in their file formats. (E.g., PowerBI's pbix files have this pattern.)
Was the second zip less than 64KB in size? If not then engrampa was being extra wrong about looking for the header.
* Since the signature is put somewhwere between zip headers instead of separate file there needs to be little bit of parsing to locate it and the hash isn't straight forward hash of whole file because it needs to exclude the signature part.
edit: turns out they later also added a version of external signature
Per the blog + spec:
A ZIP file MUST have only one "end of central directory record".
You couldn't do a "find XXX | grep YYY | zip" pipeline; you would instead often add files to an archive one-by-one. So the format consists of appendable records sharing the same structure, and adding new files is simply a matter of appending records to the end.
Now sometimes you would want to quickly locate and extract just one file out of the entire archive, and scanning the entire archive would be ridiculously slow on a system with <1MB/s throughput a some kilobytes of RAM. Hence, the central directory at the end of the archive.
Some other time, the power would go down while you are appending to the archive, or a bad block would pop up on your HDD, hence redundancy between the central directory and the records.
Also, writing archivers back then was very different from what we do now. There were no unit tests, no abstraction layers, and very limited debuggers. Software was hand-coded in assembly, and adding an extra field or a check here or there was much less trivial than it is now.
You can always find something that was designed under completely different constraints and call it a bad design. But in reality it only shows that the author hasn't done enough research on the topic.
The format was developed in 1989, at which time find(1) and grep(1) and command lines pipes had been around for at least 10 years (more, depending on how you measure it).
It wasn't that you couldn't run pipelines like this, but more that zip's creators didn't care about Unix.
> Software was hand-coded in assembly
I'm not sure what version of software history you've been reading, but I can assure you that I personally was not hand-coding assembly in 1989, even though I was writing device drivers, PostScript RIPs and other stuff at that time.
PKZIP, which originated the Zip format, was developed on MS-DOS, which did not have find(1) and grep(1), nor did it have pipes (IIRC):
* https://en.wikipedia.org/wiki/PKZIP
It was a third-party who created utilities that ran under Unix-y systems:
* https://en.wikipedia.org/wiki/Info-ZIP
Unix had its own file archive formats at the time, but it seems that folks on the PC side of things either didn't know about them, or didn't care for them:
* https://en.wikipedia.org/wiki/Tar_(computing)
I did note that likely what mattered the most was that pkzip's creators didn't care about (or maybe even know about) Unix.
COMMAND.COM supported find(1) in the form of DIR /S and grep(1) in the form of FIND. It also supported Unix-like pipe syntax, though because DOS wasn't multitasking it was implemented under the hood by redirecting to a temporary file and running the programs sequentially.
None at all? 80's C compilers were garbage. Most example code for device drivers (sound cards, video cards, etc) was in ASM back then.
I've written on the order of a half dozen device drivers in my life for sound cards, motor controllers, DSP boards and sensors, and have never used asm for any of it.
Also, note that in 1989, Xenix and others ran quite happily on the sort of hardware you're citing, so MS-DOS is/was the actual "limiting factor".
Remember that the 'rest of the computing world' didn't really catch up to the early UNIX workstations until Windows 2000/XP.
I was. You missed out on some fun.
That's how tar works. Just the other day I was making a backup and streamed 5gb of data as a tar archive, opening it in the archiver took a lot of time.
> That's how tar works. Just the other day I was making a backup and streamed 5gb of data as a tar archive, opening it in the archiver took a lot of time.
That's because tar is the tape archiver, specifically written for a medium whose very nature is linear reading/writing and where skipping around was highly impractical or even impossible.
In fact that was one of the reasons ARJ got more love than zip, given its flexibility and ability to split archives across floppies.
Unless you were coding COM files, demoscene or commercial games, there was enough software being written in Turbo and Quick Basic, Turbo/Quick/TMT Pascal, Clipper/FoxPro/DBase III, Turbo/Quick C/C++, Turbo/TMT Modula-2,....
Assuming you're commenting on the GP's point about UNIX pipelines, I think your comment here is a little disingenuous. Akin to that famous HN comment mocking the then start up Dropbox saying "you can just use FTP...." while completely missing the reason Dropbox would then go on to be successful for.
> In fact that was one of the reasons ARJ got more love than zip, given its flexibility and ability to split archives across floppies.
Did ARJ actually get more love than ZIP or is that just your anecdotal recall of the era. I seem to recall the opposite with ZIP being far more prevolent. Remember ZIP also supported splitting archives over multiple floppies.
Interestingly ARJ's website is still live and not been updated in 10 years[1]. It might be ugly by today's standards but man is it better designed for getting the core information out. We've definitely lost something when sites became more about presentation and less about information.
I think it wasn’t so much about the qualities about the zip format itself rather that WinZip software wasn’t as good compared to WinRAR. ARJ was command line, at least that is how I used it, it was well documented and easy to use.
It may have been that the zip format was more common in the DOS 5 1/4 inch floppy days because zip (1989) existed prior to both both ARJ (1992) and RAR (1993)
The advent of Windows was when ZIP started to push them out.
Back in the MS-DOS days, you couldn't use zip across floppies unless you bought the commercial Pkzip version, the shareware one did not offer it.
Whereas ARJ offered all what pkzip was capable of, for free, and every high school kid with PC at home got it on their floppies collection.
I'm not sure I understand the Dropbox analogy. In what way was 4DOS "the Dropbox of MS-DOS shell utilities"?
But...people don't write Dropbox for themselves? That's an existing product that someone else made (unless "you" refers to a Dropbox employee of course).
It literally is disingenuous for the reasons I'd already outlined. And then saying "but you could always install this other lesser known DOS-like command shell" doesn't exactly help your claim.
Sure, with a little effort pretty much anything is possible in IT (that's a large part of the reason I was inspired to work in IT growing up). But my point was if MS-DOS requires a new command shell or other hacks then your whole argument of it supporting pipelines is rather stretched. Literally the whole point of pipelines is as a quick and common way of passing information between programs and nothing you've details is quick nor common.
But you'll come along and find some other tedious argument because that is what you do....
> Back in the MS-DOS days, you couldn't use zip across floppies unless you bought the commercial Pkzip version, the shareware one did not offer it.
But the file format did.
> Whereas ARJ offered all what pkzip was capable of, for free, and every high school kid with PC at home got it on their floppies collection.
Every high school kid would have had the commercial version of pkzip too -- albeit not legally. The amount of pirated software that would get swapped around in the 80s and 90s was pretty ridiculous.
In fairness to you though, I guess the preference of ARJ vs ZIP is always going to be an unprovable subjective point. I can't go back in time to prove your community of friends used one over the other any more than your anecdotal evidence being proof of my group's, or any others, preferences. And to be fair, it was a long time ago now and memory is fallible. I remember using pkzip a lot to package software I was writing but I might just have forgotten about all those times I used ARJ.
> You might think this is nonsense but you have to remember, pkzip comes from the era of floppy disks. Reading an entire zip file's contents and writing out a brand new zip file could be an extremely slow process. In both cases, the ability to delete a file just by updating the central directory, or to add a file by reading the existing central directory, appending the new data, then writing a new central directory, is a desirable feature. This would be especially true if you had a zip file that spanned multiple floppy disks; something that was common in 1989. You'd like to be able to update a README.TXT in your zip file without having to re-write multiple floppies.
For example, the need to be able to skip unknown blocks is fairly important and imcompatibility of parsers was a known technical problem in the early 1980s already.
I think the author is aware, at the time that PkZIP was created he was writing NES games and using data compression to fit more content into the games.
There were various "archive" ZIP files that had grown to several gigabytes each. The developer was using Info-ZIP "zip.exe" to "append" files to the archives multiple times per hour (>100). Like virtually every ZIP implementation, Info-ZIP just re-writes the entire file when appending. As a result, I was seeing multiple terabytes of I/O churn each day as new ZIP files were being created to replace the old ones (which was killing VM replication across a 500Mbps pipe, and is what drew my attention to begin with).
While I would have loved to find a ZIP implementation that just appended the file and a new central directory I found it easier to just rename the ZIP files every day to minimize the churn. Still, I can't believe there isn't a ZIP implementation out there that can do this.
Do GZip files have bit-based lengths?
But the answer in this case is Huffman coding. The most common sequences compress to just a few bits, but backreferences and such are not byte-aligned, they just follow literals in the bitstream. Aligning things to bytes would significantly reduce the compression ratio.
Where this becomes a problem is when someone thinks "I'll compress my logs as I write them out". The next problem is what to do on restart. It's tempting to think "I can just append to the existing file because I can concatenate two gzip files". But if the restart is after a crash rather than a clean exit you get log salad.
> -g, --grow: Grow (append to) the specified zip archive, instead of creating a new one. If this operation fails, zip attempts to restore the archive to its original state. If the restoration fails, the archive might become corrupted. This option is ignored when there's no existing archive or when at least one archive member must be updated or deleted.
I tested this on Linux and it appears to work as advertised. I assume the reason why it is not enabled by default is because "If the restoration fails, the archive might become corrupted".
edit: python zipfile in append mode also appears to avoid rewriting the zip file. libarchive doesn't seem to support appending to zip files.
I didn't see this option on the version of Info-ZIP they were using but that would definitely do what I wanted. I thought I looked at current man pages but I guess I didn't.
Also I'd preferably use zstd because it's a significant improvement in many ways and would afford you the option of placing a custom compression dictionary at the front of the tar archive.
When Jean-loup Gailly and Mark Adler started infozip a lot of discussion was done to try to guess what the future would hold, but it was just that. Guessing.
The fact we are still using it all is a major miracle.
But I think criticizing the 30-year-old 'design' is pointless. So this is a good article with a bad title.
What is much more helpful is suggesting, in detail, how Zip might be extended to remain backwards compatible while moving forward to adopt a more sensible future format. (Though 30 years later, some guy is going to write an article saying how terrible that new design is).
The 80s BBS scene was built around things like ARC, LHarc, and ZOO. ARC essentially got killed by its creators SEA thanks to their suing PKWARE over PKARC. Many BBS operators banned the use of ARC after that lawsuit in protest (a similar furor would erupt 5-6 years later over GIF).
It was ultimately a private person who founded a small company who wrote a tool for DOS and distributed it as shareware. There were other tools like it (ARJ and LHA were other popular ones). PKZIP disappeared, but its format ZIP became the de-facto standard and survived decades later, which was likely the result of some random circumstances.
So it shouldn't be surprising that the format isn't necessary ideal looking decades back.
(One could add that ZIP encryption is a big mess. The original ZIP encryption is insecure, there are two incompatible less insecure modes that still are vulnerable to malleability attacks and there's no way to have any modern form of encryption with an AEAD in ZIP.)
Unfortunately, as far as I can see even that standard doesn't forbid self-extracting archives, non-empty zip file comments, or local file headers which don't match the central directory. It would be great if that standard were updated to also forbid these problematic features of the ZIP format, so that archive files created following that standard could be parsed without ambiguity. Then it would be a matter of updating other ZIP-using standards like OpenDocument to reference this simplified standard.
I think is better to use a separate format for archive and for compression (like .tar.gz is, and some others).
(If you do not need the metadata, I like the Hamster archive format, which is: It is a sequence of lumps, where each lump consists of: null-terminated ASCII filename, 32-bit PDP-endian data size, and then the data. (That is the entire specification.))
(Another thing about ZIP: If you have a truncated ZIP file (possibly due to a disk error on a floppy disk), I have found that bsdtar can read it; other programs I have tried are unable to read truncated ZIP files.)
That has the bad property that it does not allow for easy random access. With the ZIP format, you can directly read an arbitrary file within the archive without having to decompress or even read the compressed data of any of the other files, and many applications of the ZIP format (for example, JAR archives) depend on that property. With .tar.gz and similar, you have to decompress all data before the desired file.
Of course this doesn't fully solve the metadata issue and it's definitely more complex than just using the zip format. But on the plus side, it allows you to avoid using the zip format and the compression algorithm is significantly more performant and flexible.
The advantage of a user-supplied dictionary is that it can be transmitted once, out of band. With this approach it can also be larger than what would make sense for an individual file and it can take advantage of similarity between files.
But if you just send it along with the compressed data anyway, you’re not buying anything extra.
Bear in mind that the alternative (in context) is throwing a bunch of independently compressed files at something like tar. In the event that the files are small, that will absolutely destroy the compression ratio. If the files are large enough or you only have a few then you could always omit the custom dictionary.
One possibility is to make each file as a separate compression block, which is possible if using concatenatable compression formats; then, the same format can be used for solid and for non-solid compression.
However, then, you will need a index if you want random access. I don't know what compression formats have such a feature as a index, but I thought of some ideas of a (concatenatable) compression format; one possible thing to add is a optional index block (which links the compressed and uncompressed offsets of the beginning of blocks which do not depend on any previous blocks).
If something like tar.gz had been standard then, then that wouldn’t have been possible.
- a file named “TOC” with a well-specified format that contains the byte offsets of each file in the archive
- a file named “TOC-offset” with content the 8-byte offset of the “TOC” entry (keeping this separate means the file can be fixed-size. That makes reliably finding it easier)
Preferably, both files also would have some checksum. The TOC itself could have a pointer to the previous TOC. If so, it could be a diff.
Supporting tools could easily check for the presence of a TOC and either overwrite it with new files and wrote a new one, or just append new files and write a new TOC pair.
Attackers could thwart this by creating an archive without a TOC that has a file ending in what looks like a TOC/TOC-offset pair.
An alternative would be to write the TOC-offset file at the start. That’s more robust against such attacks, but requires the ability to overwrite the start of a file.
In general, it's also a good idea to keep data and metadata distinct, to avoid potential conflicts (and hence exploits). In this case having the archive metadata as a file inside the archive seems fraught with dangers. We could avoid this by e.g. having two levels of archive, e.g. the first level contains TOC files and data files (e.g. tar files); the 'payload' files are kept in the data files. That way, there's no way for payload to be interpreted as metadata, or vice versa.
That may have been sufficient 30 years ago...
gzip format does that too, but gunzip extracts concatenated files as one file with everything concatenated.
Zip Files: History, Explanation and Implementation
Hans Wennborg
Not sure exactly how it works for ZIP64 but it should be similar.
— in C, minimal-style library: https://github.com/richgel999/miniz, look especially at https://github.com/richgel999/miniz/blob/master/miniz_zip.h
— in JS, writing only and no compression: https://github.com/PaulCapron/pwa2uwp/blob/master/src/zip.js (255 lines including blanks & comments)
BTW, jart, let me thank you here: redbean is boldly innovative, refreshlingly witty. An inspiring work of art from a true hacker!
Also nice that it's platform independent. Same file works on big/small endian machines.
1. It has good forward / backward compatibility stories;
2. Schema is your file format documentation;
3. Don't need to decode every parts;
4. Can skip for unknowns (if you use union type);
5. Many language bindings.
The only downside, and this is a big one: validate a flatbuffers payload method doesn't have nearly as much language coverage as the encode / decode methods. (i.e. C++ has it, but a lot of other languages don't have validate method implemented).
Case study - Apache Arrow uses flatbuffers to persist its metadata: https://github.com/apache/arrow/tree/master/format
On top of that, anything intent on allow archives should provide for checksums and error correction (Reed-Solomon or similar).
This is a more easily accessed derivitive of that specification.
It has a nice central directory, an extensibility mechanism, and even supports incremental saving of large documents in a reasonably sane way. Another fancy feature only typically seen in video container formats is the ability to interleave chunks to enable streaming loads or efficient storage on mechanical drives.
Unfortunately, layering XML on top of ZIP is hideously inefficient for all sorts of reasons.
IMHO, an ideal generic format should have some of the high-level API features of OPC, but implemented with a low-level storage encoding that is:
1) Linear in the sense that no O(n^2) or even O(n log n) algorithms are never needed (such as XML parsing). Reading the data should be a single-pass, with all buffer sizes known in advance. It's okay to sacrifice write speed, because writing is less common than reading, and often scales 1:many. E.g.: a file written once is often read many times. It's almost never the case that a file is repeatedly overwritten and read only once. Similarly, writes can be cached using write-back trivially, but reads are often uncachable and are on the critical path for end-user perceived latency.
2) Inherently paralleliseable. A mistake inherited from the 1990s era of file formats is strictly sequential formats that can never be decoded faster than whatever a single CPU core can do. At a very small compression efficiency loss of about 1-5%, huge wins can be achieved by simply restarting the compressor every 64KB or so. Similarly, instead of a naive hash function, use a Merkle tree, which can be verified in parallel.
3) Safe in the sense that overlapping block references are never allowed, and the decoder code should be generated from the schema in such as way that this is strictly enforced. This would allow languages like Rust or C# to safely return "span" references to the file data buffer without having to break it up into a million tiny allocations. If aliasing is allowed, this breaks the memory model of several languages.
[0] https://johnloomis.org/cpe102/asgn/asgn1/riff.html
[1] https://en.wikipedia.org/wiki/Resource_Interchange_File_Form...
It does seem a bit weird to think that something as highly technical (and thus: subject to becoming stale quickly) as a multimedia file format has withstood the test of time so well - while "simpler" files like Office documents from 15 years ago are almost unreadable today.
But at least the multimedia files (and archive formats) were designed to be interchangeable between applications. '.DOC' word docs were basically a binary dump of its internal data structures - there was no intention of it being interchangeable with any other application, and it showed.
> Microsoft engineer looks at crack pipe...
“Yea, let’s put an entire filesystem inside this new document format”
It was there for a reason - you could embed documents/excelsheets/external apps that supported OLE into each other (think adding an editable excel graph to your document). And the FAT-like OLE file format allowed you to do partial updates to a file in place, which was pretty useful given the slow write speeds of hard disks or even floppy drives. Fully rewriting the whole ZIP/DOCX file on ever Save like we can do now wasn't an option then.
I'm not sure if, given the requirements driven by that embedding use case, the engineers could have come up with something better in the 90s.
That’s not the reason. Writing to disk and reading it back is bad, but keeping the whole file in memory is far worse and people seem to forget that routinely.
If you read from the beginning, you can stream the decompression. If you are already taking care not to clobber important files on your filesystem, this isn’t much worse.
The other advantage in the modern era of this format is that you can stream compression of the archive too. Need to send someone a bunch of files? You can start writing the file to the client while still compressing the data. If you’re very brave, you can start sending it while you’re still globbing the file system or pulling database records. This means you aren’t turning CPU time on the source into latency at the receiving end.
Edit: Extracting shows a warning regarding data after the last record. Yep, appears to have been truncated.
I only wish that compression table was at the forefront of its file format.
Not to be confused with the other JAR: http://www.arjsoftware.com/jar.htm
tar is nicely simple, if it weren't for the fact that there is so many legacy variants to support.
But one good thing that RAR has is recovery records to protect against data corruption, nice feature for a long term archiving format. Not sure if other formats have this too.
This can be done with any format by having an additional .par2 file for the recovery information. [0]