- A valid zip does not begin with, as you would normally expect, a magic number.
- A valid zip can contain arbitrary prepended data.
- A valid zip can contain sections with no identifier value.
- A valid zip can contain arbitrary data between sections.
- A valid zip is validated starting with a tail section located at the end.
- A valid zip can have a valid tail section that contains a valid tail section.
- A valid zip can contain arbitrary appended data.
I don't get it: Why design a format that's so hard to parse? Implementing a single-pass streaming parser is impossible. It should be a basic requirement for most file formats. /usr/bin/unzip cannot even extract from standard input. I'm sure the implementer didn't feel like receiving user complaints about exhausted memory.