Python – Writing large ZIP archives without memory inflation
github.com
github.com
My https://datasette.io/ application offers features to export relational data, using Python asyncio under the hood. It can currently stream an arbitrarily large table out as CSV, which is a great format for this because it can be generated without buffering the entire thing in memory.
I wrote a bit about that here: https://simonwillison.net/2021/Jun/25/streaming-large-api-re...
zipfly makes me think that maybe I could do things like "stream all of the tables from this database as a zip file full of CSVs" - and have it work for giant databases again without using a great deal of memory.
UPDATE: Actually it looks like zipfile wouldn't work here, because the library is designed to work from files that are already on disk - it doesn't look like I could feed it a few lines of CSV data at a time, I'd have to write the entire CSV out to disk first.
Edit: Dang it, GPLv3, sorry (vs Apache 2 of your project).
The two previous discussions with comments may offer enough to get you what you need: https://news.ycombinator.com/item?id=23198233 https://news.ycombinator.com/item?id=25616513
Edit 2: https://github.com/kbbdy/zipstream, no license (yet)
https://github.com/d136o/StreamingUnzip
Basically, if you have a big zip file with many files in it (csvs for example), you can pipe out the decompressed data…
It’s a bit obtuse to use since it calls for the end chunk of a zip archive (it may come from s3 for example).
Funny to see the usual great programming school divide so succinctly put :)
Anyway, already seen a number of times https://news.ycombinator.com/from?site=github.com/buzonio
The only thing that may eat up memory is the final name directory in the "footer".
No compression also has the added benefit that you can pre-compute the final size of the zip file if you know what goes into the zip already (it's just a matter of summing up the sizes, and the sizes of the info entries and final footer entries), which can be nice when sending the result over http, as it enables you to send a Content-Length, instead of not sending one because you don't know it yet, using chunked transfer encoding and letting the browser show some "20MB of unknown downloaded".
Then it really depends on the implementation. A lot of libraries will buffer in an entire input file, compress it in one go and store the result back to disk or some buffer, and things like that.
But it is perfectly possible to implement a zip writer that implements a stream/callback interface to read the result compatible with whatever your language is using (e.g. C++ istream (yeahyeah :P), C# Stream, python file-like objects, node Readable streams, etc).
With the python standard library zipfile implementation, while it is possible to open "file-like objects" (i.e. objects implementing read/write/seek, like a BytesIO buffer), it isn't really possible easily to drain the data after just a chunk of it is done (or have it output to network/http streams, as these usually do not or cannot implement seek). The OP implementation solves this by defining everything you want in the zip beforehand and then using python generators to get the data out in chunks that you can then write to whereever you want sequentially.
Maybe because Windows still can't unpack TAR out-of-box, 20 years after entering the 21 century?
(Full disclosure: I'm the main developer of stream-unzip)
If I had to build an endpoint that generated ZIPs on the fly, I would explicitly build it to be asynchronous and efficient as possible.
The only issue is that you might want to attach a buffer somewhere that might dump to disk past a certain amount, if the source of your data can’t be “paused.”