Python – Create large ZIP archives without memory inflation
github.com
github.com
I'm also perplexed by the goal as this seems to just call zipfile.write under the hood, which already streams to a zip file without accumulating a memory buffer?
[0] https://github.com/BuzonIO/zipfly/blob/master/zipfly/zipfly....
Of course, by crafting a special file-like object you could avoid this too, but perhaps a bit less elegantly.
The "usual"/naive solution (if you stay within the python ecosystem) is to compress the files and write to a BytesIO or other in-memory file like object, and then have your framework serve it. The naive solution leads to writing the whole file to memory before serving (thus memory inflation).
This library just looks like a pretty straightforward way to implement the same idea, but with chunking to bound memory usage. At the bottom, it's doing the same thing, but using generators to yield chunks at a time.
It's a useful utility for that context. Nothing groundbreaking, it's something that most intermediate and higher developers could stitch together in probably a few days (especially if they had to brush on on DEFLATE and generator protocol), but it's nice to not have to.
This is a valid concern and a good enough reason to have such a library. I've written a similar thing for uploading large files to S3 (via Django) by streaming them without the file ever touching the file system (S3ChunkUploader). The reason was the large files were deemed security-sensitive and the containers were limited to 2GB in disk space. Just uploading 4 500MB files at the same time would be an attack vector.
Examples:
* https://hackage.haskell.org/package/conduit-extra-1.1.7.3/do...
* http://hackage.haskell.org/package/streaming-utils-0.2.0.0/d...
Python seems a curious choice. Compression is computationally intensive.
If you want that speedup on the command line without Python, check out pigz. It's gzip with parallelism. Easy 10-20x speedup for some jobs.
It runs on a small embedded device that can stream zip archives many times larger then the disk or system ram without any issue.
Example Python Falcon Proof of Concept:
https://gist.github.com/kylemanna/1e22bbf31b7e5ae84bbdfa32c6...
Other then what Python's zipfile buffers in memory, my implementation shouldn't use much more then a os.pipe()'s buffer (typically 64kB?).
I need to open a very large CSV file in Python, which is around 25GB in .zip format. Any idea how to do this in a streaming way, i.e. stopping after reading the first few thousand rows?
Replace the `file_paths` list in my proof of concept with your large file(s), delete the rest (lines 61-68, 77-79) and it should just work.
import io, csv, zipfile
max_lines = 10
with zipfile.ZipFile("data.zip") as z:
for info in z.infolist():
with z.open(info.filename) as f:
reader = csv.reader(io.TextIOWrapper(f))
for i_line, line in enumerate(reader):
if i_line >= max_lines: break
print(line)