Python – Writing large ZIP archives without memory inflation
github.com
github.com
Proof of concept: https://gist.github.com/1e22bbf31b7e5ae84bbdfa32c68e03a9
So I looked to creating the zip files on the fly when requested & streaming them to the client onthe other end withou having to create temporary files and/or a lot of memmory consumption on the server. I got it working on a prototype and even found some articles from others that got it working - it is not that hard, really.
When you really need something done it’s sometimes very helpful to be able quickly get it done.
Also, most places in a zipfile you might not have context to write in a streaming fashion are predictable in size/position so you can throw in a placeholder and seek() back to it later.
Kind of a shame someone has to specifically implement a low memory usage library. That implies other implementations went the lazy route.
The CRC field of the local file header is similar, but its optionality is less painful than the length field
In other words, there isn't a single 'good' ZIP implementation, it all depends on what you're aiming for. AFAIK the built-in Python zipfile module always fully populates the local file header.
I covered that in the sentence that described these as things you could fill with a placeholder and seek() back to later.
Streaming and seeking are antithetical, an implementation gets to pick one, and I imagine for most users streaming would be the edge case. It's debatable whether a library that ensures a fully spec-conforming output is generated at the expense of memory or IO is better or worse.
(FWIW this is written from the perspective of someone who for perverse reasons once wrote a streaming ZIP reader. In order to build such a thing, the ZIP writer had to be buffering/seeking. You can't have both without giving up ability to store uncompressed assets without needless overhead (e.g. JPEG-in-deflate))
I disagree, for a library creating zip files. It's the logical way to do it. Otherwise, you're slurping potentially big files into memory, or making un-needed copies. I've also written zip library code.
Isn't that what bit 3 in the general purpose bit flags if for? When that's set, the CRC and compressed/uncompressed file lengths are written in a Data Descriptor block after the member file data.
Streaming zip creation is well supported by the format though.
I've been using this for a javascript bookmarklet (< 2000 characters) which automatically downloads all images from a web page on click.
I've used that "central directory" approach even over the internet with a great success (downloading just that segment instead of the whole archive to get a list of what is inside).
But how did you make a bookmarklet to produce a zip archive, even non-compressed?
Not quite sure I understand the question. Bookmarklets are bookmarks which contain JavaScript code, for example:
javascript:alert('Hello, World!');
You can put that in a link
<a href="javascript:alert('Hello, World!');">Bookmarklet</a>
and then you can right-click the link to bookmark it. Now, when you want to run that JavaScript code on some website, visit the website and click on the bookmark.
Do you, or do you do something much simpler?
To create the archive data, I used an Uint8Array where I wrote the bytes into.
To download the images, I used XMLHttpRequest.
> I also believed that the bookmarklets are limited to the security context of the web page
That seems to be true unfortunately. However, besides the images on the same domain, it should also be possible to download some more images from other domains by allowing cross origin requests if the remote server cooperates and sets the respective header. but I have not looked into that yet.
The bookmarklet is here if you have more questions: https://github.com/983/FileDownloader
It is a bit larger than 2500 characters now since previously the code was optimized for size. Now it is slightly more readable.
Being able to download individual parts seems good: you could have seperate sha256 info for files for instance.
Bookmarklets are executed in the context of the currently opened website, so downloading images from the same origin is usually not a problem. What can be an issue are remote images. I tried setting `Access-Control-Allow-Origin: *`, but then other downloads would fail, so I left it as it was since it worked for the websites I was interested in.
Here is the code: https://github.com/983/FileDownloader
I’m a long time python user but haven’t ever had to work with zip files and didn’t know the answer.