42 karma · joined January 26, 2022
I only later decided to look into also parallelizing the tar extraction as well. This originally started with looking into making a change to the GNU tar source code itself. However upon cloning I realized there were a lot of design choices in tar that made the assumption of single threaded execution (global state, etc). At this point I decided it would just be easier to re-implement the tar extraction myself using golang's "archive/tar" package.
It should also be possible to split off and write the raw download to disk in the background with something like an io.TeeReader object [0]. You could then checksum it after the fact like you normally would.
Hoping this is enough to release it on Github.
2) Agreed, still need to look more into this, although it's more involved and entails changing more of our pipeline for packaging and distributing these images.
3) I need to look into how this would handle files that haven't been downloaded yet. From what I know about httpfs filesystems, I'm not sure how much would need to be done to let this block until file needed is downloaded vs the normal behavior of calling out to get the file being requested.
5) Could you go into more detail here? Not sure I understand how file boundaries, etc can help with latency/bandwidth.
To be completely honest this all started as a quick hackathon project just to speed up downloading before piping to GNU tar. Only later did I consider also re-implementing the tar extraction in a parallel manner. If considering changing the overall packaging method from tar to something else there's a lot more ways to consider going about this (including squashfs).
I'd say one of the nice things about this is that tar is fairly ubiquitous, not just in container images. For example, we also have tarball build artifacts in Jenkins jobs that could benefit from this tool.