To be completely honest this all started as a quick hackathon project just to speed up downloading before piping to GNU tar. Only later did I consider also re-implementing the tar extraction in a parallel manner. If considering changing the overall packaging method from tar to something else there's a lot more ways to consider going about this (including squashfs).
I'd say one of the nice things about this is that tar is fairly ubiquitous, not just in container images. For example, we also have tarball build artifacts in Jenkins jobs that could benefit from this tool.
Ultimately, you'll need to measure this to know for sure, and those results will likely only be valid on a given hardware configuration.
OverlayFS also has a "copy_up" function, where the file is copied at the initial write. Once the copy is done, I'd expect write access to be fast. Again, you'll need to measure this.
The setup could probably look like:
container read/write -> OverlayFS([mutable fs as overlay] -> [squashfs layer as underlay] -> [squashfs layer as underlay])
1) either mount tar directly
2) or use squashfs and mount that...both would be mounted with some sort of rw overlay
3) underneath do some lazy httpfs type filesystem, so filesystem can be mounted while it's downloading
4) parallelize the underlying download ala aria2c
5) provide metadata to aria2c-alike downloader with file boundaries, so extract further latency/bandwidth savings
2) Agreed, still need to look more into this, although it's more involved and entails changing more of our pipeline for packaging and distributing these images.
3) I need to look into how this would handle files that haven't been downloaded yet. From what I know about httpfs filesystems, I'm not sure how much would need to be done to let this block until file needed is downloaded vs the normal behavior of calling out to get the file being requested.
5) Could you go into more detail here? Not sure I understand how file boundaries, etc can help with latency/bandwidth.
6)I forgot to do the best part of the optimization..do feedback-guided-optimization. You can then repack the squashfs/tar file to have all the frequently-accessed files together
Hard to change the entire container ecosystem away from tarballs tho so this is still useful. Is it available anywhere? I didn’t see a link?
Hoping this is enough to release it on Github.
Real value for me would be to integrate it with containerd or similar to speed up image layer pull times, I imagine that isn’t shelling out to a separate binary tho.
What language is it implemented in?