Pack: A new container format for compressed files
pack.ac
pack.ac
I wonder how pack compares to it, but its home page and github don't tell much.
Wake me up if a simple standard comes to pass that neither has user/group ID, mode fields, nor bit-packed two-second-precision timestamps or similar silliness. Perhaps an executable bit for systems that insist on such things for being able to use the download in the intended way
(I self-made this before: a simple length-prefixed concatenation of filename and contents fields. The problem is that people would have to download an unpacker. That's not broadly useful unless it is, as in that one case, a software distribution which they're going to run anyway)
Other than having timestamps isn't this a ZIP file? No user id, no x bit, widely available implementations... Not very simple though I guess.
I wrote a ReadableStream to Zip encoder (with no compression) in 50 lines of Javascript.
The problem is that zip is finicky and extremely poorly documented. I had to look at what other implementations do to figure out some of the fields. About at least one field, the spec (from the early 90s or late 80s I think) says it is up to you to figure out what you want to put there! After all that, I additionally wrote my own docs in case someone coming after me needs to understand the format as well, but some things are just assumptions and "everyone does it this way"s, leading to me having only moderate confidence that I've followed the spec correctly. I haven't found incompatibilities yet, but I'd also not be surprised if an old decoder doesn't eat it or if a modern one made a different choice somewhere.
It's also not as if I haven't come across third party zip files that the Debian command line tool wouldn't open but the standard Debian/Cinnamon GUI utility was perfectly happy about. If it were so well-documented and standard, that shouldn't be a thing. (Similarly, customers on macOS can't open our encrypted 7z pentest report files. The Finder asks for the password and then flat-out tells them "incorrect password", whereas in reality it seems to be unable to handle filename encryption. Idk if that is per the spec but incompatibilities are abound.)
If you're not sure what the spec is trying to say, then either the PKZip binaries or the Info-ZIP zip/unzip source code is your usual source of truth.
When one unzip works but another unzip app doesn't, then you can usually point the finger at the last zip app that modified the zip file. There's some inconsistency in the zip file.
Running "unzip -t -v" on the zip file in question may yield more info about the problem.
You don't have to include an OS-specific extra field unless you want the information in that specific extra field to be available by the party trying to extract the contents of the zipfile.
Sometimes you want to include data and sometimes you don't for different reasons in different contexts. It's not a data handlers job to decide what data is or isn't included, it's the senders job to decide what not to include and the receivers job to decide what to ignore.
The simplest example is probably just the file path. tar or zip don't try to say whether or not a file in the container includes the full absolute path, a portion of the path, or no path.
The container should ideally be able to contain anything that any filesystem might have, or else it's not a generally useful tool, it's some annoying domain-specific specialized tool that one guy just luuuuuvs for his one use-case he thinks is obviously the most rational thing for anyone.
If you don't want to include something like a uid, say for security reasons not to disclose the internal workings of something on your end, then arrange not to include it when creating the archive, the same way you wouldn't necessarily include the full path to the same files. Or outside of a security concern like that, include all the data and let the recipient simply ignore any data that it doesn't support.
Even if a given decoder could, though, most users wouldn't be able to use that and so they'd get files from 1970 or 1980 if I don't want to pass that on and set it to zeroes, so better is if the field can be omitted (as in, if the header wasn't fixed length but extensible like an IP packet). So I'd still like a "better" archiving format than the ones we have today (though I'm not familiar with the internals of every one of them, like 7z or the mentioned squashfs so tell me if this already exists), but I agree such a format should just support everything ~every filesystem supports
os and filesystem features differ all over the place, and there will be totally new filesystems and totally new metadata tomorrow. There is practically no common denominator, not even the basic ascii for the filename let alone any other metadata.
So there should just be metadata fields where about the only thing actually parrt of the spec is the structure of a metadata field, not any particular keys or values or number or order of fields. The writer might or might not even include a filed for say, creation time, and the reader might or might not care about that. If the reader doesn't recognize some strange new xattr field that only got invented yesterday, no problem, because it does know what a field is, and how to consume and discard fields it doesn't care about.
There would be a few fields that most readers and writers would all just recognize by convention, the usual basics like filename. Even the filename might not be technically a requirement but maybe an rfc documents a short list of standard fields just to give everyone a reference point. But for instance it might be legal to have nothing but some uuids or something.
That's getting a bit weird but my main point was just that it's wrong to say an archiver shouldn't include timestamps or uids just because one use of archive files is to transfer files from a unix system to a windows system, or from a system with a "bob" user to a system with no "bob" user.
For unzip, -D skips restoration of timestamps.
For unrar, -ai ignores file attributes, -ts restores the modification time.
There are similar arguments for omitting these when creating the archive, they set the field to a default or specified value, or omit it entirely, depending on the format.
The comment I'm replying to suggested that since one use case results in metadata that is meaningless or mis-matched between sender and receiver, the format itself should not even have the ability to record that metadata.
Frankly, I don't even want any of these things on my mounted filesystem either.
> The problem is that people would have to download an unpacker.
Your archive format just needs to be an executable that runs on every platform. https://github.com/jart/cosmopolitan is something that could help with that. ("Who would execute an archive? It could do anything," I hear you scream. Well, tell that to anyone who has run "curl | bash".)
tar --create --owner=0 --group=0 --mtime='2000-01-01 00:00:00' \
--mode='go-rwxst' --file test.tar /bin/dash /etc/hosts
tar --list --verbose --file test.tar
-rwx------ root/root 125688 2000-01-01 00:00 bin/dash
-rw------- root/root 1408 2000-01-01 00:00 etc/hostsIf only 7zip could also create them on Windows (it apparently can WIM which seems a direct Windows-native counterpart, also mountable on Linux).
I even dream about the days when a file main stream would be pure data and all the metadata would go to EAs. Imagine an MP3 file where the main stream only records the sound but no ID3, all the metadata like the artist and the song names are handled as EAs and can be used in file operation commands.
This also can be made tape-friendly and eliminate need in TAR. Just make sure files (streamms/EAs are written contiguous, closely-related streams go right near, compression is optional and the ToC+dictionary is replicated in a number of places like the beginning, the middle and the end).
As you might have guessed I use all the major OSes and single-OS solutions are of little use to me. Apparently I'd just use SquashFS but it's use is limited on Windows because you can hardly make or mount one there - only unpack with 7zip.
(ADS are super commonplace and the closer analogue to posix xattrs.)
I am looking forward to write my own cross-platform app which would rely on attaching additional data and metatata to my files.
"need to be in kernelspace" does not sound very scary because a user app probably doesn't need to do this very job itself - isn't there usually an OS-native command-line tool which can be invoked to do it?
Example: https://alexwlchan.net/2019/working-with-large-s3-objects/
squashfs may or may not be able to do it with as few roundtrips (I don't know the details of its layout), but S3 necessarily provides the capabilities for random access otherwise you'd have to download the entire content either way and the original query would be moot.
There's the caveat of the zipfile itself may have stuff that's not mentioned in the actual central directory of the zipfile.
First, zip files already have a central directory so why would you do that?
Second, you seem to be missing the subject of this subthread entirely, the point is being able to selectively access S3 content without downloading the whole archive. If you sequentially read through the entire zip file, you are in fact downloading the whole archive.
A zipfile can be treated as a stream of data and processed as each individual zip entry is seen in the download/read. NO random access is required.
Just enough memory for the local directory entry and buffers for the zipped data and unzipped data. The first two items should be covered by the downloaded zipfile buffer.
If you want to process the zipfile ASAP or don't have the resources to download the entire zipfile first before processing the zipfile, then this is a valid manner to handle the zipfile. If your desired data occurs before the entire zipfile has been downloaded, you can stop the download.
A zipfile can also be treated as a randomly accessed file as you mentioned. Some operations are faster that way - like browsing the each zip entry's metadata.
for sosreports (archives with lots of diagnostic commands and logfiles from a linux host), I wanted to find a file format that can both used zstd compression (or maybe something else that is about as fast and compressible, currently often uses xz which is very very slow) -and- that lets you unpack a single file fast, with an index, ideally so you can mount it loopback or with fuse or otherwise just quickly unpack a single file in a many-GB archive.
You'd be surprised that this basically doesn't exist right now. Theres a bunch of half solutions, but no real good easily available one. Some things add indexes to tar, zstd does support partial/offset unpacking without reading the entire archive in the code but basically no one uses that function, it's kindof silly. There are zip and rar tools with zstd support, but they are not all cross compatible and mostly doesn't exist in the packaged Linux versions.
squashfs with zstd added mostly fits the bill.
I was really surprised not to find anything else given we had this in Zip and RAR files 2 decades ago. But nothing so far that would or could ship on a standard open source system managed to modernise that featureset.
(If anyone has any pointers let me know :-)
`pack -i ./test.pack --include=/a/file.txt`
or a couple files and folders at once:
`pack -i ./test.pack --include=/a/file.txt --include=/a/folder/`
Use `--list` to get a list of all files:
`pack -i ./test.pack --list`
Such random access using `--include` is very fast. As an example, if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.
- It is read-only; Pack is not. Update and delete are not just public yet, as I wanted people to get the taste first.
- It is clearly focused on archiving, rather than Pack wanting to be a container option for people who want to pack some files/data and store or send them with no privacy dangers.
- Pack is designed to be user-friendly for most people; CLI is very simple to work with, and future OS integration will make working with it like a breeze. It is far different from a good file system focused on Linux.
- I did not compare to squashfs, but I will be happy to see any results from interested people.
My bet is on Pack, obviously, to be much faster.
- being read only is mostly a benefit to an archive. Back in the days when drives had been small, I occasionally wanted to update a .rar, but in the last ~5 years I can't remember a case for it.
- it's fine, but don't think that others' use cases are invalid because of your vision
- mount is also a CLI interface
I was trying to change a single file in squashfs container recently and could not find a way to do that.
My use case is extremely redundant data (specific website dumps + logs) that I want decently quick random access into, and I was unhappy with either the access speed, quality/usability or even existence of libraries for several formats.
Glancing over the code this seems to use the following setup:
- Aggregate files
- Chunk into blocks
- Compress blocks of fixed size
- Store file to chunk and chunk to block associations
What I did not see is a deduplication step for the chunks, or an attempt to group files (and by extend, blocks) by similarity in an attempt improve compression.
But I might have just missed that due to lack of familiarity with Pascal.
For anyone interested in this strategy, take a look at ZPAQ [1] by Matt Mahoney, you might know him from the Hutter Prize competition [2] / Large Text Compression Benchmark. It takes 14th place with tuned parameters.
There's also a maintained fork called zpaqfranz, but I ran into some issues like incorrect disk size estimates with it. For me the code was also sometimes hard to read due to being a mix of English and Italian. So your mileage may vary.
[1]: http://mattmahoney.net/dc/zpaq.html [2]: http://prize.hutter1.net [3]: https://github.com/fcorbelli/zpaqfranz
It supports zstandard compression, random access, and it's very robust.
I'm happy to see a fellow enthusiast. Your deduction is on point. And also, Pack is smart; it skips non-compressible files like MP3 [1], so you do not need to choose the "Store" option to have a faster option, and it speedup decompression too. Pack is the first to achieve this, being faster than Store options. Yes, it was a surprise to me too.
ZPAQ is great, and I study the Hutter Prize competition. Pack is on another chart, which is why I proposed CompressedSpeed [2]. The speed of getting to compression needs to be accounted for. You can store anything on an atom if you try hard enough, but hard work takes time. Deduplication step may get added, but in Hard Press [3].
I am curious to see the results of Pack on your data. You can find me here or o at pack.ac.
[1] It is based on content rather than extension; any data that is determined not to be worthy of compression, will be stored as is. And as a file can get chucked, some parts can get compressed and some cannot. Imagine that part of the subtitle in a MKV file can get compressed, and the Video part gets skipped. Although these features will get more updates over time, if they don't cost time,. Pack focus is being seamless and not the most compressed; there are already great works in the field, such as the noted ZPAQ.
[2] CompressedSpeed = (InputSize / OutputSize) * (InputSize / Speed). Materialized compression speed.
[3] You can choose --press=hard to ask for better compression. Even with Hard Press, Pack does not try to eat your hardware just to get a little more; it goes the optimized way I described.
It has a ton of comparison with existing tools in the README -- zpaqfranz included -- and it seems to be the best there is.
OS containers could use an update too, though. They're often so big and tend to use multiple tar.gz files.
If I understand correctly, they're suggesting Pack, which both archives and compresses, is 30x faster than creating a plain tar archive. That just sounds like you used multithreading and tar didn't.
Either way, it'd be nice to see [a version of] Pack support plain archival, rather than being forced to tack on Zstd.
Being better than that is not a hard bar.
IE
/foo.txt 21
This is the foo file
/bar.txt 21
This is the bar file
That makes it super hard to deal with as you essentially need to navigate the entire tar file before you can list the directories in a tar file. To add a file you have to wait for the previous file to be added.Using something like sqlite solves this particular problem because you can have a table with file names and a table with file contents that can both be inserted into in parallel (though that will mean the contents aren't guaranteed to be contiguous.) Since SQLite is just a btree it's easy (well, known) how to concurrently modify the contents of the tree.
Since tar is a Tape ARchive, the way tar operates makes sense (as it was designed for both output to file and device, i.e. tape).
Yes! It is! But it's awful at archive files, which is what it's used for nowadays and what's being discussed right now.
Over the past 50 years some people did try to improve tar. People did develop ways to append a file table at the end of an archive file. Maintaining compatibility with tapes, all tar utilities, and piping.
Similarly, driven people did extend (pk)zip to cover all the unix-y needs. In fact the current zip utility still supports permissions and symlinks to this day.
But despite those better methods, people keep pushing og tar. Because it's good at tape archival. Sigh.
Modern tape archives like LTFS do the same thing as well.
Zip format was not designed for piping.
(Too long to post here)
> Development machine with a two-year-old CPU and NVMe disk, using Windows with the NTFS file system. The differences are even greater on Linux using ext4. Value holds on an old HDD and one-core CPU.
> All corresponding official programs were used in an out-of-the-box configuration at the time of writing in a warm state.
7-Zip was used as others, just gave it a folder to compress. No configuration.
As requested, here are some numbers on tar.zst of Linux source code (the test subject in the note): tar.zst: 196 MB, 5420 ms (using out-of-the box config and -T0 to let it use all the cores. Without it, it would be, 7570 ms) Pack: 194 MB, 1300 ms Slightly smaller size, and more than 4X faster. (Again, it is on my machine; you need to try it for yourself.) Honestly, ZSTD is great. Tar is slowing it down (because of its old design and being one thread). And it is done in two steps: first creating tar and then compression. Pack does all the steps (read, check, compress, and write) together, and this weaving helped achieve this speed and random access.
Benchmark 1: tar -c ./linux-6.8.2 | zstd -cT0 --zstd=strat=2,wlog=24,clog=16,hlog=17,slog=1,mml=5,tlen=0 > linux-6.8.2.tar.zst
Time (mean ± σ): 2.573 s ± 0.091 s [User: 8.611 s, System: 1.981 s]
Range (min … max): 2.486 s … 2.783 s 10 runs
Benchmark 2: bsdtar -c ./linux-6.8.2 | zstd -cT0 --zstd=strat=2,wlog=24,clog=16,hlog=17,slog=1,mml=5,tlen=0 > linux-6.8.2.tar.zst
Time (mean ± σ): 3.400 s ± 0.250 s [User: 8.436 s, System: 2.243 s]
Range (min … max): 3.171 s … 4.050 s 10 runs
Benchmark 3: busybox tar -c ./linux-6.8.2 | zstd -cT0 --zstd=strat=2,wlog=24,clog=16,hlog=17,slog=1,mml=5,tlen=0 > linux-6.8.2.tar.zst
Time (mean ± σ): 2.535 s ± 0.125 s [User: 8.611 s, System: 1.548 s]
Range (min … max): 2.371 s … 2.814 s 10 runs
Benchmark 4: ./pack -i ./linux-6.8.2 -w
Time (mean ± σ): 1.998 s ± 0.105 s [User: 5.972 s, System: 0.834 s]
Range (min … max): 1.931 s … 2.250 s 10 runs
Summary
./pack -i ./linux-6.8.2 -w ran
1.27 ± 0.09 times faster than busybox tar -c ./linux-6.8.2 | zstd -cT0 --zstd=strat=2,wlog=24,clog=16,hlog=17,slog=1,mml=5,tlen=0 > linux-6.8.2.tar.zst
1.29 ± 0.08 times faster than tar -c ./linux-6.8.2 | zstd -cT0 --zstd=strat=2,wlog=24,clog=16,hlog=17,slog=1,mml=5,tlen=0 > linux-6.8.2.tar.zst
1.70 ± 0.15 times faster than bsdtar -c ./linux-6.8.2 | zstd -cT0 --zstd=strat=2,wlog=24,clog=16,hlog=17,slog=1,mml=5,tlen=0 > linux-6.8.2.tar.zst
Another machine has similar results. I'm inclined to say that the difference is probably mainly related to tar saving attributes like creation and modification time while pack doesn't.> it is done in two steps: first creating tar and then compression
Pipes (originally Unix, subsequently copied by MS-DOS) operate in parallel, not sequentially. This allows them to process arbitrarily large files on small memory without slow buffering.
Anyway, I do not expect an order of magnitude difference between tar.zst and Pack; after all, Pack is using Zstandard. What makes Pack fundamentally different from tar.zst is Random Access and other important factors like user experience. I shared some numbers on it here: https://news.ycombinator.com/item?id=39803968 and you are encouraged to try them for yourself. Also, by adding Encryption and Locking to Pack, Random Access will be even more beneficial.
Yes it is that much faster, and a good part of it is because of the multi-thread design, but as a reminder, WinRAR or 7-Zip are too multi-thread, and you can see the difference. To satisfy your doubt, I suggest running Pack for yourself. I am looking for more data on its behaviour on different machines and data.
Can I ask why do you need a version without ZSTD? If you are thinking that compression slows it down, I should say no. Pack is the first of its kinds that "Store" is slowing it down. Because its compression is smart, it will skip any non-compressible content.
On the same machine and the same Linux source code test:
Pack: 194 MB, 1.3 s
Pack (With no Press): 1.25 GB, 1.8 s
.pack vs zst-7z with the same compression settings would b interesting. That will be pure container overhead
Portability: Receiver (or future you) needs to know what you used, and what version even.
Speed: If you want to do the archive part first (tar) and then compress (gz), you will get much lower speed (as shown in the note).
Reliability: Most people use tar with gz anyway, but if you use it with not so popular algorithm and tools, you will risk having a file that may or may not work into the future.
Pack plan is to use the best of time (Zstandard) and if an update is needed in years to come, it will add support for the new algorithm updates. All Pack clients must only write the latest version (and read all previous versions) and that makes sure almost all use the best of their time.
I wonder if this line is an array in-situ?
Split(APath, [sfodpoWithoutPathDelimiter, sfodpoWithoutExtension], P, N)[1] https://www.freepascal.org/docs-html/current/ref/refse83.htm...
> It is written in the Pascal language, a well-stabilized standard language with compatibility promises for decades. Using the FreePascal compiler and the Lazarus IDE for free and easy development. The code is written to be seen as pseudocode. In place of need, comments are written to help.
Use the language that makes you money and encourages you to write code that addresses domains requiring more than just bolting together framework pieces. AI can do a measurable chunk of that work.
It had a nice unit system with separate interface & implementation sections. This was very nice. The unit files were not compatible with anything else, including previous versions of Delphi - this was not nice, especially since a lot of libraries were distributed in compiled form.
The compilation speed was amazingly fast. This is one thing that was unequivocally better than C at this time.
There were range types (type TExample = 1..1000), but they more of a gimmick - turns out there are very few use cases for build-time limits. There were some uses back in DOS days when you'd have hardcoded resolution of 640x480, but in the windows time most variables were just Integer.
Arrays had optional range checks on access, that was also nice. We'd turn them off if we felt programs were too slow.
Otherwise, it was basically same as C with a bit of classes - custom memory allocation/deallocation, dangling pointers, NULLs, threads you start and stop, mutexes and critical sections. When I finally switched from Pascal to C, I didn't see that much difference (except compilation got much slower)
Maybe you'd say that Borland did something wrong, and Wirth's Modula-2 would be much better than Borland's Pascal, but I doubt this.
Lazarus is the best IDE for Pascal, being completely free, open source and cross platform.
Wirth languages are about constraints. For instance, when I started writing code in TSM2 and Stonybrook, my general impression was that they both emitted 10-30% more compile time bugs than BP did. If that's too much of a hassle for C programmers, well ok.
Also to add, all the wordiness of Wirth languages, the block delimiters, yes I get it. But all this stuff is just another constraint for sake of correctness. M2, being case sensitive, is even worse about this than Pascal. But the point is to make you look at your code more than once, to proofread it and think about what's going on, because the syntax screams at you a little bit. Of course, with compiler defines, you can turn pascal into C and assume the responsibility for yourself. That's what runtime debuggers are for anyway, yes?
Ok, whatever, but we're missing the point that Wirth was trying to get across, which is to turn the language itself into implicit TDD, starting with first line of code written. C/++ may give you speed, but for the average programmer, all that speed is taken back in the end, due to maintenance costs. IMO, M2 was even better at shifting maintenance costs left of what the C tack did than pascal in the value delivery stream.
Sure, mission critical code can be written even in C/++. Most of SpaceX's code is a C++ codebase. So how did they pull that off? IMO, what they did was write C in the spirit of what Wirth was trying to accomplish. For the sake of maintenance costs, speed is now less of a metric thanks to hardware advances, and correctness is far more of an issue. Which makes sense, because all business is mission critical now and all business runs on more and more software. Would you turn off bounds checks in the compiler now? How about for the programmer who you'll never meet who is writing autonomous driver code for the car you drive?
Way too much money was wasted on the near-sighted value of C. Time to move on, according to Rust developers, who undoubtedly have an impressive background as C/++ programmers. So yeah we all have to follow this C dominated narrative even today, and my charge is, this narrative has retarded the art of programming. So I stick to my original proposition: Whatever you think is great about Rust, like ownership and borrowing, would have been available in production M2 code 20 years ago if we had just given Wirth languages a chance to advance the art in the commercial world. But that narrative would have been too wordy and constraining.
* There is already a "native" Sqlite3 container solution called Sqlar [0].
* Sqlite3 itself is certainly suitable as a base and I wouldn't worry about its future at all.
* Pascal is also an interesting choice, it is not the hippest language nor a new kid on the block, but offers its own advantages as being "boring" and "normal". I am thinking especially of the Lindy effect [1].
All in all a nice surprise and I am curious to see the future of Pack. After all, it can only succeed if it gets a stable, critical mass of supporters, both from the user and maintainer spectrum.
Here is the latest sqlar result on Linux source code on the same test machine in warm state:
sqlar: 268 MB, 30.5 s
Pack: 194 MB, 1.3 s
Very good result compared to tar.gz. And much better than ZIP, considering sqlar gives random access like ZIP and unlike tar.gz. I considered sqlar as a proof of concept, and it inspired me to create Pack as a full solution. I always agreed with the great drh (creator of SQLite) points about SQLite as a file format, and Pack is a try to demonstrate that.
I made Pack to give people a better life (at least behind their desks), and as you do, I hope people get to use it and find it useful.
If you actually care about the future, spec out your own database format and use that instead. It could even be mostly a copy of Sqlite3, but at least it would be part of the spec itself.
If I can't extract .pack archives 3 decades from now, the use of SQLite 3 will be the reason.
The Library of Congress has a page that goes into some depth with respect to their sustainability analysis for the format.
https://www.loc.gov/preservation/digital/formats/fdd/fdd0004...
Hot take: SQLite has bugs and quirks.
Do you need to generate sqlite bindings for every language/runtime? E g. Cloudflare workers
No, you will only need Pack, everything is built into it. Pack is built for Windows and Linux, and more will come. You will be able to run it on almost all CPUs.
* Able to write to a pipe/socket - lets you not waste space or time by writing to disc something that you intend to transmit over a pipe or TCP socket anyhow. It's almost a "virtual" archive and it should be possible to make one that is far too big to fit into memory - because as you send each bit of it you deallocate that memory. At the receiver each bit can be written to disc or extracted and then that memory is reused for the next bit - so the archive never fully "exists" but it does the job of serialising some data. An example could be piping the output of tar to an ssh command which untars it on a remote machine.
* Metadata has to be with the file data - not stuck at the end of the file - because you need to be able to start work without waiting till the file is fully received through your pipe. You don't want to be forced to have space to store the archive and the extracted files (may be a huge archive).
* Choice of compression - lzop is super fast such that using it can sometimes give slightly better performance than writing uncompressed data. OTOH that might not be your concern and XZ might suit you by compressing much more thoroughly. Either way it's very nice to have compression that works across multiple files - which is especially helpful when compressing a lot of small files such as source code.
* Ability to encapsulate - should be able to put the packed data into any imaginable container like an encryption or data transmission protocol without insisting that the entire archive has to be fully read before members can start to be extracted/processed. This is essentially the same as the pipe/socket requirement.
I'm not saying that these things matter to everyone - I have just found them incredibly useful in a few critical situations. The world of ZIP users on Windows seems to be sort of blind to them - thinking firmly in that box.
Seems to me like this is you being blind to certain use cases, and so stuck in your streaming oriented box that you can not conceive of other use cases where a streaming format is actively detrimental.
- Piping is really easy and it will get added to Pack. It is matter of time, until these features get added as they will be added based on popularity and piping is not that popular for most people. But I get you and I will add it for you.
- Metadata is not stored in Pack. I don’t want the metadata of my machine attached to a file. It’s a never-ending nightmare to match source and destination OS metadata. There will always be something missing, and Pack tends to get everything perfect or nothing. Storing metadata adds extra weight that most users don’t care about and complicates the ability to store other types of data alongside files. It may get added as an option in the future if many people need it.
- Pack uses Zstandard under the hood. Great compression speed and ratio. In my opinion, it is the leading algorithm in the field and makes it a proper choice to use instead of DEFLATE, used in ZIP or GZIP.
- At this point you are telling tar features. tar is not random access, Pack is. As an example, if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.
This is not something I've wanted yet personally but that's just random chance. When I do need it I will know which tool to use! Thanks.
It could be handy to be able to mount a pack like a filesystem.
Someday, it can be used as a virtual drive. I leave it to future people.
As mediocre as zip and tar are, you can cobble together read/write support without even needing a library. With sqlite, your only real option is to bundle sqlite itself, and while it's relatively lightweight, it's far from trivial.
zip has support for zstd, and if you wanted to make it go faster, you could embed some index metadata.
I can't see any specs for their format, not even a description of the sqlite tables.
CREATE TABLE Content(ID INTEGER PRIMARY KEY, Value BLOB);
CREATE TABLE Item(ID INTEGER PRIMARY KEY, Parent INTEGER, Kind INTEGER, Name TEXT);
CREATE TABLE ItemContent(ID INTEGER PRIMARY KEY, Item INTEGER, ItemPosition INTEGER, Content INTEGER, ContentPosition INTEGER, Size INTEGER);
According to the `.indexes` directive, there are... no indexes. What's the point of sqlite if you're not going to index things?All the data is stored in one big blob (the "Value" column of the "Content" table), with the metadata storing offsets into it. It looks like there's still the possibility of things being split over multiple blobs (to circumvent the 2GB blob size limit)
Summary:
- Custom sqlite magic bytes makes the format incompatible with all existing sqlite tooling.
- No support for file metadata.
- There's no version field (afaict), making future format improvements difficult.
Edit: A previous version of this comment had a much longer list of complaints, but after taking a closer look, I retract them. I was looking at the MediaKit.pack file as an example, which, due to being relatively small, packed all its files into a single BLOB. I was under the mistaken impression that the same approach was taken for larger files, but after some further testing I see that they're split up into ~8MB chunks.
Though, if you have lots of small files (say, a couple of kilobytes each) then random access performance could suffer.
- SQLite tooling: You will not need it unless you are debugging something, then you can change the header or just use the `--activate-other-options --transform-to-sqlite3` parameter to transform a Pack file to SQLite3, and use the `--activate-other-options --transform-to-pack` to go back. This way, you get a true SQLite3 database that you can browse as you wish. For most people, mixing Pack with SQLite was just a call for problems for the SQLite team (imagine people coming and asking to fix their Pack file from the team; that would not be fair) and a harder future for Pack to update.
- Metadata is not stored in Pack. I don’t want the metadata of my machine attached to a file. It’s a never-ending nightmare to match source and destination OS metadata. There will always be something missing, and Pack tends to get everything perfect or nothing. Storing metadata adds extra weight that most users don’t care about and complicates the ability to store other types of data alongside files. It may get added as an option in the future if many people need it.
- There is a version field. It is currently in Draft 0, and it is written using a custom VFS. Look here for more information: https://github.com/PackOrganization/Pack/blob/main/Source/Dr...
- All future versions of Pack must handle previous versions and must only write the latest version. So any files created right now (Draft 0) will be read correctly for ever to come.
- Each Draft proposal will get its own version, and if it gets final, it will be set to final.
- Two byte after 'Pack' header in little endian as (1 (Draft) shl 13 + 0 (version 0) = 8192). Final would be 0, so the first Final version will be 0 shl 13 + 1 = 1. and the second will be 2. It is by design, so any Draft version gets a higher number, preventing future mixups.
- 8 MB chunks are the default; Pack may choose smaller or bigger (16 MB for many small files or 32 MB for Hard Press).
- Random access is proper as unpacking steps take into account what you want and decompress a Content just once for many neighbouring files. But even for reading just one file, here is an example: if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.
I believe you are making a mistake by preventing Pack from writing archives that are compatible with prior versions.
Not including metadata is an opinionated stance, but I can certainly get behind it, especially as a default. 99% of the time I do not care about metadata when producing a file archive.
Compatibility with existing SQLite tooling is not just useful for debugging, it is extremely useful for writing alternative implementations. If you want Pack to be successful as a format and not just as a piece of software, I think you should do everything you can to make this easier.
In my experimentation, I wrote a simple python script to extract files from a Pack archive. Conveniently, sqlite is part of the python standard library, but in order to make it work with that version (as opposed to compiling my own) I had to edit the file header first, which is inconvenient and not always possible to do (e.g. if write permissions are not available).
Despite that inconvenience, it took less code than a comparably basic ZIP extractor, which is cool!
I worry that requiring a custom VFS will make it harder for people to produce compatible software implementations.
I think your concerns about people contacting SQLite for support are overblown. I assume you've heard the `etilqs_` story[0], but in this case, you need to use a hex-editor or a utility like `file` to even see the header bytes. I think anyone capable of discovering that it's an SQLite DB will be smart enough not to contact SQLite for support with it.
The `Application ID`[1] field in the SQLite header is designed with this exact purpose in mind
> The application_id PRAGMA is used to query or set the 32-bit signed big-endian "Application ID" integer located at offset 68 into the database header. Applications that use SQLite as their application file-format should set the Application ID integer to a unique integer so that utilities such as file(1) can determine the specific file type rather than just reporting "SQLite3 Database".
It's convenient that `Pack` is 32 bits long ;)
[0] https://github.com/mackyle/sqlite/blob/18cf47156abe94255ae14...
[1] https://www.sqlite.org/pragma.html#pragma_application_id
Did you compile it for yourself? Any problem or steps you used, I will be happy to hear, o at pack.ac or GitHub, as it is hard to follow the building here.
As a reminder, Pack Draft 0 has Compatibility with SQLite tools; the only needed step is to change the first 16 bytes. Again, you can use `--activate-other-options --transform-to-sqlite3` with the CLI tool, and you will get a perfectly working SQLite file.
VFS is not needed; they can change the header after writing; VFS was just cleaner to me.
My first work was using application_id, after a while, it did not feel right to me, so I changed it for good. It allows easier future development, fewer problems for file type detection, a decreased chance of mistaken change (you already saw many negative comments on using SQLite as a base), and the support reason: just yesterday I was reading a forum post about people asking for support on software because it was using SQLite. application_id seems like a great choice if you are doing a DB-related task or making a custom DB for transfer on wire, to communicate between internal and semi-public tools. Using it for a format that could potentially get to an innumerable count seemed unwise.
No index, as they take space, and I wrote the queries considering SQLite automatic indexes. They will be created on demand, at unpacking time. All the unpacking processes are made to read and decompress content just once, so there are no worries about slowdowns.
I suggest trying Pack for yourself and seeing the speed. Or deeper, use `--activate-other-options --transform-to-sqlite3` to transform a Pack file to SQLite3, create your own indexes, and use `--activate-other-options --transform-to-pack` to convert it to Pack and then try unpacking. You will not see any worthy difference.
Yes, Contents are like packages of raw data from a chunk, a whole, or many of the items (files or data). They may be compressed if needed (With Zstandard). ItemContent table helps to find the needed Item parts.
The Content structure circumvents any BLOB limit, but it is also made to give better compression while keeping random access.
Let's imagine: You want to read a ZIP file. Will you write your own reader? I seriously doubt it, as the work, stabilising, and security (random memory access as an example) would be issues. But let's think we are couraginous. OK, we read rather not so simple format and carefully read the binary. Now, will you write your own DEFLATE and Huffman coding? Again, a bigger doubt.
I would argue that if someone cares enough to reimplement ZIP, it would at worst be twice as hard to write a Pack reader from scratch with no ZSTD or SQLite. And for those serious people, reading a format that lets them store better and faster would be a prize that is hard to say no to. But I get your point, and if you are in a desert and need something to put together fast before going out of water, tar may be a good choice.
You're right to call out security though, because the multiple implementations cause security issues where they disagree, my favorite example being https://bugzilla.mozilla.org/show_bug.cgi?id=1534483 . Although arguably this is a symptom of ZIP being a poorly thought out file format (too many ambiguous edge-cases), rather than a symptom of it being easy to implement.
Anyone needing to reimplement Pack, can do it, very easily, if not easier than implementing ZIP, IF they use SQLite and Zstandard. Maybe a day of work or less. If they want to rewrite (reading part of) them too, it will be a couple of days of work.
Overwrite with "-w". I've never seen a tool not use "-f"
Not reserving "-h" for help text is also an interesting choice. Makes me think of the mantra "be conservative in what you send out but liberal in what you accept". Per that philosophy, both "--help" and "-h" should be accepted because neither gained a decisive majority in usage and so people might try either. It's not like you'd know what to use because it hasn't told you yet
Forcing use of a long option "--press=N" (for the zstd level setting) is also new/unique terminology for what is usually "-N" (like "-1" to "-9")
(Basic) drop-in compatibility with every other tool from gzip to zstd would also have been nice, but archiving and compression are different things and everything from zip to 7z to tar works in unique ways so this makes enough sense I guess. Still, could have been useful
It's still better than tar or ps, so if it catches on that's still a step forwards in terms of command line standards
It seems to be "read the code" or nothing - which is fine until they update the code... It's great (probably should be mandatory) to have a reference program, but if they're promoting it as a container format, something along the lines of an RFC would be helpful.
The choice of parameters was solely done to be clear, and not what people used to. -f meaning force is not clear; -w meaning overwrite, seemed like a better logical choice, to me.
Nice point on -h. Yes, I did not want to go crazy. After all, almost all (CLI) people use pack as `pack ./test/`. Options are for advanced people like you. Most people will use the OS integration that will be published later on.
--press=hard is the only option there is. There may be more, but with Pack you do not need to choose a level (like 1..9 with ZIP). Just let Pack do its thing, and you will be happy. Hard Press is there for people who want to pack once and unpack many times (like publishing), and it is worth spending extra time on it. Even then, Pack goes the sane way and does not eat your computer just for a kilobyte or two.
Well, no. Sometimes i want maximum compression while having a lot of CPU and wall clock time. And sometimes being fast is more important than compression level. Also managing server utilization is needed. Level thing is there for a reason
https://raw.githubusercontent.com/facebook/zstd/master/doc/i...
- for automatic tasks like making backups via cron: piping, correct exit codes, level of effort configuration, silent modes for reduced logging.
Second one is likely to fly first, can be used isolated on company level if file format is stable
dnf info pack
Summary : Convert code into runnable images
URL : https://github.com/buildpacks/pack
License : Apache-2.0 and BSD-2-Clause and BSD-3-Clause and ISC and MIT
Description : pack is a CLI implementation of the Platform Interface Specification
: for Cloud Native Buildpacks.If you’re not sure, grab a very basic template from either Hugo, MkDocs, or GitHub pages. They’re all pretty well tested and hard to break.
I was able to initially construct an archive of the linux tree (that failed in decompression), but subsequently went to rebuild it and the tool is repeatedly producing this output even in an otherwise cleaned up environment:
Runtime error 203 at $0000000100009B72
$0000000100009B72
$00000001000225B6
$0000000100023153
$000000010002319A
Runtime error 203 at $0000000100009B72
$0000000100009B72
$00000001000225B6
$0000000100023153
$000000010002319A
Runtime error 203 at $0000000100009B72
$0000000100009B72
$00000001000225B6
$0000000100023153
$000000010002319A
Runtime error 203 at $0000000100009B72
$0000000100009B72
$00000001000225B6
$0000000100023153
$000000010002319A
Runtime error 203 at $0000000100009B72
$0000000100009B72
$00000001000225B6
$0000000100023153
$000000010002319A
The first output did work, sadly I deleted it. It was over 400mb, so approximately double the size of a zip or a tar.zst of the same files.zstd natively compresses a tar of this set to 209mb in 800ms in multi-threaded mode, or 3.5s in single threaded mode.
I suspect that sqlite is being held incorrectly (access from multiple threads, with multi-threading disabled), and the vfs lock forwarding is broken on Windows.
Build.sh can be used for Windows too, using MSYS2 UCRT64.
Data being packed was an unzipped copy of linux-master.zip fetched from GitHub unpacked with windows zip, selecting skip for the overlapping case files.
To be clear, you can run Pack as: `pack.exe ./linux-master/`
Plus there's no discussion against zstd itself and its container format.
ZStandard is Pareto optimal.
For the argument why, I really recommend this investigation.
https://insanity.industries/post/pareto-optimal-compression/
I wrote a similar general purpose pack way back in the early 90's in TopSpeed Modula-2 and run from the command line. Needed to span multiple disks and self launch. Algorithm was fast, but not nearly the same compression ratios. Wore out Mark Nelson's classic, "The Data Compression Book" along the way.
That seems fishy...
I know many people including myself are curious on why you wrote this in Pascal?
Also, what are the main reasons Pack is faster than the tools you compare Pack with?
About Pascal: It makes me happy. It looks clean and pseudocode-like. It helps readers from around the world with different languages understand. I am happy that Pack made people curious about this old but great goodie.
Speed, Here are some reasons: - Pack does all the steps of pack or unpack (read, check, compress/decompress and write) together and this weaving helped achieve this speed and random access. It is by far the fastest speed I get to see reading or writing random files from file systems, as fast or faster than asynchronous read operations or OVERLAPPED on Windows. To a point, it is limited to file system. For example, on NTFS, Pack can pack Linux code base in around 1.3 s; similar is done on ext4 in 0.96 s.
- It is based on a heavily optimized code base, standard library, and the FreePascal compiler, which produces great binary.
- Multi core design: even mobiles have multi-core CPUs these days. Choosing threads based on the content and machine, it does not eat your machine.
- Speed-configured SQLite. SQLite is much faster than most people think it is.
- Configured the already rapid Zstandard.
In summary, standing on the shoulders of giants while trying hard to improve reliability, speed and user experience is a sign of respect for them.
Would just remove that "accordion" functionality completely or make it always expanded on mobile breakpoints or whatever. Or just move that entire "About pack" section to be on the main page "below the fold" as the first thing people are going to want to do is find out *what it is* :)
If that is not enough, let me know.
7-Zip with the ZSTD patch is good too, but Pack is much faster at handling many files.
Testing packing the Linux code base (81K files and 1.25 GB) on Windows with NTFS:
7-Zip + Patched with ZSTD (-m0=Zstd): 6.453 s, 194.9 MB (Creating the header takes too much time)
Pack: 1.3 s, 194.5 MB
1) regular archive formats vs pack
2) various containers with the same zstd inside.
Exposure comes from enthusiasts like you.
I did not want to focus the point on speed, or say, "Look, others are bad". They are great; my point was, "Look what we can do if we update our design and code". Pack value comes from user experience, and speed is being one. I was not following the best speed or compression; I wanted an instantaneous feeling for most files. I wanted a better API, an easier CLI, improved OS integration (soon), and more safety and reliability. Tech people (including me) care so much about speed
I am happy about the results, but Pack offers much more that I like others to see.
In fact, SCL has exactly two libraries, SQLite and Zstandard, so presumably it's the same developer https://github.com/SCLOrganization/Libraries
It is the point: if you trust a project based on "who" made it, my friend, that is the start of the big problem we are facing in this current situation of tech. Just look at the code, build it yourself, and check the license.
Pack is made to be a private option; future locking end encryption options will solidify that. Trusting the author is not the correct way to verify the security and safety of such a tool.
Basically I want a shared encoding dictionary. Is there an easy solution for this?
The use case is maintaining an archive of .crx files.
Having a shared encoding dictionary is called compression.
If you want one across many zip files that you cannot modify, then compress them into another archive.
(At least... this is my assumption based on how I understand the formats to work. I do need to measure and verify this.)
So basically I'm wondering if there is some way to tell the "outer" zip to reuse the encoding dictionaries of each smaller zip, or to somehow intelligently merge their encoding dictionaries rather than treating each inner zip like an opaque blob.
Are there compression formats that do share an encoding dictionary across multiple files? I guess tar + gzip might do that?
However, those compressors generally only have a small context window, so they'll only be able to take advantage of relatively nearby redundancies in the archive. They won't help you if the common substrings are separated by many megabytes of unrelated data. So the order in which you pack the files matters.
In theory, you could do a two-pass approach, where you first scan the entire set of files to create a (relatively small) shared dictionary that's as useful as possible, and then uses that dictionary to compress each file independently. I don't know of any archive format that does that, but you could roll your own using the zstd command line compressor.
Your best bet is a lossless transform that undoes the huffman coding in the zip files, converting the compressed streams from effectively uncorrelated bitstreams to largely similar byte streams, and then pass that through a large-window compression algorithm (zstd?).
Similar techniques are used in ChromeOS for delta updates.
You could maintain a shared dictionary among all the archives.
If you have a way to post process the archives, you may be able to abuse multipart zip files to split them and reconstitute them later.
You should try them for yourself.
It looks clean and pseudocode-like. It helps readers from around the world with different languages understand.
Each binary has its own build script that you can use for yourself. Binaries are used for static builds and to ease future needs. https://github.com/PackOrganization/Pack/blob/main/Libraries...
I know you don't have a duty to look around for your answer, but you too don't have a duty to say yuck to a project that has been done with a lot of effort. I am ok with your comment, but maybe go easy on the next project.
yeah, no thanks. SQlite3 automatically means:
- Single implementation (yes, it's a nice one but still a single dependency)
- No way to write directly to pipe (SQlite requires real on-disk file)
- No standard way to read without getting the whole file first
- No guarantees in number of disk seeks required to open the file (relevant for NFS, sshfs or any other remote filesystem use)
- The archive file might be changed just by opening in read-only mode
- Damaged file recovery is very hard
- Writing is explicitly not protected against several common scenarios, like backup being taken in the middle of file write
- No parallel reads from multiple threads
Look, sqlite3 is great for it's designed purpose (embedded database). But trying to apply for other purposes is often a bad idea.
You can run SQLite with an In Memory database, I use it quite a lot for unit tests.
You can do this with both tar and ZIP. If all you have is SQLite, you need to fully create the local database (be it in a file or in memory) before you can transmit it somewhere else to be stored or unpacked.
pack archive | ssh host2 pack extract
Which is a weird gripe - in this use case, `tar` makes most sense to use. It's not like this format claims to be a full replacement for tar or anything close to that.
> Most popular solutions like Zip, gzip, tar, RAR, or 7-Zip are near or more than three decades old. While holding value for such a long time (in the computer world) is testimony to their greatness, the work is far from done.
> Pack tries to continue this path as the next step and proposes the next universal choice in the field.
About piping, it can be done, and it is on my list. I will finish features based on their popularity and making sense.
As Pack has random access support, you can choose a file in a big pack, and it can stream it out to the output. It is already able to unpack partially to your file system (using --include="file path in pack"); streaming/piping it would not be a problem.
Sqlite supports parallel reads from multiple threads.
It even supports parallel reads and writes from multiple threads.
What it doesn't really support is parallel reads and writes from multiple processes.
> SQlite3 automatically means:
...
> - No parallel reads from multiple threads
Note that when SQLite is compiled with SQLITE_THREADSAFE=0, the code to make SQLite threadsafe is omitted from the build. When this occurs, it is impossible to change the threading mode at start-time or run-time.
SQLite’s APIs are often hazardous in these ways, it really should be erroring rather than silently ignoring the fullmutex flag, but alas.I'd take this more seriously if the format was documented at all, but so far it appears to be "this implementation relies on sqlite and zstd therefore it's better", without even a specification of the sql schema, let alone anything else.
The github repo contains precompiled binaries of zstd and sqlite. The sqlite builds appear to have thread support disabled so not only will it be single writer it'll be single reader too.
The schema is missing strictly typed tables, and the implementation appears to lack explicit collation handling for names and content.
The described benchmark appears to involve files with an average size of 16KB. I suspect it was executed on Windows on top of NTFS with an AV package running, which is a pathological case for single threaded use of the POSIXy IO APIs that undoubtedly most of the selected implementations use.
It's slightly odd that it appears to perform better when SQLite is being built with thread safety disabled (https://github.com/PackOrganization/Pack/blob/main/Libraries...) and yet the implementation is inserting in a thread group: https://github.com/PackOrganization/Pack/blob/main/Source/Dr.... I suspect the answer here is that because the implementation is using a thread group to read files and compress chunks, it's amortizing the slow cost of file opens in this benchmark using threading, but is heavily constrained by the sqlite locking - and the compression ratio will take a substantial hit in some cases as a result of the limited range of each compression operation. I suspect that zstd(1) with -T0 would outperform this for speed and compression ratio, and it's already installed on a lot of systems - even Windows 11 gained native support for .zst files recently.
The premise that we could do with something more portable than TAR and with less baggage is somewhat reasonable - we probably could do with a simple, safe format. There are a lot more key considerations to making such a format good, such as many you outline, such as choices around seeking, syncing, incremental updates, compression efficiency, parallelism, etc. There is no single set of trade-offs to cover all cases but it would be possible to make a file format that can be shared among them, while constraining the design somewhat for safety and ease of portability.
- Single implementation: Sure. Working with SQLite convinced me that nobody cared to reimplement it, as it worked so well that nobody wanted or needed to rework it. I may write an unpacker just to prove that it is not hard at all to read SQLite format. The complicated part is the SQL engine (and many other features that are not used in Pack), and for Pack, you can live without it.
- SQLite does not require a disk. It has a memory option. Pack can have piping and properly will. I did not implement it because, well, it is too new, and I do what I feel is needed first. You can subscribe to the newsletter on the site (https://pack.ac/notes) or follow GitHub.
- Of course you can read SQLite without reading the whole file. It is a database, not a tar file.
- SQLite is highly optimized to read the lowest amount of data, and it has layers of smart caching. There is a reason it is used on almost any device that has a computer on it, even smartwatches.
- Of course the archive is safe from changes in unpacking. It will be opened in read-only mode, guarded by the OS and file system, and Pack also uses code isolation, which prevents calling write on any file.
- There are a lot of tools that help repair a damaged SQLite file. Pack too is guarded with transactions. The file will not get corrupted unless the disk goes corrupt; the mentioned tools come handy then. And in today's world of SSDs, the risk is shrinking rapidly.
- On unpack, Pack reads, decompresses, checks, and writes in a multithreaded. So yes, parallel reading is possible and done in Pack.
- I suggest trying Pack for yourself. It gives you the feeling you need to have to be sure.
A less common tar's feature is packing on compress -- stuff like "ssh remote tar cvf - ... > local-file.tar", which skips temporary file on remote machine, and also saves lots of time in transfer.
But for both of those, sqlite's "memory" won't help you there - memory or not, you still need to have the entire file to read it. So if you just store file contents in the sql database, then you have to fetch everything up to the latest byte before you can get any data out.
Maybe you can have index in sqlite, and append data as-is... but where would you put that index?
if you put it in front (like squashfs), you need to produce entire metadata before writing first data byte.. and that should include compressed sizes too (assuming you want to support random extraction), which means you cannot stream file out until you finish compressing all the data. And also sometimes you will not be able to add files to the archive without rewriting the whole archive (if the index grows and you didn't leave enough padding). This might be OK, but definitely should be mentioned.
If you put it at the end (like zip), you will be able to stream file out during compression, but fast decompression would be impossible. Also, you'll forego any sqlite transitional guarantees - since the database will be created in-memory, and only written at the very end once all the files are written.
So frankly, I don't see how you can win on a streaming front, unless you really have a custom format and "sqlite3" is just a small part of it.
(Another problem is there is not even a short spec - how is sqlite3 used, what is your schema, and so on. And I am sorry, but I am not going to read the source code just to figure this stuff out).
Really seems to make sense! For another fun compression trick: https://github.com/mxmlnkn/ratarmount