HNHacker News
TopNewBestAskShowJobs

OttoCoddo

18 karma · joined May 16, 2022

Creator of Pack: https://pack.ac
submissionscomments
OttoCoddo··on Pack: A new container format for compressed files
Hey,

The choice of parameters was solely done to be clear, and not what people used to. -f meaning force is not clear; -w meaning overwrite, seemed like a better logical choice, to me.

Nice point on -h. Yes, I did not want to go crazy. After all, almost all (CLI) people use pack as `pack ./test/`. Options are for advanced people like you. Most people will use the OS integration that will be published later on.

--press=hard is the only option there is. There may be more, but with Pack you do not need to choose a level (like 1..9 with ZIP). Just let Pack do its thing, and you will be happy. Hard Press is there for people who want to pack once and unpack many times (like publishing), and it is worth spending extra time on it. Even then, Pack goes the sane way and does not eat your computer just for a kilobyte or two.

OttoCoddo··on Pack: A new container format for compressed files
More documentation will be published soon. For now and about CLI: https://pack.ac/cli-documentation
OttoCoddo··on Pack: A new container format for compressed files
Reading many files (81K in this test) is way slower than reading just one big file. For bigger files, Pack is much faster. Here is a link to some results from a kind contributor: https://forum.lazarus.freepascal.org/index.php/topic,66281.m...

(Too long to post here)

OttoCoddo··on Pack: A new container format for compressed files
Using 32GB of RAM, but it is far more than they need.

7-Zip was used as others, just gave it a folder to compress. No configuration.

As requested, here are some numbers on tar.zst of Linux source code (the test subject in the note): tar.zst: 196 MB, 5420 ms (using out-of-the box config and -T0 to let it use all the cores. Without it, it would be, 7570 ms) Pack: 194 MB, 1300 ms Slightly smaller size, and more than 4X faster. (Again, it is on my machine; you need to try it for yourself.) Honestly, ZSTD is great. Tar is slowing it down (because of its old design and being one thread). And it is done in two steps: first creating tar and then compression. Pack does all the steps (read, check, compress, and write) together, and this weaving helped achieve this speed and random access.

OttoCoddo··on Pack: A new container format for compressed files
Me.

It is the point: if you trust a project based on "who" made it, my friend, that is the start of the big problem we are facing in this current situation of tech. Just look at the code, build it yourself, and check the license.

Pack is made to be a private option; future locking end encryption options will solidify that. Trusting the author is not the correct way to verify the security and safety of such a tool.

OttoCoddo··on Pack: A new container format for compressed files
Thank you so much for the kind words, refreshing.

Here is the latest sqlar result on Linux source code on the same test machine in warm state:

sqlar: 268 MB, 30.5 s

Pack: 194 MB, 1.3 s

Very good result compared to tar.gz. And much better than ZIP, considering sqlar gives random access like ZIP and unlike tar.gz. I considered sqlar as a proof of concept, and it inspired me to create Pack as a full solution. I always agreed with the great drh (creator of SQLite) points about SQLite as a file format, and Pack is a try to demonstrate that.

I made Pack to give people a better life (at least behind their desks), and as you do, I hope people get to use it and find it useful.

OttoCoddo··on Pack: A new container format for compressed files
Hello David, and thank you for your comment, analysis, and the issues you opened. I will get to them all.

- SQLite tooling: You will not need it unless you are debugging something, then you can change the header or just use the `--activate-other-options --transform-to-sqlite3` parameter to transform a Pack file to SQLite3, and use the `--activate-other-options --transform-to-pack` to go back. This way, you get a true SQLite3 database that you can browse as you wish. For most people, mixing Pack with SQLite was just a call for problems for the SQLite team (imagine people coming and asking to fix their Pack file from the team; that would not be fair) and a harder future for Pack to update.

- Metadata is not stored in Pack. I don’t want the metadata of my machine attached to a file. It’s a never-ending nightmare to match source and destination OS metadata. There will always be something missing, and Pack tends to get everything perfect or nothing. Storing metadata adds extra weight that most users don’t care about and complicates the ability to store other types of data alongside files. It may get added as an option in the future if many people need it.

- There is a version field. It is currently in Draft 0, and it is written using a custom VFS. Look here for more information: https://github.com/PackOrganization/Pack/blob/main/Source/Dr...

- All future versions of Pack must handle previous versions and must only write the latest version. So any files created right now (Draft 0) will be read correctly for ever to come.

- Each Draft proposal will get its own version, and if it gets final, it will be set to final.

- Two byte after 'Pack' header in little endian as (1 (Draft) shl 13 + 0 (version 0) = 8192). Final would be 0, so the first Final version will be 0 shl 13 + 1 = 1. and the second will be 2. It is by design, so any Draft version gets a higher number, preventing future mixups.

- 8 MB chunks are the default; Pack may choose smaller or bigger (16 MB for many small files or 32 MB for Hard Press).

- Random access is proper as unpacking steps take into account what you want and decompress a Content just once for many neighbouring files. But even for reading just one file, here is an example: if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.

OttoCoddo··on Pack: A new container format for compressed files
Thank you for the check.

No index, as they take space, and I wrote the queries considering SQLite automatic indexes. They will be created on demand, at unpacking time. All the unpacking processes are made to read and decompress content just once, so there are no worries about slowdowns.

I suggest trying Pack for yourself and seeing the speed. Or deeper, use `--activate-other-options --transform-to-sqlite3` to transform a Pack file to SQLite3, create your own indexes, and use `--activate-other-options --transform-to-pack` to convert it to Pack and then try unpacking. You will not see any worthy difference.

Yes, Contents are like packages of raw data from a chunk, a whole, or many of the items (files or data). They may be compressed if needed (With Zstandard). ItemContent table helps to find the needed Item parts.

The Content structure circumvents any BLOB limit, but it is also made to give better compression while keeping random access.

OttoCoddo··on Pack: A new container format for compressed files
I guess you are overestimating the "cobble together read/write support without even needing a library."

Let's imagine: You want to read a ZIP file. Will you write your own reader? I seriously doubt it, as the work, stabilising, and security (random memory access as an example) would be issues. But let's think we are couraginous. OK, we read rather not so simple format and carefully read the binary. Now, will you write your own DEFLATE and Huffman coding? Again, a bigger doubt.

I would argue that if someone cares enough to reimplement ZIP, it would at worst be twice as hard to write a Pack reader from scratch with no ZSTD or SQLite. And for those serious people, reading a format that lets them store better and faster would be a prize that is hard to say no to. But I get your point, and if you are in a desert and need something to put together fast before going out of water, tar may be a good choice.

OttoCoddo··on Pack: A new container format for compressed files
A couple, but simple.

No, you will only need Pack, everything is built into it. Pack is built for Windows and Linux, and more will come. You will be able to run it on almost all CPUs.

OttoCoddo··on Pack: A new container format for compressed files
Good point, thank you. Note that SQLite format is very simple https://www.sqlite.org/fileformat.html
OttoCoddo··on Pack: A new container format for compressed files
It was hard to believe for me, too. And I didn't stumble upon it; I looked for it closely, and that was a point in the note. People did not look properly for nearly three decades. Many things have changed, but we computer people are still using the same tools. I am not saying old is not good; the current solutions are great, but what are we, if we don't look for the better?

Yes it is that much faster, and a good part of it is because of the multi-thread design, but as a reminder, WinRAR or 7-Zip are too multi-thread, and you can see the difference. To satisfy your doubt, I suggest running Pack for yourself. I am looking for more data on its behaviour on different machines and data.

Can I ask why do you need a version without ZSTD? If you are thinking that compression slows it down, I should say no. Pack is the first of its kinds that "Store" is slowing it down. Because its compression is smart, it will skip any non-compressible content.

On the same machine and the same Linux source code test:

Pack: 194 MB, 1.3 s

Pack (With no Press): 1.25 GB, 1.8 s

OttoCoddo··on Pack: A new container format for compressed files
That is the best joke I've heard all day. Thank you for the laugh :)
OttoCoddo··on Pack: A new container format for compressed files
Yes, and I hope that is a good surprise. As you can see, you can create fast and readable codes with it.
OttoCoddo··on Pack: A new container format for compressed files
Hey fellow enthusiast.

- Piping is really easy and it will get added to Pack. It is matter of time, until these features get added as they will be added based on popularity and piping is not that popular for most people. But I get you and I will add it for you.

- Metadata is not stored in Pack. I don’t want the metadata of my machine attached to a file. It’s a never-ending nightmare to match source and destination OS metadata. There will always be something missing, and Pack tends to get everything perfect or nothing. Storing metadata adds extra weight that most users don’t care about and complicates the ability to store other types of data alongside files. It may get added as an option in the future if many people need it.

- Pack uses Zstandard under the hood. Great compression speed and ratio. In my opinion, it is the leading algorithm in the field and makes it a proper choice to use instead of DEFLATE, used in ZIP or GZIP.

- At this point you are telling tar features. tar is not random access, Pack is. As an example, if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.

OttoCoddo··on Pack: A new container format for compressed files
You can use Pack for those cases too. --press=hard creates a more compressed pack for cases of pack once, unpack many.
OttoCoddo··on Pack: A new container format for compressed files
Thanks for the note, and sorry for the inconveniences. I did not expect this many users in the iOS world. The site is very new and needs custom work; it will be updated soon.
OttoCoddo··on Pack: A new container format for compressed files
Thank you!

About Pascal: It makes me happy. It looks clean and pseudocode-like. It helps readers from around the world with different languages understand. I am happy that Pack made people curious about this old but great goodie.

Speed, Here are some reasons: - Pack does all the steps of pack or unpack (read, check, compress/decompress and write) together and this weaving helped achieve this speed and random access. It is by far the fastest speed I get to see reading or writing random files from file systems, as fast or faster than asynchronous read operations or OVERLAPPED on Windows. To a point, it is limited to file system. For example, on NTFS, Pack can pack Linux code base in around 1.3 s; similar is done on ext4 in 0.96 s.

- It is based on a heavily optimized code base, standard library, and the FreePascal compiler, which produces great binary.

- Multi core design: even mobiles have multi-core CPUs these days. Choosing threads based on the content and machine, it does not eat your machine.

- Speed-configured SQLite. SQLite is much faster than most people think it is.

- Configured the already rapid Zstandard.

In summary, standing on the shoulders of giants while trying hard to improve reliability, speed and user experience is a sign of respect for them.

OttoCoddo··on Pack: A new container format for compressed files
Thank you for the detail check. I should thank the syrup too :)

I'm happy to see a fellow enthusiast. Your deduction is on point. And also, Pack is smart; it skips non-compressible files like MP3 [1], so you do not need to choose the "Store" option to have a faster option, and it speedup decompression too. Pack is the first to achieve this, being faster than Store options. Yes, it was a surprise to me too.

ZPAQ is great, and I study the Hutter Prize competition. Pack is on another chart, which is why I proposed CompressedSpeed [2]. The speed of getting to compression needs to be accounted for. You can store anything on an atom if you try hard enough, but hard work takes time. Deduplication step may get added, but in Hard Press [3].

I am curious to see the results of Pack on your data. You can find me here or o at pack.ac.

[1] It is based on content rather than extension; any data that is determined not to be worthy of compression, will be stored as is. And as a file can get chucked, some parts can get compressed and some cannot. Imagine that part of the subtitle in a MKV file can get compressed, and the Video part gets skipped. Although these features will get more updates over time, if they don't cost time,. Pack focus is being seamless and not the most compressed; there are already great works in the field, such as the noted ZPAQ.

[2] CompressedSpeed = (InputSize / OutputSize) * (InputSize / Speed). Materialized compression speed.

[3] You can choose --press=hard to ask for better compression. Even with Hard Press, Pack does not try to eat your hardware just to get a little more; it goes the optimized way I described.

OttoCoddo··on Pack: A new container format for compressed files
- As far as I know, squashfs is a file system and not an archive format; the "FS" in the name shows the focus.

- It is read-only; Pack is not. Update and delete are not just public yet, as I wanted people to get the taste first.

- It is clearly focused on archiving, rather than Pack wanting to be a container option for people who want to pack some files/data and store or send them with no privacy dangers.

- Pack is designed to be user-friendly for most people; CLI is very simple to work with, and future OS integration will make working with it like a breeze. It is far different from a good file system focused on Linux.

- I did not compare to squashfs, but I will be happy to see any results from interested people.

My bet is on Pack, obviously, to be much faster.

OttoCoddo··on Pack: A new container format for compressed files
Hello to all. I am the author, and I just saw this post and am happy to see this exciting discussion. Let me try to show my respect for it and answer as well as possible.
← PreviousPage 2 of 2