HNHacker News
TopNewBestAskShowJobs

OttoCoddo

18 karma · joined May 16, 2022

Creator of Pack: https://pack.ac
submissionscomments
OttoCoddo··on SQLite: 35% Faster Than the Filesystem
With the right code, NTFS is not much slower than ext4, for example. Nearly 3% for tar.gz and more than 30% for a heavily multi-threaded use like Pack.

https://forum.lazarus.freepascal.org/index.php/topic,66281.m...

OttoCoddo··on SQLite: 35% Faster Than the Filesystem
SQLite can be faster than FileSystem for small files. For big files, it can do more than 1 GB/s. On Pack [1], I benchmarked these speeds, and you can go very fast. It can be even 2X faster than tar [2].

In my opinion, SQLite can be faster in big reads and writes too, but the team didn't optimise it as much (like loading the whole content into memory) as maybe it was not the main use of the project. My hope is that we will see even faster speeds in the future.

[1] https://pack.ac [2] https://forum.lazarus.freepascal.org/index.php/topic,66281.m...

OttoCoddo··on Pack: A new container format for compressed files
Thank you for the notes. I am well aware of the levels and Pack uses custom configuration to match its inner design. Maybe more level come, or maybe not. But to be clear, Pack supports any valid Zstandard content, and this levels we are discussing are about Pack CLI chosen for better user experience. Any other client can produce and store any valid content for chose level or configuration, and other clients can read it.
OttoCoddo··on Pack: A new container format for compressed files
Then `--press=hard` would be the choice for you.
OttoCoddo··on Chrome Feature: ZSTD Content-Encoding
Sane choice, albeit delayed, as Zstandard has been the leading standard algorithm in the field for quite some time. I tested most of them for developing Pack [1] and Zstandard looked like the best alternative to the good old DEFLATE.

The dictionary feature [2] will help design some new ways of getting small resources.

[1] https://news.ycombinator.com/item?id=39793805

[2] https://facebook.github.io/zstd/zstd_manual.html#Chapter10

OttoCoddo··on Postgres vs. File Systems: A Performance Comparison (2022)
Readers may find Pack interesting: https://news.ycombinator.com/item?id=39793805
OttoCoddo··on Pack: A new container format for compressed files
Thank you for the new numbers. Sure, it can be different on different machines, especially full systems. For me on Linux and ext4, Pack finishes the Linux code base at just 0.96 s.

Anyway, I do not expect an order of magnitude difference between tar.zst and Pack; after all, Pack is using Zstandard. What makes Pack fundamentally different from tar.zst is Random Access and other important factors like user experience. I shared some numbers on it here: https://news.ycombinator.com/item?id=39803968 and you are encouraged to try them for yourself. Also, by adding Encryption and Locking to Pack, Random Access will be even more beneficial.

OttoCoddo··on Pack: A new container format for compressed files
Easy to build using this document: https://pack.ac/source

Each binary has its own build script that you can use for yourself. Binaries are used for static builds and to ease future needs. https://github.com/PackOrganization/Pack/blob/main/Libraries...

I know you don't have a duty to look around for your answer, but you too don't have a duty to say yuck to a project that has been done with a lot of effort. I am ok with your comment, but maybe go easy on the next project.

OttoCoddo··on Pack: A new container format for compressed files
Yes, I am. Pack can hold millions of files with no problem. One field in which it shines is that, aside from being fast at processing large amounts of data, it can process many small files much faster than similar tools or even many popular file systems.

About piping, it can be done, and it is on my list. I will finish features based on their popularity and making sense.

As Pack has random access support, you can choose a file in a big pack, and it can stream it out to the output. It is already able to unpack partially to your file system (using --include="file path in pack"); streaming/piping it would not be a problem.

OttoCoddo··on Pack: A new container format for compressed files
It makes me happy.

It looks clean and pseudocode-like. It helps readers from around the world with different languages understand.

OttoCoddo··on Pack: A new container format for compressed files
You got a point. Although with that that option comes a great cost: We will lose portability, speed and even reliability.

Portability: Receiver (or future you) needs to know what you used, and what version even.

Speed: If you want to do the archive part first (tar) and then compress (gz), you will get much lower speed (as shown in the note).

Reliability: Most people use tar with gz anyway, but if you use it with not so popular algorithm and tools, you will risk having a file that may or may not work into the future.

Pack plan is to use the best of time (Zstandard) and if an update is needed in years to come, it will add support for the new algorithm updates. All Pack clients must only write the latest version (and read all previous versions) and that makes sure almost all use the best of their time.

OttoCoddo··on Pack: A new container format for compressed files
What are the parameters you gave to the CLI program? This issue seems interesting, as these files on Windows 11 were tested countless times.

To be clear, you can run Pack as: `pack.exe ./linux-master/`

OttoCoddo··on Pack: A new container format for compressed files
You can do that with Pack:

`pack -i ./test.pack --include=/a/file.txt`

or a couple files and folders at once:

`pack -i ./test.pack --include=/a/file.txt --include=/a/folder/`

Use `--list` to get a list of all files:

`pack -i ./test.pack --list`

Such random access using `--include` is very fast. As an example, if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.

OttoCoddo··on Pack: A new container format for compressed files
Pack binary? Can you tell what machine and what steps?

Build.sh can be used for Windows too, using MSYS2 UCRT64.

OttoCoddo··on Pack: A new container format for compressed files
It is not disabled, as you think. It is secured by the internal code in Pack. Almost all parts of Pack are multithreaded.
OttoCoddo··on Pack: A new container format for compressed files
Thank you!

Exposure comes from enthusiasts like you.

I did not want to focus the point on speed, or say, "Look, others are bad". They are great; my point was, "Look what we can do if we update our design and code". Pack value comes from user experience, and speed is being one. I was not following the best speed or compression; I wanted an instantaneous feeling for most files. I wanted a better API, an easier CLI, improved OS integration (soon), and more safety and reliability. Tech people (including me) care so much about speed

I am happy about the results, but Pack offers much more that I like others to see.

OttoCoddo··on Pack: A new container format for compressed files
I am happy to hear that, and I really appreciate your interest.

Did you compile it for yourself? Any problem or steps you used, I will be happy to hear, o at pack.ac or GitHub, as it is hard to follow the building here.

As a reminder, Pack Draft 0 has Compatibility with SQLite tools; the only needed step is to change the first 16 bytes. Again, you can use `--activate-other-options --transform-to-sqlite3` with the CLI tool, and you will get a perfectly working SQLite file.

VFS is not needed; they can change the header after writing; VFS was just cleaner to me.

My first work was using application_id, after a while, it did not feel right to me, so I changed it for good. It allows easier future development, fewer problems for file type detection, a decreased chance of mistaken change (you already saw many negative comments on using SQLite as a base), and the support reason: just yesterday I was reading a forum post about people asking for support on software because it was using SQLite. application_id seems like a great choice if you are doing a DB-related task or making a custom DB for transfer on wire, to communicate between internal and semi-public tools. Using it for a format that could potentially get to an innumerable count seemed unwise.

OttoCoddo··on Pack: A new container format for compressed files
Did you compile it for yourself? Any problem or steps you used, I will be happy to hear, o at pack.ac or GitHub, as it is hard to follow the building here. I should prepare more documents on how to build it. I suspect that there is a problem with the custom build. Error and speed issues are not something you see in the official build.
OttoCoddo··on Pack: A new container format for compressed files
You are one of the bravest. And you know that, using SQLite as the base storage, rules out many of the security problems we can face.

Anyone needing to reimplement Pack, can do it, very easily, if not easier than implementing ZIP, IF they use SQLite and Zstandard. Maybe a day of work or less. If they want to rewrite (reading part of) them too, it will be a couple of days of work.

OttoCoddo··on Pack: A new container format for compressed files
I answered this question here: https://news.ycombinator.com/item?id=39801083

If that is not enough, let me know.

7-Zip with the ZSTD patch is good too, but Pack is much faster at handling many files.

Testing packing the Linux code base (81K files and 1.25 GB) on Windows with NTFS:

7-Zip + Patched with ZSTD (-m0=Zstd): 6.453 s, 194.9 MB (Creating the header takes too much time)

Pack: 1.3 s, 194.5 MB

OttoCoddo··on Pack: A new container format for compressed files
Thank you very much! I guess you are friend with colors ;)
OttoCoddo··on Pack: A new container format for compressed files
Indeed. Benchmarking Pack as a file system was fun. It is near 10X faster to let you iterate all the files compared to what I get from NTFS (warm with all caching on for both solutions).

Someday, it can be used as a virtual drive. I leave it to future people.

OttoCoddo··on Pack: A new container format for compressed files
I answered this question here: https://news.ycombinator.com/item?id=39801083 If that is not enough, let me know.
OttoCoddo··on Pack: A new container format for compressed files
You may like to read https://pack.ac/note/pack and test it for yourself.
OttoCoddo··on Pack: A new container format for compressed files
Hello, and thank you for the notes. Unfortunately, your points seem to be mostly wrong, so let me clarify them a little. Do not worry; many people misunderstand SQLite and its abilities.

- Single implementation: Sure. Working with SQLite convinced me that nobody cared to reimplement it, as it worked so well that nobody wanted or needed to rework it. I may write an unpacker just to prove that it is not hard at all to read SQLite format. The complicated part is the SQL engine (and many other features that are not used in Pack), and for Pack, you can live without it.

- SQLite does not require a disk. It has a memory option. Pack can have piping and properly will. I did not implement it because, well, it is too new, and I do what I feel is needed first. You can subscribe to the newsletter on the site (https://pack.ac/notes) or follow GitHub.

- Of course you can read SQLite without reading the whole file. It is a database, not a tar file.

- SQLite is highly optimized to read the lowest amount of data, and it has layers of smart caching. There is a reason it is used on almost any device that has a computer on it, even smartwatches.

- Of course the archive is safe from changes in unpacking. It will be opened in read-only mode, guarded by the OS and file system, and Pack also uses code isolation, which prevents calling write on any file.

- There are a lot of tools that help repair a damaged SQLite file. Pack too is guarded with transactions. The file will not get corrupted unless the disk goes corrupt; the mentioned tools come handy then. And in today's world of SSDs, the risk is shrinking rapidly.

- On unpack, Pack reads, decompresses, checks, and writes in a multithreaded. So yes, parallel reading is possible and done in Pack.

- I suggest trying Pack for yourself. It gives you the feeling you need to have to be sure.

OttoCoddo··on Pack: A new container format for compressed files
As far as I know, PAQ strives to give the best compression for archival purposes. Pack, on the other hand, tries to give the best compression, while keeping speed as instantaneous as possible.

You should try them for yourself.

OttoCoddo··on Pack: A new container format for compressed files
Hell yeah indeed! You can find me on the forum too if you liked to talk Pascal: https://forum.lazarus.freepascal.org/index.php/topic,66281.0...
OttoCoddo··on Pack: A new container format for compressed files
Sorry. Site is very new, and custom made, and needs to be worked on mobile.
OttoCoddo··on Pack: A new container format for compressed files
I just do not want to follow Wirth's law.
OttoCoddo··on Pack: A new container format for compressed files
It looks clean and pseudocode-like. It helps readers from around the world with different languages understand. FreePascal compiler is very good too.
Page 1 of 2Next →