https://forum.lazarus.freepascal.org/index.php/topic,66281.m...
18 karma · joined May 16, 2022
https://forum.lazarus.freepascal.org/index.php/topic,66281.m...
In my opinion, SQLite can be faster in big reads and writes too, but the team didn't optimise it as much (like loading the whole content into memory) as maybe it was not the main use of the project. My hope is that we will see even faster speeds in the future.
[1] https://pack.ac [2] https://forum.lazarus.freepascal.org/index.php/topic,66281.m...
The dictionary feature [2] will help design some new ways of getting small resources.
[1] https://news.ycombinator.com/item?id=39793805
[2] https://facebook.github.io/zstd/zstd_manual.html#Chapter10
Anyway, I do not expect an order of magnitude difference between tar.zst and Pack; after all, Pack is using Zstandard. What makes Pack fundamentally different from tar.zst is Random Access and other important factors like user experience. I shared some numbers on it here: https://news.ycombinator.com/item?id=39803968 and you are encouraged to try them for yourself. Also, by adding Encryption and Locking to Pack, Random Access will be even more beneficial.
Each binary has its own build script that you can use for yourself. Binaries are used for static builds and to ease future needs. https://github.com/PackOrganization/Pack/blob/main/Libraries...
I know you don't have a duty to look around for your answer, but you too don't have a duty to say yuck to a project that has been done with a lot of effort. I am ok with your comment, but maybe go easy on the next project.
About piping, it can be done, and it is on my list. I will finish features based on their popularity and making sense.
As Pack has random access support, you can choose a file in a big pack, and it can stream it out to the output. It is already able to unpack partially to your file system (using --include="file path in pack"); streaming/piping it would not be a problem.
It looks clean and pseudocode-like. It helps readers from around the world with different languages understand.
Portability: Receiver (or future you) needs to know what you used, and what version even.
Speed: If you want to do the archive part first (tar) and then compress (gz), you will get much lower speed (as shown in the note).
Reliability: Most people use tar with gz anyway, but if you use it with not so popular algorithm and tools, you will risk having a file that may or may not work into the future.
Pack plan is to use the best of time (Zstandard) and if an update is needed in years to come, it will add support for the new algorithm updates. All Pack clients must only write the latest version (and read all previous versions) and that makes sure almost all use the best of their time.
To be clear, you can run Pack as: `pack.exe ./linux-master/`
`pack -i ./test.pack --include=/a/file.txt`
or a couple files and folders at once:
`pack -i ./test.pack --include=/a/file.txt --include=/a/folder/`
Use `--list` to get a list of all files:
`pack -i ./test.pack --list`
Such random access using `--include` is very fast. As an example, if I want to extract just a .c file from the whole codebase of Linux, it can be done (on my machine) in 30 ms, compared to near 500 ms for WinRAR or 2500 ms for tar.gz. And it will just worsen when you count encryption. For now, Pack encryption is not public, but when it is, you can access a file in a locked Pack file in a matter of milliseconds rather than seconds.
Build.sh can be used for Windows too, using MSYS2 UCRT64.
Exposure comes from enthusiasts like you.
I did not want to focus the point on speed, or say, "Look, others are bad". They are great; my point was, "Look what we can do if we update our design and code". Pack value comes from user experience, and speed is being one. I was not following the best speed or compression; I wanted an instantaneous feeling for most files. I wanted a better API, an easier CLI, improved OS integration (soon), and more safety and reliability. Tech people (including me) care so much about speed
I am happy about the results, but Pack offers much more that I like others to see.
Did you compile it for yourself? Any problem or steps you used, I will be happy to hear, o at pack.ac or GitHub, as it is hard to follow the building here.
As a reminder, Pack Draft 0 has Compatibility with SQLite tools; the only needed step is to change the first 16 bytes. Again, you can use `--activate-other-options --transform-to-sqlite3` with the CLI tool, and you will get a perfectly working SQLite file.
VFS is not needed; they can change the header after writing; VFS was just cleaner to me.
My first work was using application_id, after a while, it did not feel right to me, so I changed it for good. It allows easier future development, fewer problems for file type detection, a decreased chance of mistaken change (you already saw many negative comments on using SQLite as a base), and the support reason: just yesterday I was reading a forum post about people asking for support on software because it was using SQLite. application_id seems like a great choice if you are doing a DB-related task or making a custom DB for transfer on wire, to communicate between internal and semi-public tools. Using it for a format that could potentially get to an innumerable count seemed unwise.
Anyone needing to reimplement Pack, can do it, very easily, if not easier than implementing ZIP, IF they use SQLite and Zstandard. Maybe a day of work or less. If they want to rewrite (reading part of) them too, it will be a couple of days of work.
If that is not enough, let me know.
7-Zip with the ZSTD patch is good too, but Pack is much faster at handling many files.
Testing packing the Linux code base (81K files and 1.25 GB) on Windows with NTFS:
7-Zip + Patched with ZSTD (-m0=Zstd): 6.453 s, 194.9 MB (Creating the header takes too much time)
Pack: 1.3 s, 194.5 MB
Someday, it can be used as a virtual drive. I leave it to future people.
- Single implementation: Sure. Working with SQLite convinced me that nobody cared to reimplement it, as it worked so well that nobody wanted or needed to rework it. I may write an unpacker just to prove that it is not hard at all to read SQLite format. The complicated part is the SQL engine (and many other features that are not used in Pack), and for Pack, you can live without it.
- SQLite does not require a disk. It has a memory option. Pack can have piping and properly will. I did not implement it because, well, it is too new, and I do what I feel is needed first. You can subscribe to the newsletter on the site (https://pack.ac/notes) or follow GitHub.
- Of course you can read SQLite without reading the whole file. It is a database, not a tar file.
- SQLite is highly optimized to read the lowest amount of data, and it has layers of smart caching. There is a reason it is used on almost any device that has a computer on it, even smartwatches.
- Of course the archive is safe from changes in unpacking. It will be opened in read-only mode, guarded by the OS and file system, and Pack also uses code isolation, which prevents calling write on any file.
- There are a lot of tools that help repair a damaged SQLite file. Pack too is guarded with transactions. The file will not get corrupted unless the disk goes corrupt; the mentioned tools come handy then. And in today's world of SSDs, the risk is shrinking rapidly.
- On unpack, Pack reads, decompresses, checks, and writes in a multithreaded. So yes, parallel reading is possible and done in Pack.
- I suggest trying Pack for yourself. It gives you the feeling you need to have to be sure.
You should try them for yourself.