Honey, I shrunk the NPM package
jamiemagee.co.uk
jamiemagee.co.uk
In part (3), because its running only on lib/npm.js which is 13KB, you are getting skewed results which aren't directly applicable to the compression of npm-9.7.1.tar. Brotli excels at compressing small Javascript files, as this is where its dictionary provides the most benefit. The benefits of the dictionary for a large tar file will be negligible.
However, in the npm-9.7.1.tar scenario we still expect Brotli level 11 to produce slightly smaller files than Zstd level 19. Likely ~5% smaller. But we do expect Zstd to provide significantly faster decompression speed.
The interesting part of the article to me is: okay, now that we have a better compression option, how do we deploy it to an existing ecosystem? And suddenly we're looking at a 4 year migration path!
No easy win at all.
Amdahl's Law in action.
For anyone else who didn't know.
As opposed to... just using the npm-9.7.1.tar file that the other tests were just using (sorta barely internally to tar, but I don't think tar does any fancy streaming or anything if you're passing something via --use-compress-program, certainly nothing that would skew the results more than replacing all of npm.tar with just npm.js.)
In my local install, npm.tar is 25MB. npm.js is 16KB. It may not change the final outcome, but the data in the article do not support the conclusion. I would strongly suggest tarring up the npm directory and rerunning lzbench.
The analysis would have been much simpler, clearer and more focused on the compressor if they just benchmarked the compression with an already-generated tar archive.
(Notably zip is different where files are compressed individually in the archive)
Create a shared Brotli dictionary (or zstandard or whatever) based on the top NPM packages by download bandwidth and then have all npm packages compressed using it.
I think this can be done server side by npmjs.org, where NPM packages are recompressed in this fashion after upload using the shared dictionary, thus it is an optional feature and fully backwards compatible.
Riffing on this idea because of this new chrome feature, which does this in a flexible fashion: https://chromestatus.com/feature/5124977788977152
EDIT: This recompression of packages may be insecure as the digital signature of the package no longer aligns, but then the trick is to sign the contents of the package rather than package itself.
Wait, does npm have digital signatures at all? I sort of assumed it did, but does it really?
Maybe a small dictionary would save you 20% though.
It would be a fun experiment to figure out the compression ratio of say the top 1000 packages given a shared brotli dictionary of X size. Just keep increases the dictionary size until you see diminishing returns.
My estimate is based on NPM packages contain a bunch of super stereotypical files that if they are used to create the dictionary likely result in amazing compression ratios: package.json, package-lock.json, CSS, Tailwind, Bootstrap .gitignore, LICENSE, .eslint, README, React/Vue/Angular code...
Although the bandwidth savings would only be the size of that dictionary.
Still, I can't help but think there's ways to improve size & bandwidth usage by a lot, even besides using a different compression algorithm; a non-HTTP transfer method, for example.
https://gist.github.com/klauspost/2900d5ba6f9b65d69c8e
brotli literally already has a tokens for function/return/throw/indexOf(/.match/.length/etc.
Also verify after decompress is not without tradeoffs. On one hand we have folks like github who can't change the version of zlib because people rely on identical .tar.gz. https://news.ycombinator.com/item?id=34586917
On the other hand we have a whole lot of iffy stuff you can do to make programs decompressing content use large amounts of resources https://en.wikipedia.org/wiki/Zip_bomb which makes "decompress this potentially untrusted file so that I can validate it's safe to use" hard.
People rely on the same checksums for the same files. They’re perfectly fine with changing the method for new files.
I guess you could base it on the commit time, but this is user supplied.
Yeah, I see it already has a lot of JavaScript, HTML and CSS content. Interesting. I didn't realize it had an existing web-focused token library, and figured it was more like zstd, 7z and zlib, which I believe have none.
I would love to do the experiment if I had time. I wonder what is the laziest way to do it?
https://github.com/facebook/zstd/blob/dev/programs/zstd.1.md...
Not sure about brotli though.
There’s a smarter mechanism than that: npm already caches packages locally twice, both in node_modules and in a global location.
In reality you’re probably never going to install every dependency from scratch every day, so you’re already using either cache on every install.
What I’d like to see instead is a pre-resolution of packages done by the registry. I have a list of 30 dependencies, please resolve the tree and send it all over instead of forcing me to do a waterfall of fetches.
I am definitely not saying download 100MB to save a few MB per user. That is just dumb.
I find NPM in the last year or so is faster to resolve that yarn. Although I think bun is even faster and may have better global caching.
IIRC the main "innovation" that Brotli brings to compression is a default dictionary trained on web content. So if you are replacing this you best use the better algorithm.
I make a periodic practice of searching our node_modules folder for files that shouldn't be there and reporting bugs against the offending projects.
Usually that's been pretty effective, and now the total cruft is around a megabyte whereas before it was somewhere north of 50MB all told. (Conditions apply).
coverage reports, test results, build detritus, etc.
The one I'm still debating, because it's becoming a serious problem for a couple of our libraries: should the tests be included in .npmignore or kept along with the library? I'm not sure what the right answer is there. Test sizes especially including fixtures can creep up quite a lot over time. I know what I'd like it to be, but I'm not sure I can win that argument with a bunch of maintainers on different projects.
What are the reasons it would be a good idea to include the tests with the release/distribution of a package?
Seems like something you don't care about when you're just using it as a library, unless you want to modify something in it, but then you'll clone the library straight from a repository anyways, which includes the tests.
For packages where I do include tests, I've had at least one user request that I remove tests so that the footprint of the Docker image they're building is smaller.
Both are entirely reasonable requests, but package repositories don't really provide a good way of accommodating both at the same time, for instance, by allowing a separate upload of the dev gubbins such as tests.
It's easy enough to have things work in pre-prod and fail in prod without running slightly different code between them.
I think there is a solution to this, but it's going to require that we change to something a lot more beefy than semver to define interface compatibility. Semver is a gentlemen's agreement to adhere to the Liskov Substitution Principle. We are none of us gentlemen, least of all when considered together.
You have a software project, with a build process, and the "output" or final product of that project is the library that gets uploaded to NPM.
If they are packaging a software library, they should do it from the project's repository, not from one of its output artifacts.
They would probably reject a request if someone who was downstream of their work decided to repackage their stuff and asked them to include tests and other superfluous content on their packages.
All of which, if you ask me, is the correct way of doing any kind of packaging. Following that, IMO the same should be done for JavaScript libraries: the packaging should be done by cloning the project repo and adding whatever packager-specific files in there.
Notice in your link how in the upper part it says: [ Source: jq ], where "jq" is a link to the source description. In that page, the right hand section will now contain a link to the Debian copy repository where packaging is done:
https://salsa.debian.org/debian/jq
You can clone and explore the branches of that repo.
(Maybe you are a Debian maintainer, in any case I'm writing this for whoever passes by and is curious about how I think JS or whatever else should be packaged if done properly)
That said, in old Java dependency management (i.e. Maven), you could often find a source file and a docs file alongside a compiled / binary release, so that you get the choice.
But this can also be done with NPM libs already; the package.json shipped in the distribution contains metadata, including the repository URL, which can be used to get the source.
The recent HN discussion about “that one npm maintainer” confirms please hold onto the most painful ideas.
I'd like to see how well a custom dictionary trained against a few hundred npm packages could work against arbitrary extra npm packages. My hunch is that there are a lot of patterns - both in JavaScript code and in the JSON and README conventions used in those packages - that could help achieve much better compression.
We looked at Brotli as well but decompression speed at acceptable ratio was the most important factor for us, that plus the far superior docs and evangelism sealed the deal for zstd.
Don’t push devDependencies.
I honestly don’t understand why devDependencies come with the package. Most packages don’t distribute their source files, so what use are the dev dependencies?
If you want to do dev, clone the got repo and get everything. To me `npm publish` is for distributing the production library, plus source maps and types—that’s it.
NPMJS.org probably cares. And having smaller downloads for everyone else would speed things up a bit.
BTW I saw this recently about shared brotli dictionaries for delivering JS, which is nice: https://chromestatus.com/feature/5124977788977152
Yeah, who cares about 4TB/week? What is this, the '90s?!
They shrinked a package size by almost 40%. No way this isn't significant improvement. Hell, at their scale 4% improvement is big.
A 10 gigabit link can transmit 3240 Terabytes of data in one month.
Picked a popular package at random, webpack. npm says version 5.88.2 released 3 months ago has 5,992,398 downloads in the last 7 days.
I don't know how anyone can look at that see it as anything other than a massive failure.
Fast connections and free bandwidth have caused people to completely ignore the fact that every time some CI pipeline runs, npm goes off and downloads 100MB of dependencies. Dependencies that haven't changed since the pipeline last ran 30 seconds ago.
npm could fix this by aggressively rate limiting clients that have already downloaded the same package multiple times, but I guess as long as the vc funding is paying the bandwidth bill it's not a problem, and those "millions of downloads" make you look good.
The vast majority of those are from CI on ephemeral cloud instances.
Do you think CI should not be run?
Or CI should be run, but not on ephemeral cloud instances?
Or CI should be run on ephemeral cloud instances, but the packages should be cached using a separate service from npmjs.com (e.g. S3)? If so, what makes this other service preferable?
Yes, you should vendor external dependencies.
A build should ideally not require internet access to complete.
People learned nothing from leftpad.
You've got a non-internet CI with non-internet source code repository with non-internet vendored dependencies??
Technically possible, but I call BS.
Vendored dependencies are pulled down from an internal s3 bucket (and cached locally) before the build starts, the rest of the build runs with no internet access.
Look how nix does this, it's basically the same.
> Fast connections and free bandwidth have caused people to completely ignore the fact that every time some CI pipeline runs, npm goes off and downloads 100MB of dependencies. Dependencies that haven't changed since the pipeline last ran 30 seconds ago.
Maybe it's just me, but I've always thought it was well known best practice to cache your deps[0].I'm pretty certain that this can be achieved with most CI/CD tools.
https://docs.github.com/en/actions/using-workflows/caching-d...
Have you seen most CI systems in the wild? Majority of projects just wipe everything and do a fresh install, on every single run.
Maybe it will be counter-productive to shrink the packages, since it would encourage that behavior of not caring about local cache even more.
*I don't mean it's a simple endeavor but at least it's simple to describe.
tar archives don't support random access. So it would be slow to scan for each file in the archive as you load it.
From Project Management to Data Compression innovator: https://podcasts.apple.com/us/podcast/corecursive-coding-sto...