NPMprune: Remove unnecessary files from node_modules to optimize storage
github.com
github.com
- Some packages contain non-JS files for good reasons, and they may break in subtle unpredictable ways when you mess with the contents of their package.
- Node.js will happily run JavaScript files even if they're not "*.js": A file like "hello.alsdfhlshdfl" works just fine as long as its content parses. There is no guarantee that your dependencies (and their recursive dependencies) don't statically or dynamically load files with completely arbitrary filenames.
- If you distribute packages with license files stripped this way, you are violating licenses that require the license to be distributed along with the code.
If this is actually a major issue for you, consider instead sending PRs to upstream to tidy up their package. This will also benefit other users.
- The patterns used to find files are specific enough to target only those files that are well known to be useless at runtime.
- The license texts of these libraries can be copied and merged into a main LICENSE file.
- Have you seen the number of modules installed by most major libraries? Making a pull request for each of them is humanly impossible and counter-productive. It's easier to use a simple script that releases dozens of MB in a few seconds.
Traditionally, this wasn't an acceptable way to think about projects we engineers were being paid lots of money to build.
As you note, a project may hoover in some absurd number of dependent libraries and you have no tooling that tells you which of those might fall in the 1% and what code paths in those 1% intersect with call stacks in your project. You have no idea what impact blindly deleting some "They're probably unnecessary" files in somebody else's code will have on your application and no insight into how to make sure your testing unearths problems. It's an invitation to phantom bugs of unknown scope and the most frustrating kind of debugging effort that comes from chasing those kinds of phantoms.
It's already bad enough that people don't read and review their dependent code with the eye they bring to PR's from their on-team colleagues, but to then go futzing around and deleting things in the unread dependencies because you have a hunch that it's no big deal is about as far from software engineering as you can get.
Right so most projects end up with 100's (random one I have is 700+) modules. Which would mean multiple breakages.
The worst part isn't the breakage - it's not knowing where or when it breaks, and because it could be missed when it's being bundled it can happen in production.
The bundling step should effectively be doing the file pruning for you (or even parts of files) and you can be a lot more confident that won't miss things.
node_modules are generally big (580MB in my case), but I don't know why you'd trade 580MB of storage for reliability. For us the 580MB will get bundled under 1MB for our web application, essentially all dev machines will be 512GB+ at this point anyway.
> Have you seen the number of modules installed by most major libraries?
Chances of success, negligible.
Translation: take 99% to the power of 'a lot' and what do you get?
So for the typical enterprise crapware where the app template installs about 2,000 packages for a React Hello World, how many broken modules is that?
NPMs package.json has a `files` field which allows you to define which files are included on an npm install: https://docs.npmjs.com/cli/v6/configuring-npm/package-json#f....
This also extends to an .npmignore file that works similar to a .gitignore file.
Compare https://github.com/express-rate-limit/express-rate-limit/blo... to https://www.npmjs.com/package/express-rate-limit?activeTab=c...
Agree with you about the other points.
Having all of this stuff makes it possible to ctrl+click on functions in my libraries and read the corresponding source code. That’s a godsend during development - well worth a few extra kb of files in the npm module.
tsconfig.json:
"declaration": true,
"declarationMap": true,
"sourceMap": true,
...
package.json (assuming typescript compiles src/ to dist/): "files": [
"dist/*",
"src/*"
],Be super careful of removing large swaths of files. Out of 150,000 node modules in your manifest, I'm willing to bet at least one of them is doing something by reading one of these non-source files.
Looking at the script source, it's just matching globs, so there isn't much smarts to this. I'm sure it works most of the time, but yeah..
Do JS packages need some kind of .prodignore file similar to other .ignore files?
So with a flag passed, after doing an npm install, there's a extra cleanup step that removes explicitly marked files that aren't needed for running in prod?
(Not a fully formed idea, I'm sure I'm not thinking of drawbacks with this)
Edit: this sort of exists as the .npmignore file?
https://docs.npmjs.com/cli/v10/using-npm/developers#keeping-...
I did notice small issues with some libraries (react testing library IIRC)
Was introduced at work and it's a game changer. The monorepo support (via "workspace:*") is absolutely clutch too.
As my ol' grandpappy used to say, "Why wonder? Let's go search the Internets!"
Npm already does it at the package registry with ignore/npmignore files, and that's the package authors choice. How much storage can you really save? 50MB? 200MB? is it really worth the risk of running rm on some glob pattern and cross your fingers the packages don't require any of the deleted files?
I tested it recently on a clean install of Strapi: about 250 MB are freed up. Storage is cheap but that still represents a lot, especially inside a Docker image.
The patterns used to find files are specific enough to target only those files that are well known to be useless at runtime.
You are taking a big risk of subtle breakage right now, and a big risk of breakage as you change your project code in the future, as you may start to invoke a code path that needs that resource in the future.
>
> wget -qO- https://raw.githubusercontent.com/xthezealot/npmprune/master... | sh -- -p
Serious question: Is this the norm now? Are people actually executing unversioned wget'd shell scripts from random github users as part of their deployment workflow?
I could do some tricks where I sent different files based on user agent, but still... most people aren't inspecting the download anyway before running it.
Not from githubusercontent you couldn't. Which I'd say is where the majority of these scripts are hosted.
https://www.idontplaydarts.com/2016/04/detecting-curl-pipe-b...
I'm on mobile and this refuses to load due to SSL/HSTS problems. It's an interesting approach if you can get it to load.
Point being... someone clever could make a bad day for the minority of users who do check first (but don't save/run what they explicitly checked)
Running curl, saying "yup that looks good", then adding a pipe achieves little with actual malice. The payload can be completely hidden until it's too late
The phrasing you use is important for any hope at safety - review/run the same downloaded thing.
For about the last 15 years
It looks like that practice goes back at least 40 years.
curl ... | sudo bash
Useless files will still be there.
Also, when you create a Docker image, you avoid packing in dev tools that aren't absolutely essential (such as Yarn).
I used to use it, but at some point the hassle of the occasionally breaking package wasn’t worth it.
Are people that concerned about the size of a directory on their machines?
Try this bash one-liner :)
find / -name node_modules -print0 | xargs -0 rm -rfAlso, when you create a Docker image, you avoid packing in dev tools that aren't absolutely essential (such as pnpm).