Node_modules: One character saved 50 GB of disk space
mainmatter.com
mainmatter.com
The 50GB figure is the number of Node modules a developer has installed on their local machine that are duplicated between multiple Node projects/repos checked out on that machine. Even for the notoriously bloated JS ecosystem that seems well above average, even assuming any given developer has >1 project checked out at once.
For reference, I'm a developer primarily working in JS. I use an old 2017 MB Air with the smallest disk (120GB) and have many many Node projects checked out (including random GH FOSS I've contributed to once). I don't use pnpm & I've never had disk space issues.
Don't get me wrong, pnpm is cool. I've started trying it out and will likely convert a a lot of stuff. But 50GB is extreme even for Node.
Depends entirely on how many projects you regularly work with.
pnpm is objectively better than npm in many other way too though. It does all the things that npm finally realized they needed to do in version 8 by default.
Only taking specific issue with the hyperbole hook in this article.
One? Oh lord you don't want to look at my work machine then. Probably a couple dozen projects pulled down, and once added, they aren't getting removed until I have to swap PCs.
1. I'm not an average case.
2. Even I've never hit 50GB
3. (not mentioned in my original post but...) if you did reach 50GB that's likely going to be because you've a load of old projects lying around needing deleting. Using a new tool in other to retain that dysfunction isn't exactly a great recommendation.
Pnpm is great for other reasons: not recommending against it, just calling out the ridiculous title here.
For those who haven't read it: it's about using `pnpm` instead of `npm`. Saved you time and possible frustration.
> Believe it or not, this is basically how everything in the Python and Ruby worlds work
And those 2 paragraphs are not really true. Python and ruby environments don't exist "these days" - they've been available longer than npm existed. (virtualenv 2007, npm 2010)
The system/project split exists in the same way npm --global / npm exists. The only real difference is that you can't have different versions installed in the same environment at the same time - not the other things implied by the post.
That said, I did have issues with the same paragraph, because php’s composer has been doing it the correct way since forever.
You're mixing layers. Pip uses setuptools to install packages inside a virtualenv. You need something that manages environment/dependencies and something that installs them. Sometimes they're the same thing, like with poetry or pipenv.
How do you know what to use? Read the project readme. Same as people choosing yarn or npm.
Actually no. If a repository has a package.json file I can run any package manager I want and it’ll work.
I said the same thing about some ruby gems years ago and thankfully that’s a little bit sane now.
I don’t use JS that often. But recently I looked at the dependencies for some library I was using and I was astonished at the literally hundreds of tiny modules that were being used.
And it gets even worse - those tiny little modules have their dependencies too.
Amazing.
Change one minor version number and everything breaks. Forget one dependency and npm will not tell you that a dependency is missing but instead it'll complain that Steam has a broken link in the home directory (this is a known open issue for years and the only two solutions are to uninstall Steam or to use a Docker container)
Needless to say I am not a fan of large web projects.
This is one reason I like Deno's idea of having a standard library (I just wish Ryan would have proposed that for Node directly instead of creating a brand new runtime).
My least favorite is assign.
Not only does JavaScript feature that natively (though I suppose the library may predate widespread support for Object.assign), the 1-liner assign flips the order of the parameters!
assign({ a: true }, { a: false }) -> { a: true }
Object.assign({ a: true }, { a: false }) -> { a: false }
And most of them are just straight-up pointless! Like, let's introduce a dependency for decrement lolEDIT: In looking up whether the 1-liners assign predated widespread Object.assign support, I found that their implementation - confusingly named extend[1] at first - literally used Object.assign from the very beginning. And they still chose to mess with the parameter order. For shame lol
[0] https://github.com/1-liners/1-liners
[1] https://github.com/1-liners/1-liners/blob/7c1f8d51df4b4b3e0a...
Here's a package that basically does that: https://www.npmjs.com/package/number-precision
Not entirely unreasonable as all `number`s are floats by default in JS, but the implementation of the entire package (https://github.com/nefe/number-precision/blob/master/src/ind...) is less than 100 lines of code and actually contains a method called "minus".
The very worst JS packages I've seen have got to be is-odd and is-even. 430,796 and 202,268 downloads every week, I kid you not!
You might not like it, or the library (I don't write JS on purpose, so no opinion), but it's right there in the README.
My mental model of assign was backwards, and it took me a while to comprehend why their implementation would be data-last.
So, the data is the object that is being modified? And the idea is it lets you write stuff like this more easily:
const addStuffToObject = stuff -> object -> assign({ "some_stuff": stuff }, object);
const addWeirdStuff = addStuffToObject("weird stuff");
const weirdObj = addWeirdStuff({ "foo": "bar" });If you have a function isEqualTo(a, b), you can curry(isEqualTo, 5) and filter with it.
`assign` could be used in the same way for assigning/overwriting the same field in an array of Objects, and so on.
[0] https://github.com/lookfirst/mui-rff/blob/master/package.jso...
I’d rather use node with all its ugly bits than Python or .NET or C++ with all their ugly bits. I’d rather use Rust over any of them but rarely is it the right tool for the job with what I do.
1. Looks like it attracts a lot of talent these days
2. It's basically so well designed that its smooth learning curve has enabled generations of people to develop cs skills/get jobs/build stuff independently
3. While this and the previous points are not here to state that the JS universe is perfect, I have seen far worse stuff haunting the industry in the past: Php and Java alone have produced thousands of so-called professionals who still have trouble distinguishing db from backend from frontend, and have no remote idea of what dependency management is whatsoever.
Since node_modules is mostly text, this has amazing results and can be applied to the deduplicated pnpm store as well.
As someone who's worked pretty heavily in both ecosystems – it's definitely not something I think about every day on the Python side, but Python dependency conflicts are very annoying... while in Node they're mostly not a big deal except in a small set of cases where peer dependencies show up.
Until you have to figure out that the reason something doesn’t work is that dependency v1 is storing the data that dependency v2 is trying to use, and it complains about missing data that you are sure is there.
I very much enjoy having those issues up front, instead of at runtime.
In java is basically a non existing problem, you CAN have dependency conflicts yes, nontheless dependency management is simple, and you keep everithing on a local central repo when using maven, which also provides a very nice dependency tree plus tools for filtering, whicg are nice, which of course you can also achieve with grep for even easier dependency conflict debugging.
Also using tools as dependencyManagement in maven allows you to replace all usages of a library across your entire application "at your own risk" which simplifies addressing security vulnerabilities
Our workaround is to put the global cache in the project so we can mount it alongside the code in the same mountpoint.
I would argue building different containers from a shared node_modules is inherently dangerous anyway. Sounds like your ”workaround” is in fact pretty much the optimal setup for quickly performing multiple similar builds.
Do you mean hardlinks?
That said, soft links should work just as well.
I have no clue why they are saying that they are hard linking directories, I sure hope they are only doing it on file level.
In most programming languages I've used on Linux, symlink expansion is on by default, creating all the problems you can think of.
Yes, you can use whatever API call readlink relies on to prevent loops, but you can also keep a set of inode numbers on a file system and stop processing on duplicates. In both cases you need to do all kinds of workarounds.
The best argument I've heard is that hard linked directories break the acyclic graph property of the file system but even there I'm not so sure if that's really a problem in an environment where most tools recurse into soft links anyway.
I suppose the C folks that like to do all the hard things themselves would get annoyed by having to add another check?
I don't think the kernel forbids directory hard links per se; NTFS has junction points which are somewhere between hard links and soft links for directories, and NTFS support made it into Linux. I don't know how the kernel exposes directory junctions in Linux, but they're different from soft links (in that they'll be processed in the server for SMB file servers, as opposed to symlinks which are resolved on the client).
This is where a poor choice was made and everything went wrong.