We check our node_modules folder into source control
jackfranklin.co.uk
jackfranklin.co.uk
To respond the two major criticisms:
1) “It takes a lot of space”
Don’t be so sure. Text diffs and compresses well. I have a 9-year old Node repo that I’ve been vendoring from the beginning and it’s only grown 200MB over that time. (Granted, I’m fairly restrained in my use of dependencies. But I do update them regularly.)
But even if it does take a lot of space… so what? If your dependencies are genuinely so huge that this is a problem, then vendoring may not be right for you. But you could also use one of the many techniques for managing the size of your repo. Or just acknowledge that practices are contextual, and there’s no such thing as “best practice”—just a bunch of trade-offs.
2) “It doesn’t work well with platform-specific code”
This can cause some pain if you’re in a multi-platform environment. The way I deal with it (in Node) is by installing modules with --ignore-scripts, comitting the files, running “npm rebuild”, and then adding whatever shows up to .gitignore. I have a little shell script that makes this easier.
This is only an issue for modules that have a platform-specific build, which I try to avoid anyway. But when it comes up, it can be a pain in the butt. I find its pain to be less frequent and more predictable than the pain that comes from not vendoring modules, though, so I put up with it.
Bonus) “It’s not best practice”
Sez who? Dogma is for juniors. “Best practices” are all situational, and the only way to know if a practice is a good idea is to examine its tradeoffs in the context of your situation.
If consistent build inputs is your concern may I ask why using lock files wasn’t enough? That’s a problem they were designed to solve.
(Perhaps they've solved this in the last couple of years. I've been staying away from that ecosystem... too much growing in it...)
On reflection, it seems more like I'm papering over network issues. Perks of working in an enterprise company I guess.
In a discussion about Skub, no one need to explain why the Skub-powered approach isn't good enough. It is the duty of anyone pushing Skub to explain exactly what makes Skub so special to the point that we need to have Skub in our lives.
It wouldn't be necessary if the dependency tree was a pure function of package manifest.
https://developer.okta.com/blog/2019/12/16/semantic-versioni...
I don't check in my dependencies in my current project because I don't need to; but in earlier projects, I or we did, for various good reasons; and it worked perfectly well, and was extremely convenient for new developers.
Best guess is something like “ship required source or binaries along with your end product.” Like static linking but extended to dynamic languages and source control.
- "npm add package" puts in a "^ver", which is bad practice
- there is no good infrastructure to pull hash based blobs out of the ether in case npmjs is offline
npm-shrinkwrap has solved repeatability forever, people just didn't always use it. Auto-upgrading dependancies is the big problem, which should have never existed because it is not principled. I'd go further and say that dependancies and devDependances should only support exact versions, and peerDepenancies are the only thing that supports non-exact versions.
For your average npm shop which doesn't have infinite internet oil money, here is why the article recommendations won't work for you.
Your CI will pay the time penalty during git clone instead of npm ci. In fact, the node_modules folder will be bigger than your source folder almost immediately. And over time you won't be cloning just the head files you'll also be cloning every npm package binary ever committed. You can't undo this without investing in smarter git tooling. Which is time spent not writing features.
NPM packages which install arch-specific binaries will constantly flip flop from commits by devs on different OS's.
Nobody is safe from left-pad, not even Google, and committing your node_modules folder doesn't change that. Eventually someone is going to have to run npm i.
Running npm ci on everyone's machine is reproduceable, I don't know what OP is warning about. Package lock pins all the versions.
If you have a large enough team to invest in dev experience, there's way better ways to get the advantages of the article without the downsides. You can cache the npm ci result in a container layer for your CI/CD or use middleware like artifactory
Things like GitLab's CI runners will do a single clone then do `fetch`, `checkout`, and `clean` to checkout your repo. Git repo size isn't a huge bottleneck in CI performance.
I'll admit I don't believe everyone always need to check every dep, but we're skating close nobody checking them ever.
Once something's in package.json I don't believe anyone who says they can vouch for the security of that over time. We're all doing security theater with npm audit, dependabot, etc. Don't use npm at all if anyone's life depends on your code.
#1 You look at the dependency and do an assessment on whether it's worth including. Check.
#2 You probably require some automated checks. SAST, Depedency Scanning / SCA, maybe some DAST, etc. Check
The outstanding question though...
#1 Did anyone actually read the code of the depedency?
#2 Did anyone actually look at what the depedency itself pulls in?
#3 Are these checks re-done when you update the lock files?
#4 If nobody is doing it, who's updating the lists and rules we use to scan from?
#5 Where possible do you have the monitoring to check when an app is doing something weird? i.e. network ACLs that when they fail, cause an event, that alerts a person to investigate?
I think we're mostly agreeing here, but the wider question is why is it that folks writing the app and including the depdency don't feel responsible for these things?
On an Ubuntu x86_64 machine with an SSD, require("react") with Yarn PnP ends up 4x slower than when installed from npm (Node 14.17.6).
Yarn doesn't do any kind of compilation itself - that would be done by build tools like webpack.
https://yarnpkg.com/features/zero-installs
It’s also compares this with checking node_modules into git
If offline installs is what you want, I don’t see any advantage of node_modules compared to this feature - only disadvantages (size, noise, and cross-platform incompatibilities).
But there is nuance (there always is...), the README file in node_modules is here: https://github.com/ChromeDevTools/devtools-frontend/blob/mai... - and it makes it clear the only NPM dependencies used by the build-system or infrastructure is meant to be checked-in. Other NPM packages should not.
----------
In conclusion: the linked blog-article is clickbait that misrepresents how the Chrome team manages their dependencies.
Here are all "node_modules" folders I can find (sorry I'm not much familiar with npm/js to know more) https://source.chromium.org/search?q=file:node_modules%2F$%2...
In our industry (games) we often do that - checkin prebuilt code in the depots (typically "p4"). I'm not saying it's wrong/right, it's just what we do (not 100% fully, but almost, though people in IT/infra tend to do otherwise).
Does your team/company store art and asset blobs (even rendered FMVs?) in Perforce too, or do you have a separate asset-management system for that? If you do have two separate systems, how do you keep the asset-store and source-control in-sync?
It's been a long time (easily 20 years now) since I dabbled in any high-end creative software (like 3ds, etc), but I remember they generally all had large binary singleton project files (like Flash .fla and Photoshop .psd files) that couldn't be deconstructed and effectively diffed by any source-control system (though Flash eventually supported external ActionScript source-files), I'm curious how that affects your org's asset storage needs.
Locking of files too, but that’s only necessary on non-diffable assets.
Forgive the questions, I'm just curious about the minutia and peculiarities of the gaming-biz because I've never worked in the field.
Most game engines natively integrate with perforce so it's largely invisible.
Those that don't use the engine typically use an IDE which automatically checks out files; then they just check them in using the p4 GUI.
Even I don't use the commandline and I'm a CLI-Junkie.
It maybe wasteful, but it's the established practice (it seems). People coming from other game companies pretty much use it too (some rare exceptions). Also the automotive/chip design industry uses them - and yes mainly for the big blobby non-diffable/mergeable assets.
Locks are terrible, and yet you gotta doit sometimes, as how else would you prevent people working on the same asset.
I'm (still) terrible with git, often screw up commits, and have to google search/stack overflow to get it right (I use it mainly for simple home projects). I can only imagine the pain and suffering a non-tech person would have with git. Also the metadata is quite lot for WFH conditions. Working remotely does not always mean working from a dumb terminal (I wish).
I'm unfamiliar with P4's locking semantics; when files are "locked" does that prohibit other users from even getting a copy of the centralized file, or merely prevent users from overwriting the centralized file on push/upload? How does branching work?
If I were designing a centralized asset management system then I'd definitely add support for git-style (i.e. "many-worlds") branches instead of SVN/TFS-style "spatial" branches - but I'd also add support for some kind of "mini-branch" or deferred-conflict-resolution, whereby a file, or entire directory, can still be pushed to central storage but have multiple different representations that can be resolved/merged later, rather than immediately. So if two artists are working on the same "texture123.psd" file without realizing it then the system would let them both push (so the first artist to push would get to overwrite the file, and the second artist's push would see their file saved as "texture123.psd.v2").
There are good business reasons for having the ability to disallow changes to files in central storage, but that doesn't mean locks need to be used: it could be done by instead directing all updates/pushes to separate mini-branches, thus allowing users to push-and-forget and allowing them to defer conflict resolution while still protecting files from unwanted changes.
Some modules straight up download binaries so I don’t see being so straightforward.
I've seen this trip up people in the past, in one case a CI/CD system running Linux was used to produce a project deployed to a mostly windows environment this didn't cause any issues until the day a developer added a binary module. Honestly, I'm surprised it worked as long as it did, it took a little over a year before anyone hit that issue.
It could be if it's a CPU architecture difference. For example an M1 Mac (ARM64) vs just about every other system (x86-64).
I know we had to switch out MySQL with MariaDB locally because the official MySQL Docker image doesn't support ARM64 devices but MariaDB does. That's just another example where even if you're using Docker there could be differences.
We've also had issues where developers aren't used to case sensitivity at the file system level and things work on their Mac but fail on Linux in CI because Docker's bind mounts (often used in dev) will use file system properties from the host OS which means even if your app runs in Linux within a container it may run differently on a macOS host vs Linux.
The moral of the story here is Docker is good but it isn't a 100% fool proof abstraction that spans across Linux, Windows and macOS on every combination of hardware.
Couldn't this issue be solved by using caching? If I remember correctly, Travis CI has the option to cache certain folders between builds, meaning an npm install doesn't have to start from scratch, and can just incrementally update the cached node_modules folder (any changes are then copied to the next CI build).
One example was having code where the test suite ran fine on my local MacBook, but would fail on CI (Linux). It turned out that on Linux finding files by name is case sensitive, whereas on Mac OS it isn't, and a require statement in Node.js was referencing a file path with casing that was different to the file name's spelling.
But then again, I know there's a bunch of Nixers that pick their head up slightly, shake it, and then just go back to work on actually interesting problems instead of a millionth discussion about dealing with npm. Christ, the stuff people put up with.
EDIT: Ironic, this being here along with the CISA/log4j post where everyone is yammering about SBOM (software bill of matearials). Again, I just glance over at Nix and go, "sure, what do you want to know, I can tell you instantly if log4j is anywhere and if it's a vulnerable version (excluding non-source-built packages in nixpkgs)".
What? What if the dependency upgrade requires code changes? Master is just broker until the second MR merges?
This all sounds like terrible advice.
Yes, but a PR commit is not in master.
> There are times where this doesn't work; updating TypeScript may require us to update some code to fix errors that the new version of TypeScript is now detecting. In that case we have the ability to override the rule. As with anything in software engineering, most "rules" are guidelines, and we're able to side-step them when required.
Another idea is to host your own package cache. That would be my preference where SCM size prohibits checking in dependencies themselves.
This makes most sense, and is the common solution at larger companies. Right tool for the job and all that.
So, for example, you install the dependencies with Node 12, but have to run locally the system with Node 14.
So, i think just checkout the node_modules folder into git is not complete solution. You're avoiding the need to commit inconsistent package-lock.json, which is not hard to solve though.
Also, according to the official npm documentation, npm should be installed through nvm, thus having the ability to specify its version separately from node.
https://docs.npmjs.com/downloading-and-installing-node-js-an...
I think, instead, it would be good to move in a different direction for sub-dependencies generally. A rough sketch:
1. Packages state their dependencies as they do currently.
2. When released, they’re built and bundled by the package manager host.
3. Included in the bundle is a manifest of the dependencies used, specific imports used, and a hash for each (recursively until exhausted).
4. On install, identical code (same hash/same bundled result) is deduplicated.
5. A human readable record of dependencies and imports used is produced. It’s important that it’s human readable, because:
6. This should not be filtered out in diffs. It should be subject to review just like any other change.
All of this is pretty complex, and there are probably ways to reduce that complexity. But it has some obvious advantages:
- Only your direct dependencies are installed. Tons of bloat can be stripped out.
- Even deduplication can be performed on the package manager’s servers. And hashes aren’t a particularly expensive lookup.
- It would go a long way towards addressing audit fatigue: if your dependencies’ bundles don’t include affected code, the audit doesn’t apply; if they do, you can be reasonably confident the audit is valid.
- A (wild guess) huge amount of the time, sub-dependency changes will require little to no review. Their stable parts will seldom change, and the parts shared among several dependencies could be reviewed as one unit.
- Lock files themselves just need to track direct dependencies (and even then, only to support semver ranges).
Ideally the onus is on the package manager to provide metadata in the lockfile for the strategies it took when generating.
something like cargo-crev for npm might be a long term solution
The fact that you have so many people believing "nuke node_modules and delete package-lock.json" is a reasonable step in diagnosing an error is damning to NPM.
We don't check in our node_modules, but "use the lockfile" is not a valid counter to this article's points.
Lockfile v1 literally ignores pinned versions of dependencies if the package.json specifies a fuzzy version number[0], and the advice of the npm team was, "it's fine, everyone will just bump a major version number of npm." And to this day, I still don't know what the expected behavior is, there really isn't a list anywhere about when the lockfile is and isn't supposed to be respected. So it's not really surprising to me that people distrust version pinning, and I always feel like I'm kind of living in a different world when people say that lockfiles just solve everything.
[0]: https://github.com/npm/cli/issues/564#issuecomment-921314014
but those who care use yarn, those who even want to be correct use yarn2, and so on.
A Git bot that parses changes to yarn lock and comments with size/loc/files of all the deps etc, kinda like how coverage bots work. get all the observability benefits without the checkin cost, also wouldn’t have to split npm changed and code changes..
.. on and on, 40MB..30MB..20MB.. hundreds of them just adds up! 147 dependencies over 1M. pretty incredible/wasteful really when you look at it.
What I don't understand is why dependency download isn't a separate task you do before compilation. FIRST you download all the dependencies to stabilize that, then you build the code. The reproducibility is the big one.
IT drives me nuts that the dependencies are hidden away in maven and gradle. I have to lookup obscure "download dependencies to a lib" task configuration. Still the obsession with massive jars/wars/whatever when all you should have to update in a deploy is the difference in the libs and the main code jar.
The reason for this obscurement is pretty much "uh, it saves disk space?" which is a laughable consideration given the bloat in war files, docker images, and the like.
I agree with the poster.
This is what bazel does. It also offers a `bazel fetch` to pre-download things before going offline (ex: flight).
https://github.com/qaware/go-offline-maven-plugin
It's not perfect, but it can be useful.
I really would have hoped that the NaN module related problems will be fixed over time, but here we are in 2021 and nothing's been fixed.
As long as npm doesn't use binaries and headers, those things will stay broken. The thing that they argue with to use "always source" is kinda ridiculous when considering that probably the most of all npm packages are using webpack or another bundler before pushing their own package to npm - because npm itself has become impossible to use as a package manager alone.
I mean, a couple MB of libraries with the wrong dependencies can lead to multiple phantomjs installation, which is an inactive, deprecated, and unsecure project for years already... just because of some unit tests that have no place in a production npm package.
My hopes are that more sane developers come together, switch to ESM and implement better policies for evaluating their dependencies (e.g. blocking sources from people that have more than 1000 npm packages and brag about it).
Pikapkg was a great idea in my opinion, and I was using it before they moved the project to building astro as a platform :-/
It's a sentiment that I'm actually in agreement with. I've been coding mostly in Java for the past 22 years. Somewhere around 2010 Maven became the prevalent build tool quickly displacing the venerable Ant. With Ant we had builds that used checked in jar file dependencies. It was obvious what your builds consisted of and they were very fast once you cloned the repo.
Then came Maven and the conflict resolution hell quickly followed (esp with unpinned dependencies). Now every time I type mvn install it feels like a small adventure in its own right. I'm absolutely flabbergasted. We sacrificed simplicity and reliability for a bit of instant gratification. Bad tradeoff.
But I share the overall sentiment. What if Maven Central goes down one day? It’s just a web server like any other.
Most Java shops run their own artefact repository acting as a pull-through repository to Maven Central and other 3rd party repositories. In addition to acting as insurance against losing access to critical dependencies, it also can provide a performance boost when downloading dependencies, and helps to offload your organisation's traffic from Maven Central.
With frontend nowadays it’s sad but after six months it’s highly unlikely that my project will compile, not to mention the tooling like Vue, Vite, etc that has breaking changes.
I mean it’s scary. You write a program and it WONT run if you just give it enough time. Locking versions is not really a solution since often times the tooling itself, and IDE extensions require newer versions of packages.
Maybe you wanted to fix a typo a year down the line but oh no, now you need to figure out why Vite won’t start, why eslint dropped support for xyz, spend hpurs figuring out what you need to change in your configs, etc.
But, there is one thing I like from this, which is git diffs showing the actual final code diff when you upgrade dependencies.
Of course, this being horrible advice, it ignores how many JS packages ship minified which would make the diff as useful as binary noise. But I like it in theory. This could be a good opportunity to write a tool that replicates this specifically (and what's more, for other languages as well)
The company I worked for, used to do that with everything. In fact, one of the ways that they archived versions, was to create a bootable external hard disk clone, of the entire development machine, and store that.
If you want to be absolutely sure that you have the complete building blocks, then you don't trust your package manager. Make local dupes of the packages, and integrate them into your own version control.
* https://developer.chrome.com/docs/devtools/
* It's mirrored on GitHub: https://github.com/ChromeDevTools/devtools-frontend
* And look, THAR SHE BLOWS: https://github.com/ChromeDevTools/devtools-frontend/tree/mai...
And composer patches makes life quite easy compared to maintaining a fork. If I were to fork something I would need to handle merging every time they have a new release, run the build etc. With composer patch, a new released version is installed and the patch on top. Sure, if there's a conflict that needs to manually resolved but that's usually minimal effort since most patches are absolutely tiny, a few kilobytes at most.
I never even understood the arguments for keeping the packages out of git. Trying to save disk space these days is pointless. Maybe npm is different but composer handles about 160MB of code here. Maybe I missed the memo but these days that's nothing. My laptop shipped with a 500 000MB SSD so it's like, what, half a percent? The speed advantage , on the other hand, is absolutely undeniable, git won on speed in the first place, these script language tools can't possibly compete with a git pull on speed. git diff, as the author notes, is not at all a problem, just separate the vendor commits from your commits. And as I noted: they are useful for vendor packages.
Any reasonable CI tool will have a way to cache generated assets based on file contents, that's the way to go here IMO.
That is what happens in a DVCS, after all—in fact, it's sort of the whole point. If you're so uncomfortable with this, it might be worth asking yourself whether it was ever really the case that you agreed that DVCSes were the right approach. (Even then, still no reason to embrace package managers like some kind of paramilitary force that's subject to its own rules—better to just improve your version control system to handle things the right way, right?)
Sure you can say "well why not just make git better then it will handle any operation in any size repo imperceptibly quickly", but I think you and I both know that isn't anywhere near as easy to implement as it is to type.
Good thing I didn't type that, then. That's not what the shape of my argument looks like at all.
There's a massive leap between, "we don't want every clone of our repo per se to carry all the baggage of our dependencies' histories, so it would by nice to have some scheme for handling lightweight, shallow copies" and "... so we decided the right way to do that, rather than making that a first class feature of the version control system we're using, is to create a hack in the form of a new set of unrelated tools meant to circumvent our VCS's fundamentals completely—so from our its point of view, these controlled objects and the scheme we use for managing them are invisible and might as well not even exist."
Horrible advice. Don't break the industry practice and check-in your node_modules
For me, five years after you left the company it’s a pain the arse because all your code refers to this repository that doesn’t exist, the one with the custom packages with no source control, the dependencies which are no longer even in LTS versions of any extant OS distribution, and the Vagrantfile won’t work because it used undocumented perimeters for both Vagrant and that homebuilt hyper visor that you and your team built as a lark (that doesn’t exist outside your personal laptop).
So for me, it’s all the dependencies get checked into the repository, all the tests run before we merge to master, and we do not use any custom in-house infrastructure of any kind.
This way you can build if those services go down. Also, performance.
https://github.com/unqueued/git-cache-tag
It copies all untracked stuff (including node_modules) into a leaf tag. It is fairly easy to manage them, or find the latest one. And because they are leaves, they can be pruned and completely garbage collected when they aren't useful anymore.
I have been burnt many times by npm, and I use this script to guarantee that I have a stash of my node_modules, while also keeping my project small.
And I have diffed different snapshot tags to see which module changed that broke something.
And by leaving everything in unaltered text, it exposes it to git which does a great job at compression stuff, especially highly differential revisions of my node_modules.
A 500M node_modules from one of my projects only weighed about 100M extra, even with several snapshots. And I can just delete them anyway.
I need to work on it a lot more, it was just a quick and dirty solution when I had to work with React Native a few years ago. It doesn't handle submodules at all, and there's plenty more I'd like to do with it.
Depending on the context, if you don't want this in git history, and want to handle git submodules, there's also git-archive-all https://github.com/roehling/git-archive-all (if you like shell scripts, it is using bats for testing - it was the first time I heard of it)
Also, you clearly have not written software using Node.js on a long enough time horizon. Pinned versions don't mean anything when sub-dependencies can have transient versioning resolution occur.
The reality is that unless you can fully byte-for-byte assure what you have deployed today is what you can retrieve from an old tag, let's say weeks, months, or years from now, you don't have a replicable build.
Most people will never need to do this, sure, but serious operations who will choose Node.js to build some software and then plan to walk away from it later should not only commit their node_modules directory, but also keep a copy of the designated Node.js engine version as well. It's not likely you'll need a backup of an LTS version when you can just go retrieve it, but that's not the point:
You will encounter a scenario in your professional career where retrieval is not an option for some piece of software.
Edit: There are industries where committing prebuilts is normal and has absolute strengths, and having experienced it myself, it certainly is desirable sometimes.
This comes across as very patronising. I would assume the author is perfectly aware of what they just wrote.
I live in the five years future where the propeller head rock star programmer has moved to greener pastures (the ones where he doesn’t have to write project planning documents of any kind).
They are different worlds.
Don't let it get to you, I'll happily cash checks to work on whatever legacy spaghetti tech is in play. Hours are hours, dollars are dollars, and as long as I'm maintaining a happy ratio of those two, I don't mind what code I'm working on. I'll sharpen pencils and sweep floors for 30 hours a week, I'll even listen to your life's struggles if that's what you want me to do for that money.
Don't let me give the impression that this is because all I think about is money or that I don't care, it's quite the contrary.
I think of myself as a developer, I like to think I do a good job, when I get the opportunity to straighten the edges of a sagging beam or create a structure from scratch I take pride in that, as it is my purpose. There's no point in getting worked up about the practices of my peers because that does not serve me. This is my craft, and the person who's creations I am now steward of was also a craftsman, who had different experiences, motives, and contexts that led them to expressing their intent through the code now entrusted to me.
Imagine a television show, or movie franchise; it may have different writers over time commanding the dialogue of the main character, developing their mannerisms, polishing their pearl. These businesses and legacy products we work on are just like those characters who get passed on to new writers. Think of yourself as one of that team, carrying on a legacy, adding your flair and support. Never stop working on your pearl and use every project as an opportunity to fulfill your own desires and express yourself, while honing your skill.
I'd love to join the smoking-jacket crowd but I have student debts to pay, and my other job as a janitor actually sweeping floors to get to.
The world isn't the same place it was when punched cards, smoking jackets and bourbon in the library were a thing.
That's the purpose of lockfiles-- to pin the entire dependency tree.
If the package.json listed a fuzzy dependency and the lockfile was pinned to an outdated version, it would just be updated anyway. This was fixed in later versions with the release of the lockfile v2 format, but the fix was never backported to older versions of npm, even though those versions of npm were the recommended, default versions that shipped with LTS Node installs if you went to the main website or installed from a software repo.
I think that for a non-trivial number of people, they may not have a lot of trust for lockfiles because they tried using them and they just straight-up didn't work.
But "not best practice", qed.
Are they, really?
The less debatable point is arguing that making CICD pipelines slightly faster, but this feels like an appeal to microoptimization. Any free tier CICD system out there let's you do a single npm install and move these dependencies as far as you'd like into the pipeline as artifacts. Is a git checkout really faster that a npm install?
For example, is adding a npm dependency really invisible if you already track package.json and even package-lock.json? Those files show up in diffs, and it's hard to miss them.
Also, if the goal is to get replicated builds, isn't this handled by pinning versions and tracking package-lock?
The left_pad example is particularly ridiculous as I highly doubt that a company like Google, like any company that cares about auditing and vending dependencies, does not run its own npm proxy with cherry-picked packages.
One wonders why they aren’t checking in the binaries for their database and language runtimes. Surely this would save crucial seconds in project setup.
Dumb false equivalence, since that ("every system library and utility it depends on") is not the argument of the side you're trying to appear to offer a response to. Please refer to the HN guidelines.
A less dishonest retort would be to ask if one should check in the dependencies that are analogous to what ends up in node_modules, and the response would be, "welp, that's exactly how many app developers have been known to approach things, so 'yes'."
I don't recognize your claim that the distinction is arbitrary. Is the NPM world's distinction between package.json's "dependencies" vs "devDependencies" arbitrary? (Answer: no.)
Not only is that possible, that's also expected to be mandatory in any company that is required to monitor ad control dependencies. I know for a fact that some FANGs do manage and enforce the use of internal npm repositories, mainly because of infosec audits, and I doubt Google is not one of them.
It's already been long enough since that time that people seem to have forgotten that it even existed. Vendoring dependencies is one of those things where every once and a while a language will reintroduce the concept, and it always seems to catch people off guard. Go is a good example, although it seems to have varying advice about whether vendored dependencies should be checked in to version control. That might not be surprising considering that Go is also coming out of Google, just like this article.
It feels a little bit weird to say that this is just "industry practice" when you have the Chrome DevTools team telling you they don't do it, but :shrug:. Google does tend to be a bit of a rarity in how it treats monorepos. I'm just always interested to see how opinions on this have evolved; it's rare for me to see analysis that says, "we used to do this, and here's why we found out that it didn't work." Usually the opinions end up seeming more universalist, like the very idea of vendoring dependencies is somehow weird and unexpected, and not something that the industry was largely on board with for a decent amount of time.
How would you go about the left pad issue?
That's really not a lot for knowing the code that is being used in your codebase hasn't changed, and the bonus of having everything available should the registry go down, or something.
So I don't think it's 'horrible advice' - it's do what suits your needs best. Some people want to have everything they need to build their application in their control, on the off chance everything hits the fan.
(Also, this: "Don't break the industry practice and check-in your node_modules" - does not necessarily mean it is the best way, it just happens to be the advice from the start.
Oddly not mentioned in the article is that it allows you to build when npmjs.org is down or unreachable, which happens often enough to be frustrating, and if it happens when you're trying to deal with an emergency, it's downright infuriating.
with NPM shrinkwrap you lock down the version of the package you installed and their deps. That way you can use the same package on all envs. Helps with testing and debugging as you're removing a variable (bad deps, outdated deps, newer deps etc)
Just use locking, cache and reduce your dependencies to bare minimum you actually need.
In JS you can often find yourself pulling a massive dependency tree just for simple usecase that would be replicable in your codebase under 3 hours.
Git will pack objects after a while (initially they’re stored as-is, just compressed), but large files or enormous amounts of changes can make the repository grow to unwieldy amounts, and while git is able to perform “shallow” clones (clones which only store part of the history), not everything handles them well
This article has excellent diagrams that depict the different types of clones:
https://github.blog/2020-12-21-get-up-to-speed-with-partial-...
> Having your node_modules checked in guarantees that two developers running the code are running the exact same code with the exact same set of dependencies.
No, it doesn't, your system environment is important. Any code executed can check for environment parameters and branch accordingly. For example, a simple "if (macOSVersion === "10.10") {} else { }" would run different code branches and possibly produce different results, even when executing the same code binaries.
The reason you don't check in node_modules is that differences in system env during build time produces different build results - checking in node_modules fixes that, but does not handle system differences at runtime.
Is this really true for packages with pre/post install scripts?
It says, quite succinctly:
> I currently work at Google on the Chrome DevTools team and we check our node_modules folder into source control
Can't you ping him internally (don't you guys still use WebEx? lol) to ask him to clarify, and find out how it works for their team?
Aside: They seemed to like their Linux containers in 2014[0]. No idea if they're still using the same technique, or have maybe moved onto Kubernetes?
[0]: https://www.theregister.com/2014/05/23/google_containerizati...
One of the lead devs - years of tenure - on one of the Android backend teams had barely heard of the name and wasn't sure what it was when I brought it up 2 years ago.
Linux containers including Docker are absolutely widespread as I understand it, though.
Anyway, it's a given that I disagree with the article's specific point (i.e. to commit node_modules to source-control), however I am sympathetic to arguments about avoiding another left_pad incident, but there are better solutions to that then simplistically committing node_modules:
1. package-lock (though this is an incomplete solution: it helps to protects you from vague dependency version numbers (as it uses cryptographic hashes), but it doesn't store a copy of the npm package, and you need to make sure everyone is using the exact same npm/node tooling versions otherwise your package-lock file will be clobbered by different users checking-in wildly different `lockfileVersion` versions.
2. git LFS: every so often (once a month or so?) in a separate directory off-to-the-side, add a heavily compressed 7z LZMA archive of a snapshot of your node_modules directory (ideally in a known-good-state). This allows you to keep a repo-local copy of your important dependencies without it cluttering up your commits. While these would be monthly updates - and your actual package.json/package-lock.json dependencies may change daily or weekly - in the event of catastrophe it won't be too much work to track-down any missing dependencies or to revert the deps back to the last known-good LFS file.
3. Use tools like `offline-npm` and Verdaccio, which are NPM caching proxies. If this was 2019 and everyone was working in a central office then you'd run Verdaccio on a single box in your LAN and have everyone configure their NPM clients to route through that box, which then stores every package ever requested - you could presumably run a cron-job to ensure that package cache is backed-up somewhere safe, maybe even with git-LFS as discussed above.
I've only ever had issues with differing major versions. That is, sharing lock-files between any node v14.x should work, but expect things to break if you go to v16.
Or is there actually a fundamental design problem with lock files in general? My sinking feeling is that it's rather just the Node ecosystem's implementation of them :/
Then somewhere down the line...
> One of our dependencies that we check in is TypeScript, and every time we update that, the git diff is huge and frankly not worth looking at (beyond the CHANGELOG)
Which I assume is the official attitude towards any dependency of the same magnitude. How are you more aware of the code you are shipping? Okay you managed to give yourself a visual on how much LOC your dependencies are but is that a relevant awareness? Do I not get the same thing with a `du -h node_modules`, with a matching pretty GUI on top?
The one thing I haven't seen addressed so far is: doesn't this make them susceptible to poisoned dependencies? Say they have a dependency to a large well-known library, what's stopping a malicious contributor from adding an HTTP call, thinly disguised to prevent grep, to some server in MiddleOfFucking, Nowhere? Even if they manage to flag this from blackbox testing, they now have a problem that only they have.
I can try to answer my own question: they're Google, they can afford a team scanning for vulnerabilities like this, a team dedicated to analyzing codebases that are found to be compromised, a legal and PR team to handle the fallout if this kind of vulnerability makes it to the public.
In short: horrible advice to follow if you are not Google.
1. Most of CI systems have better options for caching node modules[1][2]. When you check node_modules, you add fixed cost(increasing) for every commit. When you use CI caching you add fixed cost(static) to only small number of builds.
2. Once you start upgrading packages repo size will continue to grow. At some point you will be forced to git filter out node_modules. You will lose the ability to run locally older commits.
3. You will need to pin version of npm/yarn because structure of node_modules depends on hoisting algorithm. Every upgrade of node will be extra painful because you also potentially need to upgrade all yours packages.
4. Platform dependent modules like fsevents, node-sass can be broken if you use different OS. You will be forced to only support a single platform(linux).
5. Impossible to resolve node_modules conflicts. Modern package manager have git conflict resolution build in. If two people update the same module to the same version they can still create a merge conflict when node_modules are checked.
6. Currently, you have plenty of good options that can achieve the same with smaller effort. You can use yarn2 with node_modules linker and local cache. This would create .yarn folder in the repo that have all modules as zip files. During install it would use these to hydrate node_modules. Alternatively, You can use pnp and have zero install but with proper support[3]
7. You lose automatic audit and dependency management. Current best practice is to use something like dependabot or renovate bot. Once you commit your package, you will no longer be able to use this effectively.
8. Most people commenting on left-pad are maybe not aware, but today npm is immutable, and you simply cannot unpublish public package[4]. Because npm.com is such vital infrastructure, it unlikely that it would ever stop working.
[1] https://docs.github.com/en/actions/advanced-guides/caching-d... [2] https://circleci.com/docs/2.0/caching/ [3] https://yarnpkg.com/features/pnp [4] https://docs.npmjs.com/policies/unpublish
Some of your points are valid, but the first sentence seems really hand-wavy and rude.
> 1. Most of CI systems have better options for caching node modules[1][2]. When you check node_modules, you add fixed cost(increasing) for every commit. When you use CI caching you add fixed cost(static) to only small number of builds.
I was wondering about this too[0], mostly because it's the setup I tended to gravitate towards for my projects. Gitlab CI and Travis can do this AFAIK, although I'm not sure how long the cache folders would be kept (especially on the Free tier).
Because people here are discussing this and entertaining this idea. This creates a level of legitimacy. After this article, there might be now countless teams transitioning to this crazy idea. Then two years later, people would continue to complain about node ecosystem because they were burned badly by projects maintained using this approach.
The problem with today word is that everyone tries to be politically correct. Everyone wants to discuss things in a civilized way based on merit and logic. In many cases, we could avoid wasting time and energy by declaring things as they are. For example, If mainstream would call anti-vaxxer stupid, we would have now more people vaccinated
There are more things that people spend large amount of time discussing when there are nothing to discuss. This thread should not have 195 comments.
Except that is has? It's a service, like everything else on the internet, it can and will go down. The choice here is whether you can carry on, or wait until it comes back up. THAT is the main benefit I see in this. (Not everybody wants to manage a local copy/proxy/internal NPM registry etc)
We had left-pad type incident where the package we depended on disappeared from npm, and didn't even realize it till the next time we setup from Git.
Since then, we have node_modules as a submodule tracking latest-commit, and it works.
I don’t think committing ‘node_modules’ is egregious, and it has benefits as the post describes, but unfortunately it’s far from _the_ solution to the problems described. Something like Google’s own Bazel is a better option.
Nowadays, most new large codebases run on TypeScript or some recent version of EcmaScript.
The rant the article is based on is not really about language itself but the tooling and ecosystem around it.