RFC: Make NPM install scripts opt-in
github.com
github.com
leftpad has no need of install scripts, nor `eval`, reflection, or access to my disk or the network. nor should it be allowed to gain them in the future, at least without a million alarm bells ringing and explicit approval.
ACLs would allow establishing "moats" of dramatically-more-difficult-to-attack libraries, and encourage libraries to voluntarily reduce their attackable surface to make them more likely to be approved. Instead we have this.
Some Linux distros also ship SELinux or AppArmor policies with packages by default, but that seems like a kind of inside-out proposition from what you're describing.
But I want this in process. Because libraries are in practice very frequently treated as black boxes (like binaries)[1], but without any ability to limit their access like we can for users/binaries/etc.
There are some WASM things that do kinda what I want and pull off neat end results, but they're kinda tough to use together, and have obvious perf/debuggability/etc costs. e.g.: https://github.com/dtolnay/watt
[1] unfortunate in many ways, but largely reasonable IMO (simpler abstractions of things we don't want to think about is the whole point of shared libraries), and I don't believe the industry is capable of reversing course on this either way.
Far from reversing course, the industry seems to be interested in accelerating on the current one— containerization as a deployment mechanism is about doubling down on the black box-ness of binaries. Paradoxically, this doubling down on ‘the whole point’ of shared libraries is also a destructive innovation with respect to the advantages of shared libraries, because it also means surrendering the fine-grained resource sharing that shared libraries, as packaged traditionally, allow.
Security is not binary, nor can it be. Air-gapping isn't even enough, as is repeatedly demonstrated. That's not a reason to be yolo.
1. Each package declares what permissions it needs in package.json (if any).
2. I declare in my project's top-level package.json what should be allowed.
3. npm install fails if any dependency's permissions exceed what I've defined in 2, and points me to the offending package(s). Then at runtime just my project's restrictions can be applied across the board.
The same kind of permissions system would also help warn you when you're importing a "does far more than you want" sub-tree. If `pad` imports `leftpad` and tomorrow `leftpad` imports the entire Kubernetes codebase, you might actually notice it and fix that instead of your builds just getting a bit bigger and hoping a reviewer notices the lockfile diff (which is probably hidden by default!). Again, not perfect, but better... and applies pressure to do better, possibly improving the ecosystem as a whole.
Future languages/tools/etc could be dramatically more intelligent about these kinds of things, but even the crappy basics don't exist in the extreme majority of cases. You could build it yourself many times, but almost nobody does that. Default behaviors matter.
The reason to disallow importing transitive dependencies is that we can write packages which use that functionality, but only expose a safer/restricted form; e.g. we might forbid dependencies on `eval`, but allow dependencies on a `mock` library, even if `mock` happens to depend on `eval`.
A sibling points out that restrictions can be broken when different packages pass callbacks to each other. Thankfully, that's exactly what capabilities are for (code is written without any "ambient authority", and instead requires capabilities to be passed in as arguments; e.g. see the Confused Deputy Problem)
So what you actually need to do is split your program into more VM units, not unlike code splitting in bundler-integrated JavaScript HTML5 routers. You can build that atop a WASM runtime, if you pass the ability to call other WASM modules into other WASM modules. And then you could pass literal capability tokens around. You just need to split your code up.
Then you can bridge them back again with trusted code, and you can add additional constraints on API calls between modules (e.g. callbacks don't work across units, or if they do, then they remember the capabilities they were instantiated with, something like that).
If there were an easier way to have your dependency management system do this code splitting for you, then it would be feasible.
Packages really shouldn't bundle their dependencies anyway.
Capability security is generally based on the following:
(1) A "capability" to perform some action is (by definition) something which is necessary and sufficient to perform that action. If that action is calling a particular function (like 'eval'), then we can treat the function's name/reference as a capability. This fits the definition since (a) we can only call functions which we have a name/reference for, and (b) anything with the name/reference to a function can call that function. For example:
function somethingWhichCanUseEval(eval) {
return eval("foo");
}
function somethingWhichCannotUseEval() {
throw "Not given a reference to 'eval', so we don't have anything we can call";
}
(Note the contrast to your statement that "runtimes can't tell which dependency a particular API call is coming from". Such functionality would actually violate (b) above, making it much harder to implement capability security!)(2) Capabilities are kept secret to begin with, and we pass them around to whatever needs them. In particular, this requires that (i) we can't ask for a list of capabilities (e.g. a list of references to every function), and (ii) we can't guess a capability (e.g. names should be cryptographically-random strings). For example:
// All guessable references to the eval function need to be shadowed, e.g.
function eval() {
throw "Error: Something attempted to call 'eval' without using the package";
}
window.eval = eval;
// etc.
// The package manager can give each package a cryptographically-random "import-name"
// When a package is imported, it's given the import-names of those packages it directly
// depends on
function myPackage(deps) {
// The 'deps' argument is a map from package names to (randomly generated) import-names
// We can import a package if we know its import-name, hence import-names are capabilities
// for importing packages
// We can import the 'eval' package by getting its import-name from the 'deps' map
const evalPkg = require(deps['eval']);
// The 'eval' package provides a reference to the 'eval' function
const eval = evalPkg.eval;
// At this point, we have the capability to call 'eval' (or pass it around to other functions, etc.)
}
(3) If we find ourselves "checking permissions" then it's already too late (those without permission shouldn't have even been capable of asking in the first place). Likewise if we try to "validate the caller's identity", since (i) that's a case of ambient authority (the "caller" may be a Confused Deputy), and (ii) it would prevent us delegating capabilities to others to act on our behalf, which in turn requires granting sweeping access to a broad range of systems 'just in case' anyone might want to use them.(4) We can wrap capabilities in a more restricted interface, which acts as a capability for some more restricted action. For example, the following code provides the capability to evaluate code of the form 'foo = bar;':
function namerPkg(deps) {
const eval = require(deps['eval']).eval;
// Regex for valid Javascript names
const jsNameRegex = /[_a-zA-Z][_a-zA-Z0-9]*/;
return {
// A function which evals '<oldName> = <newName>;'
'namer': function namer(oldName, newName) {
// Parse/sanitise the input
const oldJSNames = oldName.match(jsNameRegex);
const newJSNames = newName.match(jsNameRegex);
if (oldJSNames == null ||
newJSNames == null ||
oldJSNames.length == 0 ||
newJSNames.length == 0) {
throw "Error: namer needs to be given Javascript-compatible names";
}
const oldJSName = oldJSNames[0];
const newJSName = newJSNames[0];
// Now it's safe to call eval
eval(newJSName + " = " + oldJSName + ";");
};
};
}The package manager fetches dependencies recursively into your chosen directory running as your choose account. The runtime executes with your chosen account with whatever permissions have been granted.
Now having the runtime track whether the caller of a library is another library or your application. Or for the runtime to allow reject runtime api seems better left to the os to control with container, sandbox or transitional ACL
I do think though that as it seems to be only a runtime thing, it's probably warning you far too late. If it were baked in statically so, when you pulled in a new http lib (or just updated it), you approved the popup that said "let X access network and files?", it'd probably be a lot less painful 95% of the time. (since sometimes you'll want it more fine-grained)
---
Since I can't glean it from what I've skimmed so far: does SecurityManager let you do capabilities-like things? E.g. can you be given a SecurityManager as an argument, allowing code using that manager to temporarily access a file, while normally blocking it? Or is it something closer to "applies to a ClassLoader"? Though with enough effort and runtime cost you could make those equivalent, of course.
I have a very simple vue frontend app that I wrote a few years ago, and it somehow has >4000 dependencies (including dev dependencies). The fact that npm install could run code from all of those (which might not be obvious to a newer dev) is flat out dangerous.
Production deploys tend to be if you use the right tools. Docker images, cert signing, ACLs, network policies, etc. But we have no equivalent for developer machines. And engineers have access to lots of dangerous things. Docker alone isn't good enough.
I think longer term, we're going to find things like Qubes's VM for everything model becoming more normalised.
And builds are sandboxed, which helps with the OP issue
Basically, I'm trying to imagine a world where every Bash script needs to come with headers specifiying binary dependencies and file paths (and env vars) it wants to access. Upon invoking said script, it gets executed within a light-weight container, the binary dependencies then get installed on-the-fly and the file paths get mounted into the container (after I've confirmed that the script may access them).
There are pieces of this that have been implemented in Nix, Distri, and Flatpak.
Binary dependencies getting installed on the fly used to be something you could enable systemwide on NixOS, and parts of it still can be. There are also FUSE filesystems that can automatically build/fetch paths as they're requested. DwarfFS, for fetching debug symbols, is the first one I'd heard of, and then there were some experiments like this in Distri.
In Nixpkgs, scripts built by b Nix are created in a way that bundles in the executables they call by interpolating in full paths to deterministically built versions.
In Flatpak you have this notion of portals, where you declare file paths and other things where your program is allowed to talk to anything else. It seems a little more seamless than sharing the filesystem on Docker.
I think this is pretty achievable in the moderate term.
You get snapshots and if your drive image is small enough you can toss it on a USB SSD drive and take it wherever you need to.
Docker was never good enough. Especially if you're on a Mac. It's a Linux VM with more overhead added and greater access to the host making it less secure than just using a VM in the first place.
If you don't trust those libraries enough to run code during installation, why do you trust them enough to run code _after_ installation?
What we really need are content security policies for node. I want to define at the top level of my project exactly which file system directories and internet domains can be accessed, then have that enforced by the runtime.
For review after install I install and fire up my IDE.
For review before install I have to manually download the package and figure out the dependency tree and do that for all packages.
1. Package installation often happens under different privileges than actually running your end user app. There is unfortunately a lot of "just use --unsafe-perm to make the install work" advice out there which means a lot of people are installing packages as root even though they're not running as root. Also consider that a lot of npm packages ultimately get run in the browser, so the attack surface there should have very little to do with the files on your computer, but install scripts make this not the case. For the same reason, this means that this might be the only kind of attack that has the ability to run on your build machine, since your build machine may not actually "run" your app.
2. The fact that the code doesn't have to be run to be effective makes it quite difficult to spot, and thus increases its likelihood of sneaking into your dependency chain. For example, GitHub will often straight up just hide large `package-lock.json` file diffs, and no one wants to read through those. So a very innocuous patch version change in an otherwise unrelated bug fix could completely fool the PR reviewer: this is because there's no code "history" that would imply any sort of entryway into attack. All the code looks totally reasonable. By having to explicitly allow the package to use install scripts however, all of a sudden the PR would contain a clear indication that new foreign code that isn't represented anywhere in the commit will be run. This is huge.
3. This also takes advantage of the fact that the vast majority of npm installs aren't done by humans, but by CI and production deploys. Again, by ensuring that the attack is completely "passive", that is to say, as long as you get installed you've succeeded. A lot of times this can be triggered just by issuing a PR with automatic build and CI machines. The correct behavior for something trying to run a new script on your machine when a PR is submitted isn't to infect your CI machine, it should be for the test to fail with "package X tried to use an install script".
Ultimately, I still think we need real sandboxing built in to the runtime to really solve this problem. Is there something inherent to node's architecture that would make this impossible or is it just a ton of work? Could Chrome's sandbox and CSP implementation be piggy-backed on perhaps? Deno takes some steps in this direction but they stopped well short of what's really needed.
Could you say a little more about what's lacking in Deno here? I know it exists and I know it's taken some new security measures, but not much else.
That said, perhaps it’s just a starting point and will be expanded into something more useful in time. It would be a killer feature imo if done right so I hope they go in that direction.
For example if we detect that 1 million packages depend on leftpad then add leftpad int eh standard library, if 500K packages depend on isValidUrl,isValidEmail then put this int eh standard library. I am sure Google devs could over engineer this to make it optional and include only what you need.
Also https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
You can make it a library , name it STL-extras. So instead of having 10 packages with 10 copy of left pad, 12 packages with some copies of isOdd and other ridiculosness you have this 1 STL-extra library that is reviewed by Google and Mozilla.
You can come with other solutions to reduce the problem further, like some of the other ideas suggested, one of the ideas is to maybe make it less CV worthy to have a npm package to discourage the behavior we seen where someone creates or forks some package and then pushes it into other projects just for fame.
I apologize if I miss-understood your point, honestly is not clear what you did not like about an optional STL library that is trusted and people could use instead of 100 untrusted leftpad like ones.
To address the resistance, I would propose a compromise, namely that install scripts won't be run by default unless the publishing account is secured by 2FA, or the previous version of the package also included an install script. That should greatly reduce the attack surface, and pave the way towards requiring 2FA for all packages with install scripts as a later step.
But how does it raise the cost of attacks? I don't see why it would be harder for someone uploading a malicious package to embed said malicious code in the index.js instead of in the install scripts.
That makes it a lot harder to steal ssh keys.
Package ua-parser-js wants to run a script before installing.
The description of the package is:
"Detect Browser, Engine, and Device type/model from User-Agent data."
The reason for the pre-install script is:
"Configuring the local user agent thing for reasons."
This script has been unchanged since version 0.7.29 which was published:
14 hours ago
The hash of the script is:
0123456789abcdef
Press Y to examine the script, or N to cancel installation.
After npm echoes out the script, the user should decide whether it looks obfuscated or does anything suspicious. If the user is still unsure, they can search the web for the hash of the script to see if other people have audited it.For automated installs, such as a CI server, there would need to be a command line argument or config file entry with something like:
allow-preinstall-scripts: ["0123456789abcdef"]I think the problems of npm and the js world in general is the depth and breadth of the dependency trees and the misunderstanding of transitive dependencies. I've heard devs say that they only have a single dependency when using CRA (which IIRC pulls in 1500ish transitive dependencies). I've heard devs say that a dependency is always better (even if it replaces a single line of code) than your own code.
Besides that even simple dependencies seem to be updated even when they were "done". Many devs see a repo that has not had a commit for a few months and consider the project dead instead of done, so there is an incentive to keep updating things that didn't need updating or for dependents to switch to "actively developed" projects for their dependencies.
So yes, requiring 2fa would be great, making the install steps not able to run arbitrary code would be great, and most of all requiring repeatable builds from source would be great. I still think the problem at the core would remain, which is that the ecosystem is too hooked on excessive dependency usage and sees newness as a virtue.
Given NPM's increasing viability for large scale supply-chain attacks, there are other things to worry about still. Perhaps those other things are more fundamental to NPM's design and can't be changed overnight. It's still helpful to solve these simpler problems.
The more realistic solution would be teams of volunteers that are auditing the packages and check the differences between specific versions of those. This doesn't block all possible infected packages, but most of them, which is better than what we have now. Everything is based on trust so you can't stop this, but maybe prevent it.
Ultimately, any code inside an npm package needs to be run by default in the context of a sandbox, such as vm2 or SES. That way a developer would have to opt in to granting permissions for a package to run executable code.
https://github.com/patriksimek/vm2
https://medium.com/agoric/ses-securing-javascript-in-the-rea...
I'm working on a solution along the lines that you've suggested:
https://github.com/vouch-dev/vouch
Vouch lets users create and share reviews for NPM packages. Project dependencies can then be checked against those reviews.
Vouch uses extensions to interface with package ecosystems. Extensions currently exist for NPM, PyPi, and Ansible Galaxy.
I'm currently working on a website to index known reviews and publish official reviews.
Drop by the Matrix channel if you have any feedback or thoughts to share: #vouch:matrix.org
There should be granualar control over which packages, and which scripts you want to block.
E.g in your package.json:
`"scriptblock": { "puppeteer" : "*" } ` or ` "scriptblock": { "puppeteer" : ["preinstall","postinstall"...] } `
etc.
yeah, because they're allowed to. The only time that install scripts are sane and reliable is when they are part of a community that polices them and requires them to be minimal (e.g., in Linux distros). That's basically never the case with language-specific package managers
I run an install script for my package [0], this install script forces 2 lines in users' .gitignore files to prevent EXTREMELY sensitive data being committed into a public repo - despite me telling everyone in the community to not commit these files on god-knows how many occasions. I do this because I care about my users and I know if I tried some funny business then they would rightly call me out for it - NOT "because I'm allowed to".
[0] - https://github.com/open-wa/wa-automate-nodejs/blob/0ba7d6850...
1. The most "legitimate" use of install scripts is to build binary node addons. N-API is a much better alternative with ABI compatibility that will allow the package to just ship built, which aside from being a much friendlier and faster install experience, will also allow for more deterministic builds and better caching.
2. Many install scripts just transpile code on installation. It would be much better to instead just transpile the code on publish. This is something that requires nothing new and everyone should absolutely just do today. It's crazy that people re-download and retranspile the same code over and over.
3. Many install scripts simply download non-npm resources, like apt dependencies or things like Chromium. It would be a lot better, a decade after npm's introduction, to standardize this common use case like this instead of having everyone roll their own. We actually implemented our own version of this where we can just list apt dependencies in our package.json. It gives you so many more options on how to optimize builds and provides you so much better information about your dependencies to have this be a "declarative" piece of the puzzle as opposed to having every package figure out their own slightly different way of doing this.
4. Many install scripts just present the user with ads. NPM should also just add a package.json field for donation links or whatever and provide a good standard way of disabling this as opposed to spinning up a process just to present ads.
5. And of course, some install scripts are used maliciously.
All this to say, 10 years later, there aren't actually a bunch of "wow, what an interesting use of the install script" examples, but rather just a bunch of fairly limited common uses (that could be better handled by npm itself), and yes, some interesting stuff, but also a worrying amount of abuses, that are easily duplicated and that we are currently doing nothing to stop. Given a proper transitional period, where npm just warns you above an install script rather than failing, I think most packages could be made to not require custom install scripts. And again, this is for a small percentage of packages, nowhere near the undertaking of converting packages on NPM from CommonJS to ESM for example. And, it's not that it would be disallowed, it's just that the rare package that really did need to do something really out there would need a flag. Having actually queried all the packages, I am certain this could be done with basically no disruption for end users, certainly no disruption to most package authors that don't use these scripts, and almost no disruption for the few package authors that currently use this feature.
https://docs.npmjs.com/cli/v7/commands/npm-fund.
Point 3 is interesting, does it work across non-apt package managers?
At the very least it's not harmful. You could use something like repology or your own index to identify equivalent packages and then install them.
Really what you'd want would be a package manager-agnostic way to declare dependencies on things outside the Node universe, and then plugins for NPM that could use to assert they were satisfied in an appropriate per-package-manager way, and maybe optionally try to install them for you using the appropriate external package manager.
But the packagers for Node projects wouldn't be writing the mechanisms for that, they'd just be declaring the deps.
`install: "malware"` is a few steps closer to running than `cp.exec('malware')` in index.js.
Also the install script may run with higher privileges than the one that the program would have when it's executed. For example there is the case of installing packages globally with root privileges, ideally something that you shouldn't do, but how many people runs `sudo npm i -g something`? In that case the install script runs with root privileges.
Of course there are rare cases when they are needed: but that cases should be an opt-in, meaning that the user is prompted to run the script (if running interactively) or the user has to provide a command line option to allow specific packages to run scripts, or add an option in the package.json to whitelist specific packages.
how many Linux users copy&paste the command from random webpages without doubt? It's 100% opt-in. So it won't help at all.