Alert: NPM modules hijacked
drinchev.com
drinchev.com
Most programming-language communities manage the code-reuse/dependency-hygiene tradeoff by concentrating code into a relatively small number of large libraries, and when people want a function or two that isn't in one of the libraries they're using, they incorporate it by copy-paste. In the Javascript/NPM world, on the other hand, I see a lot of projects with dependency references to huge numbers of tiny dependencies.
Today, we're seeing one of the reasons why that's a liability. Most people will take away lessons about code signing and signatures, and presumably the NPM software is going to be improved in that regard. But the other lesson is that projects should have fewer dependencies. Using a library may be as simple as adding one line to a package configuration file, but using a library properly requires substantial due diligence on the library, its license, its bug tracker and its author.
As it is, this is the first time I've logged in for about 12 months, just to berate you.
First lesson learned is of course in regard to how package managers such as NPM should handle scenarios like this. However, I would also hope this might make some people take a harder look at their dependencies to see if everything they are referencing is both truly needed and trustworthy.
>22 Apr 2015: By this date, SourceForge Open Mirror Directory has expanded to take over popular former SourceForge projects which left the site, such as VLC and GIMP for Windows. Some projects have malware embedded despite earlier assurances that the adware system would be purely opt-in.
>16 May 2015: GIMP developer asks SourceForge to remove gimp-win from the site. They don’t. SourceForge later claims they didn’t receive this message.
Yes this particular episode is related to the unpublishing results... however the underlying issue is not only about unpublishing results. The underlying issue is:
number 1, Having thousands of dependencies is a software engineering trade off, their is a upside and a downside of that trade off. Having Thousands of dependencies is fundamentally more 'risky' than have a few. Period. End Of Story. It is ok to accept the trade off based on what you are optimizing for. It is not ok to refuse to accept the fact you are making a tradeoff.
number 2, is you should not be running production builds or builds for important components based on dependencies on the naked internet. This is also an engineering tradeoff but is firmly in the camp of what is considered 'bad practice'. This is fundamentally the same exact underlying reason why we have build servers instead of just builds from developer work stations. Developer workstations are not controlled environments, which means the builds are not guaranteed repeatable. A build server that depends on the state of the naked internet is not a controlled environment. It blows up today because of unpublishing, it will blow up tomorrow because of something else.
I think it's a worthy exercise, since the prize is no less than efficiently distributing coding effort across the globe.
Problems like we've seen today are a double edged sword. On the one hand they're troubling, but it'll focus the community on the problems, and hopefully npm will make a few changes too.
> A build server that depends on the state of the naked internet
Whether npm based or not, don't the vast majority of build servers, for whatever language or platform, rely on the state of the naked internet? Whether that's downloading binaries, compiling from source, or whatever?
A side note, if your build depends on the box it's being built on, your build process is bad. The build from any dev workstation should be identical to any other build. People have build boxes to automate building/deployment and testing and alerting people things broke; they aren't about dependency management.
If you're using a build box because of dependency management issues with dev workstations, you're doing it wrong.
No, this is misrepresenting the issue. The amount of dependencies you have is just as irrelevant as the amount of lines they contain.
What actually matters here are the following things:
1. How many different entities am I trusting?
2. How hard is it to audit this code?
3. How easy is it to 'lock in' a dependency version without missing out on updates I need?
4. How much work is it to review updates?
There's no significant difference between small and large dependencies for point 1. You can choose how many entities you trust, regardless of whether the functionality you are using is split up into many small packages from the same author, or combined into one big one.
There will usually be no difference for point 4, because the amount of updates generally relates to the amount of functionality that is being updated. This means that larger modules means less CHANGELOG files to track down, but smaller modules means less chance of an upgrade affecting something you're not using to begin with. It evens out nicely.
However, for points 2 and 3, there is a difference - in favour of small modules. A much smaller, more well-defined surface with less implicit dependencies and coupling. These are the metrics that actually matter.
EDIT: To clarify, 'smaller' and 'larger' here refers to complexity, not line count.
I don't think it's necessarily "most programming languages". I know Perl, and probably Python and Ruby, have very large ecosystems of small modules. I suspect it's more along the lines of static vs dynamic, or compiled vs interpreted.
> On the other hand, each dependency you add creates a little bit of risk: that the package will be updated in a way that breaks your application, or the package maintainer will go rogue, or the package will get hijacked.
Those specific risks are not necessarily from including a module, but are from subscribing to a module. If you include a specific version of a module and can confirm it has not changed to a degree that satisfies you (whether that's keeping a local copy to build from, or trusting the distribution system to be immutable and pegging a specific version to install), then your risk is fairly well defined. It's not zero, but it is limited to the problems the module existed with at the point you reviewed and included it.
Subscribing to a module, that is having the system download and build "the latest and greatest of whatever thing you call $foo" is not very safe, and if the build mechanism can execute arbitrary code and is automated, is insanity.
That only works if the central registry is append-only. But as we've seen packages can be removed from npmjs, which caused these problems.
The Perl, Python, and Ruby communities might have a lot of "small-ish" modules, compared to something like Boost in C++, but the Node.js community has really taken it to a new extreme with an explosion of "one-liner" modules like the absurd user-home https://github.com/sindresorhus/user-home
The rallying cry of this movement is "Modules are Cheap in Node.js!". Incidents like this help demonstrate that explosive growth of your dependency graph isn't ever cheap, in any language.
Alas due to the beauty and power of combinatorics, there will inevitably be more compound modules than atomic ones. If there are not already.
I don't see automatic combination of lots of small modules actually making this situation better in any way. Quite the opposite, actually...
The example you gave was a poor choice - read further down https://github.com/sindresorhus/user-home#why-not-just-use-t.... It exists to avoid exactly the problem that sparked the conversation around npm today.
But there are myriad other equally ridiculous packages. is-positive, for instance: https://github.com/kevva/is-positive.
Abstractions are good when they hide and generalize complexity. These modules merely conceal mundanity.
> Edge cases are often identified and accounted for, even if the responsibility of the package is tiny.
In my observation, the opposite is true -- most of these trivial packages punt on edge cases. is-positive makes no attempt to handle any. https://www.npmjs.com/package/average implements an utterly naive calculation of an array's mean that is numerically unstable in the presence of JS's everything-is-a-float semantics.
Hiding tiny, naive implementations behind dependency imports actively harms understanding of the implementation details and failure modes of the code being relied on, in addition to making your build much more fragile by encouraging you to rely on external dependencies that can and will vanish from the Internet at inconvenient times. It's a bad trade-off from every angle.
Edit: this blog post (currently on the front page!) does a really good job of capturing my utter bafflement at how completely crazy the node community has gone with this nonsense http://www.haneycodes.net/npm-left-pad-have-we-forgotten-how...
Edit 2: and none of this even touches on the complete insanity that is npm running arbitrary code in each dependency at install time. Having 2,000 one-liner trivial dependencies amounts to inviting 2,000 strangers to run arbitrary code on your build server. The node.js community appears to be a pack of complete amateurs.
https://github.com/sindresorhus/os-homedir/blob/master/index...
I'd rather `npm install` it and require it. Of course, now Node has this: https://nodejs.org/api/os.html#os_os_homedir, but it didn't and you get the point.
The problem was with npm and allowing unpublishing/hijacking, not the Node community.
A low barrier of entry for submitting modules is generally a good thing, but if it's too low, you end up filling your ecosystem with crap.
If your development environment is such that you can't test these things before going into production, and that production updates its packages on its own, then you have much bigger (and harder to fix) issues that dependency hygiene.
How about 3 years after a release? 5, 8, 10?
Programmers loves dependencies because it lets them pass the buck, but each and every 3rd party module is a potential time bomb waiting to happen.
The HN "web developers" are a different crowd.
I suspect JavaScript ended up this way because people wanted to avoid downloading anything they absolutely didn't need in the browser. Now packagers can automatically remove unused functions (WebPack 2 does this IIRC, and I think require.js did ages ago) so it's no longer a concern but somehow NPM still ended up with a million tiny libraries.
Thank god Underscore/lodash didn't use the approach of one-module-per-function.
The difference is they also provide the rollup bundle with all of Lodash. Which is what most users actually use.
My understanding of the left-pad issue was that it was frequently down the dependency tree where dev's themselves could not remove the dependency on left-pad without modifying or branching a dependency (or perhaps multiple).
If we copy and past snippets of code between packages, we lose out on potential code de-duplication from dependency flattening.
If you're really worried about your dependencies disappearing/changing beneath you just check them into git.
For those who go this route, any advice on what you can do to keep your version control history from blowing up?
So instead of checking in your dependencies into your app repo directly, you turn your dependencies folders into submodules to keep their histories separate from your app repo's history.
Makes sense! I'll be sure to give that a try on my next project.
There's a bit of extra overhead because you have to fetch them separately after cloning a repo, and updating them is separate from pulling origin/master. If you do one without the other you can probably get into unexpected version mismatches where your project doesn't play nicely with the dependencies.
Was on vacation and found out that I had lost a published package name within 24 hours of the dispute request. Broke a few production systems. Really messed up my day.
Ended up having to beg with the person who filed the original request and they eventually gave me the package back.
Honestly, the whole process was a bit personal and I felt like I was being singled out as an individual by NPM, rather than being treated like a developer who was using the service. Not a nice feeling.
That's a reasonable excuse in the 90s, before automatic updates over HTTP were common. Our industry now has decades of experience securing HTTP updates and package managers, with various Linux solutions demonstrating good practices.
And people wonder why big enterprises are scared of touching open source stuff.
some open source stuff. Most enterprises dig distributions, especially with LTS.
> This is the sort of amateur behaviour that is emblematic of the entire nodejs community.
This is pure mudslinging
https://www.reddit.com/r/rust/comments/4bm3rk/how_would_crat...
That's a strange thing to say considering how insanely complex modern hardware and software are. While the NPM setup is indefensible, there are countless ways a stranger can prevent your deployment, if that stranger works at Microsoft, Intel, or the manufacturer of any production hardware component. OS bugs, driver bugs, compiler bugs, and hardware/microprocessor bugs are a thing. At any point, you are dependent on the work of many thousands of strangers, no matter how perfect your decisions are.
It destroys days of development.
That's the part that really got me. Questionable dependency practices is one thing; but running whatever JS when a package is installed, with no warning whatsoever, is probably a very bad idea.
A year ago I never ever would've thought I'd be happier writing C than NodeJS, but here I am. Weird how that works.
My semi relevant tweet, it's just asking for it with that warning
It is insane that modules don't have signatures (that are actually verified, of course). Because npm can feed you basically anything.
It's not a problem that something else gets published under the old address. It's perfectly natural. The real problem is the trust model - that new content's accepted without even warnings.
Even if I have a dependency locked down in my NPM shrinkwrap file, it can change underneath me? That is pretty absurd and gives me zero confidence in my packages. It means I MUST commit them to source control or risk having my project completely broken some day.
I thought for sure that since NPM removed the ability to re-publish the same version of a package with different content that they also wouldn't let you remove versions of a package. It also means you should never user the "^" version specifier or risk downloading some completely different project.
They do have a "warning" at https://docs.npmjs.com/cli/unpublish: > WARNING > > It is generally considered bad behavior to remove versions of a library that others are depending on!
Seriously, what good does that do? Nobody takes warnings seriously.
If you depend on a version range like ^1.2.3 or ~1.2.3 its a different story, of course.
Moral of the story is imo always pin to exact versions and use shrinkwrap for production apps.
EDIT: This may not be true .. see below.
That's not true. I thought it was but it's not. That's the terrible part.
To demonstrate with one of the packages that was removed, run:
$ npm info andthen
You'll see that versions 0.0.1 and 0.0.2 were published at one point. However, for "versions" it only mentions 2.0.0. And of course, if you run:
$ npm install andthen@0.0.2
It blows up in your face.
> and the content of the files is suspicious
The script seems as though it might publish your entire codebase to NPM. Sorting NPM packages by creation date[1] reveals[2] a[3] few[4] potential victims. Filtering by the ISC license also seems to work.
[1]: https://libraries.io/search?order=desc&platforms=NPM&sort=cr...
[2]: https://libraries.io/npm/alaska-dev - internally hosted repo, edit: license doesn't match other projects by the same developer
[3]: https://libraries.io/npm/b3app-prototype - private bitbucket repo, edit: deleted from npmjs
[4]: https://libraries.io/npm/nodework - private bitbucket repo
It takes one argument, a package name, then attempts to publish that package.
So
cat list | xargs -I{} ./x {}
was what I used to publish the whole list.Anyway did you think of explaining your intend to doing so?
I've been doing development for years and I've never seen a developer naming files "x.sh", except if it was not malicious.
*edited npm capitalization
Then continue working like you always have.
There are some changes that you'll need of you are using native modules, but they are simple and easy to do.
Learned this trying to fix Shrinkwrap a few weeks back but seems like it's a good practice for securities sake now.
`npm shrinkwrap` exists to do that. Applications should use it to pin the versions of all of their dependencies and sub-dependencies.
Rather than artificially constrict your version ranges, NPM should support real version locking, and applications should check in their lock file.
"npm" is two things in this picture: the server and centralized service, and the software tool on everyone's build/dev machines.
Folks have historically been happy trusting the centralized npm server to behave consistently and pleasantly, and some opinions have recently shifted on that. But frankly, the server doesn't matter, in the big picture. The npm tool on your computer does. It's the one executing code on your computer with all of your local user's privileges.
This tool starts executing new code from a new author on my host as my user without any authentication except "trust the server". This is exactly the same words we would use to describe behavior of $script_kiddie_virus_of_the_week:
> "download code from C&C server; run it, thanks:D"
What's the difference here? "good intentions"?
I'd rather have something more than "good intentions" controlling what code ends up running on my computer. Wouldn't you?
I updated the blog post.
Nevertheless I find his actions dangerous and irresponsible.
I personally think the best course of action would have been for the NPM team to immediately blacklist these names (aside from left-pad, that's a separate conversation) after the entire list of unpublished packages was shared, and then make them available on a case-by-case basis.
nj48 is a known friendly who has identified himself to us. We're going to clarify later today.
Also, relevant conversation here in yesterday's thread: https://news.ycombinator.com/item?id=11340510
Relying on the notion of a "known friendly" to protect packages and namespace does not strike me as a sound practice.
As others may have mentioned, publishing packages really should be fire and forget. If something bad goes out, a replacement should be sent out. And for the life of me, I don't understand why they did not go with the <author/package> scheme.
> Some things are not allowed, and will be removed without discussion if they are brought to the attention of the npm registry admins, including but not limited to:
...
4. "Squatting" on a package name that you plan to use, but aren't actually using. Sorry, I don't care how great the name is, or how perfect a fit it is for the thing that someday might happen. If someone wants to use it today, and you're just taking up space with an empty tarball, you're going to be evicted.
5. Putting empty packages in the registry. Packages must have SOME functionality. It can be silly, but it can't be nothing. (See also: squatting.)
Ideally the developer would sign before publishing and the consumer could check the signature to validate before using.
Whilst not a silver bullet this is a kind of essential part of a secure package management solution.
What repositories were you thinking of that do require that?
I'm guessing @nj48 used a script to go over the official list of unpublished modules and attempt to generate a placeholder for each of them.
The same user also published some actual (unmanipulated) forks of the original modules, so I'm guessing this is mostly a quick move to preempt any malicious hijacking by others.
There is no reason to assume the "hijacking" is malicious. Certainly not in the scripts. The user is active on GitHub and shows no indication of malicious intent.
However it IS worth pointing out that the unpublished modules should now be treated with caution if you still rely on them because even if the replacements are identical and benevolent you probably need to take action.
Yep, that and it publishes your project to NPM!!!
edit: I was wrong, the x.sh script looks malicious but appears not to be? In any case it shows that installing NPM modules is just as unsafe as 'curl | sh'.
Ways to be safe are to first check what scripts will be run by a package: 'npm show $MODULE scripts'. There is also the '--ignore-scripts' flag[1].
[1]: https://blog.liftsecurity.io/2015/01/27/a-malicious-module-o...
As it is, TFA seems like an attempt to invent malicious rumors about a completely innocent person.
[EDIT:] you're still wrong; in order to be equivalent to "curl | sh", x.sh or its equivalent would have to be referenced in the "scripts" object in "package.json".
I don't think there's anything malicious here. The file x is simply a list of the modules that were 'given up', and x.sh is used to loop through that list and republish them as placeholders, meaning other people can't use it for malicious purposes. The author seems to be in good standing and has even republished the original code to a few modules
Now I hope NPM author will come to their senses and change the way NPM works. But the trust is broken, no question. Between stuff like that, people selling "realestate" on NPM for real money... Nodejs has been the least professional platform I have ever used. Everybody's out there to make a quick buck, nobody gives a damn, this is whole thing will collapse sooner or latter.
Finally people need to stop with the "unix philosophy" excuse. Importing 10 lines of code from a random package is not the "unix philosophy". Splitting a 100 loc module into 20 packages is not the "unix philosophy". Packages have so many dependencies it's getting ridiculous.
edit: corrected
comm -12 <(ls ./node_modules) <(curl https://gist.githubusercontent.com/azer/db27417ee84b5f34a6ea/raw/50ab7ef26dbde2d4ea52318a3590af78b2a21162/gistfile1.txt)
It will output any of the modules that @azer unpublished yesterday (that are being used by your project).I made the assumption on the top post that they were in-house proprietary software given the reference to keeping everything in git.