What happens when packages go bad?
jakearchibald.com
jakearchibald.com
This way anyone who decides to trust the new maintainer will be able to act as a "canary in the coal mine", notifying others if they run into issues. This also delays the gratification to the new maintainer. If they're truly malicious they'll need to spend maybe months / years maintaining the code / fixing bugs until they'll be able to hit pay dirt. I think most malicious devs will not want to pay this price.
EDIT: This would also act as a window for (a) folks to find other alternatives for their projects, and (b) inspire folks to build alternative options.
I think the willingness to fork over access to widely used packages isn't just a reflection on your desire to move on from the project, but it also reflects your blatant disregard of the thousands, maybe millions, of people who depend on what you've built.
"I have not looked at the readme...". Maybe this will create a market for a new type of project. The one that lets folks know the status of the packages they use in their project.
There's properties you can set in the package.json to indicate that a package is deprecated and/or has moved.
Why do you say that? The point is that you do not transfer it.
So if you wanted to stop maintaining your very own `right_pad`, you could mark it as deprecated so devs installing it get a warning in the CLI, and I could then publish my own `@klathmon/right_pad` and if you want you could endorse it in your readme if you wanted.
If I mark my package as deprecated and I don't point to yours, then malicious actors flood with copycat namespaces, and you have @klathmon/right_pad, @danshumway/right_pad, @linus/right_pad, and even the occasional phished @k1athmon/right_pad. Would that extra confusion be enough to trigger an audit for a company that wasn't planning to audit the original dependency anyway? Would an overworked engineer have the presence of mind to double check that their version has the right prefix?
There are, I'm sure, people who think it's OK to upgrade a package without worrying about the security implications, but not OK to switch to another package when the original is marked as deprecated. But I don't know that those people are in the majority. Certainly, anyone who's not already using a lock file and freezing their dependencies is probably not gonna think too hard about this.
(Somewhat to my shame) I can think of several instances where I personally have found a repo marked as deprecated on Github, went to the repo that it pointed me at, and started using it without even checking to see who the new author was.
Maybe I'm atypical with that behavior?
Then over time if it's an important enough package there will probably be discussions or blog posts about the new maintainer and what a "fantastic" or "terrible" job they're now doing.
Maybe standard libraries need to get bigger for JS.
Also since the code has to be downloaded (or uploaded to small lambdas, for example), JS developers are generally weary of depending on large libraries.
Isn't a dependency tree of thousands of small, interdependent libraries essentially just a large, distributed library?
Is there really a benefit to that versus compiling those dependencies to a single file, especially given how many NPM packages are just a single small function? The end result in many cases would probably be smaller than a lot of old JQuery plugins.
Yes and it's a thousand times worse than a single, large library.
While I do think making things easier for developers is a good goal, making the wrong things too easy (that is, making libraries that become very heavily depended on) results in the problems of Node.js (and I would argue the same is true for Rust). Maybe I'm just a curmudgeon, but I really have a sour taste in my mouth when I look at how Node.js and Rust do library management -- it's making it too easy to make a small library which doesn't really work and then people depend on it.
I once tried to write a simple IMAP cloning utility in Rust and found 4 IMAP crates -- none of them worked properly and all of them incorrectly handled several core parts of the IMAPv4 spec. I figured out later that they all appeared to be forks of one another (or they copied each others' APIs), but that doesn't help matters -- why are forked crates taking up more space in the global crate namespace? The "imap" crate doesn't implement IMAP properly!
Python had this problem (in a lesser degree) too with pip, but I think having a larger standard library and lots of time to mature allowed them to overcome it.
I noticed this problem with Perl and CPAN: there were dozens (well, several) email-sending packages, none of which were near complete. (Speaking SMTP is easy, an actual MTA is hard.)
I suspect there are similar issues in Java-land, although I haven't used Maven-derivatives enough to find out.
It's the wave of the future: make it really easy to share and use libraries, and get a crap-ton of low-quality libraries.
Another factor that probably helps is that most operating systems are built off C derivatives and thus usually are already carrying common libararies as dlls
How so?
It's a failing of the Unix model with a small silver lining in that C projects tend to go out of their way not to have to make the build any harder than it already is.
>> I think the answer is everything in your standard library plus some of the stuff you would use 3rd party dependencies for in your own project.
> Even in C with its tiny standard library you don’t see this kind of explosion.
I think one of the largest factors is that client-side JS doesn't have a good solution to the problem of dead code elimination.
There are solutions like Google Closure (the JS-to-JS complier), but it's difficult to set up and not many people use it. Instead, it seems like people have moved to lots of small dependencies so they can essentially do dead code elimination by hand.
E.g. imagine C++'s Boost libraries: That's I think ~150 libraries with dependencies among each other? Now imagine there was no need to publish them together, so each of them is an individual library that could pick any random dependency, not just primarily from inside the set of 150. And split some of them up in 4-10 pieces doing parts of their task. And external projects of course also then depend only one some subset of those parts.
A relative lack of a standard library maybe seeded the principle that people go looking for libraries for small things.
The problem is that something this basic is considered a complete library. If you know regex you could write this on your coffee break. If you don't, you could look it up and still write in on your coffee break. If you want trim and do anything else with a string, you need another dependency tree.
This is not even really a problem with javascript lacking a standard library, so much as a problem with the accepted practices of the javascript community.
When the trim library was needed, you likely also needed to support IE8.
But IE8 has a bug where \s doesn't match a non-breaking whitespace \u00a0 - http://www.nivas.hr/blog/2012/01/19/non-breaking-white-space...
And you need to decide whether you polyfill or ponyfill the function.
Yes you can theoretically do a trivial function in your break. With real browsers in practice you will often find your trivial function just isn't trivial.
But the string trim package mentioned above is that trivial.
But in the end JavaScript needs a much better standard library.
There are obviously upsides to having a large standard library, but there are also significant downsides: because of Javascript's position on the web, it is very hard to remove things from the language when we get them wrong, and because Javascript is so widely used, it is very difficult to know in advance which implementations and styles of coding should be preferred.
So for example, I'm happy that we have native Promises now, but I'm also happy that we waited until basically the entirety of external library authors had settled on an interface, and I'm even happier that we decided that deferreds were unnecessary since they could be easily recreated with normal promises.
Of course there are downsides to preferring external libraries instead of native ones, but in Javascript this is a calculated choice. And I agree that JS culture encourages developers to take this too far. Even with a tiny standard library you probably don't need a leftpad package.
It's just that those downsides look a lot worse, because in the JS community we lack good security practices about freezing packages. We also assume that NPM is secure by default, instead of a fancy wrapper around Github. We don't have a way to make sites immutable. And we have really stinking awful sandboxing in NodeJS, and (arguably) insufficient sandboxing in the browser.
I typically get a little bit of pushback on this, but I advise people who are building end-user applications or a website as opposed to a library to commit their dependencies to Git. I also advise enterprise developers to avoid using packages that haven't been audited unless they're willing to read through the source code themselves -- and if you find an audited package, download that version and check the hash. That one in particular is a hard sell, I've heard developers tell me that it's literally impossible for companies to audit their Javascript libraries. I find that mindset really, well, disappointing -- especially since those same companies seem to have no problem blaming NPM for not auditing everyone's libraries.
On the browser side, I advise people to self-host packages instead of using commercial CDNs, or at least to use subresource integrity policies[0]. That doesn't protect you from a compromised server, but it does protect you from a good number of XSS attacks.
This is annoying of course. In an environment like Linux, on the server you'd use CentOS and you'd have strong guarantees about stability because you wouldn't be installing random packages from Github. On Linux you also get better checksums around your packages. Linux sandboxing isn't very good, but it's not like Node's is any better. I think on the web, we're really behind on this stuff.
But I don't think expanding the standard library helps with any of that. I think it's a band-aid fix.
[0]: https://developer.mozilla.org/en-US/docs/Web/Security/Subres...
It also has 1103 total dependencies.
Because there's almost no friction to add a new dependency, it's often easier to add one simple thing than roll your own.
Note that nobody uses 1700 dependencies directly. Your project might use 5 deps for the 5 things it needs, but each of these libraries will use a few libraries for smaller pieces of functionality their library is composed of, and so on.
Wow. Small attack surface with 1700 dependencies. What a hard concept to swallow.
Image manipulation on the client.
There's no reason to commit your dependencies if your intent is security. Just use the lock file. The packages won't change on npm and you'll always get the same version for all sub-dependencies.
The ones who do are focused on caching private npm servers like verdaccio.
Personally I do not understand why Java libraries which are binary blobs are not targeted more often.
Java libraries are JAR -> Java Archives. You can unzip them and there is JVM bytecode inside which can be decompiled. Not really binary blobs.
Until they decide to stop hosting old versions you need, because reasons.
Though I agree, over reliance on just npm can be a bad thing (the 2016 issue over kik)
You actually see the diff on dependencies on updates. Vendoring is done in the Go world. Together with `yarn autoclean` it can be reasonable.
Perhaps, but building from a controlled body of code fully under your own control also has some big advantages over the typical package management and development culture in the JS world, which I think it's fair to say has been a pretty good representation of how not to do robust, professional software development for years.
This article is about the latest high-profile mess with event-stream, but it follows years of interrupted work because NPM was down again, several generations of tools that couldn't meet basic requirements like reproducible builds, excessive dependence on near-trivial packages like left-pad, and of course the overall problem that managing thousands of such dependencies is basically an impossible problem and inevitably leads to problems with trust, reliability, licensing and legal matters, and so on.
Distribution maintainers, developers and system engineers have to either unvendorize everything or go through all vendorized dependencies every time a vulnerability comes out and patch them.
Most live systems cannot receive a full upgrade every time a patch is published.
Use dynamic linking on languages that support it and don't vendorize things.
Even better, store them yourself: cache, mirror, zips in the repo, whatever.
This won’t protect you from an initial hack, but it means that if the dependency was good when you got it, it will stay good.
Vendoring dependencies on a Node project doesn't make it easier or harder for me to update those dependencies later. I run the same commands from the command line either way. It also doesn't make it harder for me to update dependencies of dependencies either -- quite the opposite; I wouldn't advise people to do this, but in an emergency if my dependencies are being tracked via source control rather than fresh-downloaded on every install, I can make security tweaks or change their dependencies, and those changes will actually stick between installs.
Even on a website, vendoring my dependencies doesn't mean my dependencies need to be a single blob -- I can still have a dependency deployed as a separately updatable file.
Is there a specific scenario you're thinking of where vendoring a JS dependency would prevent someone from updating it?
It just doesn't happen without your knowledge and consent (unlike the default NPM settings).