NPM Vulnerability Discussion on Twitter
solipsys.co.uk
solipsys.co.uk
If you prefer Twitter's rendering then just click on a node and that tweet will open for you.
Edited to fix a typo ... thank you Jonn. Oh how I hate auto-corrupt.
See top right: https://upload.wikimedia.org/wikipedia/commons/d/d8/Trn_cons...
If you find the trn/slashdot/HN-style threading better than the 2D graphical layout then I doubt I can explain to you why I find it (the threading) lacking.
What kind of curve is this?
All that shows up on an iPad at normal zoom is a white page. I waited quite a while thinking maybe it was just generating something client side and that was taking a while before I noticed that there were scroll bars (both vertical and horizontal) on the page and their thumb size indicated I was only seeing a tiny fraction of the page.
On my desktop it depends on what I was doing previously. About half the time my browser window won't be big enough for anything to show up in the initial view.
Thanks.
Thanks.
The white background seems to come from this:
<polygon fill="white" stroke="transparent" points="-4,4 -4,-11651.5 3898,-11651.5 3898,4 -4,4"/>
We just need to find that in the SVG, look at the points defining the rectangle, find the left side and the top side from that list of points, then add some text. The left side is the minimum first coordinate in the set of points and the top is the minimum second coordinate in the set of points. That's -4 and -11651.5 in this case.Then we just need to add text there:
<text text-anchor="left" x="-4" y="-11637.5" font-family="Times,serif" font-size="14.00">Hello, World!</text>
Using the left side for x seems fine. For y using the top results in clipped text because apparently the y for text is the baseline. I don't know how font sizes work in SVG, but it looks like moving the base line down by font-size works nicely.Here's a little Perl script that can be used as a filter that takes your SVG and adds that text line, finding the top and left by looking at the points list in the first polygon line that has fill set to "white" and stroke set to "transparent".
#!/usr/bin/env perl
use strict;
my $left = 0;
my $top = 0;
while (<>) {
print;
if (m{<polygon.*fill="white"} && m{stroke="transparent"} && m{points="(.*?)"}) {
my @points = split /\s+/, $1;
foreach (@points) {
my($x, $y) = split /,/;
$left = $x if $x < $left;
$top = $y if $y < $top;
}
$top += 14;
print qq{<text text-anchor="left" x="$left" y="$top" font-family="Times,serif" font-size="14.00">Hello, World!</text>\n};
last;
}
}
print while (<>);So I need to test whether they will overlap, and suddenly it's a bag-o-nails. Or at the least, need to test if something is up there and not both including the extra, but that's ... inelegant.
The option of having an HTML template and including the SVG will work, and that's the one I'm probably going to go with. It needs tweaking.
(Probably need slightly more than font-size, in case parts of some characters descend below the baseline).
Thanks.
Implementing your good code with no dependencies is heresy! It's a relic of the past.
</rant>.
On a serious note, you can implement a lot of things with a standard library of any language, and only time consuming machinery should be incorporated as external dependencies, but somebody didn't get the memo, I guess.
* Legacy Projects that pulled them in way back when and never migrated to the native API's
* Inexperienced developers not knowing this is part of standard JS now.
You can't really do much about the former. For the latter though, NPM would be wise to plaster a link to MDN's docs on some of these packages.
That's one line of code in your utils module...
For isEven, you first need isNumber. And that’s where the complexity lies.
({}) % 2 == 0 // false
([]) % 2 == 0 // true
"" % 2 == 0 // true
"a" % 2 == 0 // false
"1" % 2 == 0 // false
"2" % 2 == 0 // true
undefined % 2 == 0 // false
null % 2 == 0 // trueI am literally not joking. https://www.npmjs.com/package/is-even
I don't know what to say.
/**
* is-odd-and-even
* Github: https://github.com/fabrisdev/is-odd-and-even
*/
import isOdd from 'is-odd'
import isEven from 'is-even'
/**
* @param {number | string} i The number to check if it's odd and even
* @returns {boolean} True if the number is odd and even, false otherwise
*/
export default function isOddAndEven(i){
return isOdd(i) && isEven(i)
}Your isEven function should check that it's operating on a number. Eg "const isEven = (i) => typeof i === 'number' && i % 2===0;" If it's not a number then calling isEven on it doesn't make sense. If the user wants to check if a string is even then they can convert it before making the call.
A proper solution would be along the lines of:
const isRealNumber = (n) => {
if (typeof n !== 'number') return false;
if (isNaN(n)) return false;
if (n === +Inf || n === -Inf) return false;
return true;
};
const isEven = (n) => {
if (!isNumber(n)) throw new Error("Numerical error: Invalid input");
return n % 2 === 0;
}
const isOdd = (n) => {
if (!isNumber(n)) throw new Error("Numerical error: Invalid input");
return n % 2 !== 0;
}
This is the additional complexity created by dynamically weakly typed languages. And that’s why we get all these BS tiny npm packages.'isEven' really shouldn't be a dependency though, either through npm or internally. 'x % 2 === 0' works fine for integers and can be inlined, or if you're using it a bunch and want slightly cleaner code, you can define it as a lambda alongside the code that uses it. Then everyone can see exactly what it's doing with the full context.
The real issue is overabstraction. Even if you do everything "right" and have it as an internal dependency with all the correct error handling, it still makes the code a pain in the ass to read.
If the goal of isEven is to handle user input, it is woefully misnamed.
I have a Rust project I've been working on.
It has dependencies for:
axum: the web framework in use
serde, serde_json: Serialization of API responses
sqlx: To talk to sqlite
time: Standard datetime format with localisation and serde support
tera: templating
tracing: async logging
uuid: UUID generation
Of these, I feel they're reasonable dependencies. You could maybe quibble about uuid (which itself only depends on rng + serde with the features I've enabled), or tracing maybe, but they provide clear value.
Anyway, those direct dependencies expand out to 300-odd transitive dependencies.
So what do I do? Do I write a templating engine, HTTP server implementation, SQLite driver and JSON library from scratch to build my little ebook manager?
If Software Engineers were more like "real" engineers (ie PEs) they could never sign off on a project built like this.
$ npx create-react-app my-app
$ find my-app/node_modules -name package.json | wc -l
1506The actual runtime dependencies of a react app are basically just react and react-dom.
Aside from that, I guess your other dependencies have the same problem. It's not enough for one person to be mindful if they need something fully reasonable like a web server library which then depends on a million packages to build from source. Often-used dependencies could
- work like a regular package system (e.g. Debian's) and distribute (reproducible) binaries unless explicitly asked to build from source,
- only pull optional dependencies when they're used (a web server might depend on a logrotate dependency but maybe you don't use on-disk logs at all)
- and/or be more selective in what they depend on.
None of these are quick and easy to do without downsides
Source: https://twitter.com/sephr/status/1524080664106086400
I don't fully understand why packages like this are so popular. My guess is it has to be a combo of legacy projects pulling those deps in and inexperienced developers not knowing that these features are now part of standard JS.
While this would cause controversy, I think that NPM should lock all these dependencies down and not allow any modifications. The modules would basically be passthroughs to native ES6 functionality. I'm not saying NPM should lock down all legacy packages, just the ones that implemented ES6 functionality (isBetween, isEven, isOdd, forEach, etc)
Compared to that, left-pad is rocket science
An entire module for just `return value !== value;`
(Not supporting a culture of taking a dependency for this sort of thing, though)
isNaN Number.isNaN
----- ------------
{} true false
[true] true false
undefined true false
"{}" true falseTL;DR "SEO" spam to market ones own nodejs expertise, would not be out of the question
I consider 'iseven' and 'isodd' to be signs that Javascript is a hellaciously engineered piece of crap that should be avoided at all costs. They're popular because Javascript is garbage.
let isEven = theNumber % 2 == 0;
There are reasons to dislike JavaScript, but this isn’t it. The native solution is about as standard as it gets.It’s like doing a for(i=0; i<items.length:i++) { e=items[i] } to iterate over each element instead of something like foreach, or a lambda.
How is that for loop not expressing what it means? for is used to iterate over something. There are legitimate reasons to choose for even when forEach is available.
Does this not express what I mean?
if (amount > balance)
return “insufficient_funds”;
Chances are, I’ll put this in a method anyway, and that’s what a language enables me to do - to create meaning that is contextual to my program.Abstraction can help simplify complex problems, and that’s what it is good at. But when literally everything is abstracted away, it becomes harder to reason about what the program is doing without also learning about what the particular abstraction does (e.g. implicitly handling negative numbers, etc).
A well implemented standard library can be a joy to use, and some people choose languages for this reason. Some languages very intentionally provide a very lightweight standard library, and that’s just fine. All of this shouldn’t absolve the developer from understanding some of the very basic constructs of the language they choose - a lack of which will lead to bugs and vulnerabilities.
If I choose JavaScript (or it’s chosen for me), it behooves me to understand the basics vs. farming this out to some node package that just made my app harder to reason about and vulnerable to supply chain attacks for essentially no value.
jbverschoor said (edited for brevity):
> Yes. But you’re not expressing what you mean.
> It’s like doing a
> for(i=0; i<items.length:i++)
> { e=items[i] }
> to iterate over each element instead of something like foreach ...haswell replied:
> How is that for loop not expressing what it means?
It's true that the for loop is accomplishing the goal, but it is specifically running through the elements from 0 to length.
But sometimes what you want to do is, for example, apply a function to every element, and the fact that they are in a structure indexed from 0 to N-1 is not relevant. Your intent is to do a thing for every element, not to run down a list doing the thing.
There is a difference in intent, and using foreach instead of a for loop quite specifically expresses that.
Small examples like this are never convincing, but there really is a difference between "Run down this list" and "do this thing for every element". The output is the same, but the intent is different, and expressing that intent is important in larger projects.
It may be true that running from 0…length is not the most important part, but it may also be true that I need more flexibility and control while iterating, and may still choose to use for as a result.
But I think we’re getting a bit off track (and I helped get us here…).
jbverschoor‘s comment was a followup about the expressiveness of the modulo operator vs built-in convenience methods.
At least in the case of JavaScript, there isn’t an isEven available. If it was, an argument could be made that using it is more expressive. Since it doesn’t, the argument can be made that the most expressive option is the idiomatic one. Certain code fragments become idiomatic for a reason.
I interpreted the comments about for and the perceived lack of expressiveness in using it against this backdrop, and my point was just that for still has a time and a place. It may be that jbverschoor recognizes this, but I think the comparison doesn’t quite help when examining the use of modulo.
Going back to the original, there is a difference between "isEven(n)" and "0==n%2". Sometimes (and this is definitely true in some code I write) there is a conceptual difference between asking if something is even, versus asking if it leaves a remainder of 0 upon division by 2. Yes, they can be proven to be equivalent, but the underlying intent is different.
BTW, I agree entirely that in these cases it's really mostly pointless, but the principle carries over to larger examples. Trying to see that there is a difference is probably worth it, even if we agree that it doesn't really matter in these trivial cases.
Oh, and for a compiler I used to use it really did matter that you used "foreach" when you could, versus using "for ()", because it could dispatch things in parallel in the former case, and couldn't in the latter. So expressing the intent in the code itself and not just in the context really can matter.
The only nit I'll pick is re: the 2nd paragraph. I do agree that there is a difference in the expressions "isEven(n)" and "0==n%2" in isolation. But these expressions will always be surrounded by more code, and the degree of difference will entirely depend on that code and its structure. All of this can be very simply solved with a quick:
function isEven(n) { return 0==n%2 }
And let isEven = 0 == n % 2 provides about as much context as isEven(n).Admittedly, I had to be conscious of how I wrote that expression, but choosing meaningful variable names is also just part of writing code, and isEven() by itself likely doesn't provide enough context about why I'm checking for even-ness to begin with. These considerations all need to go into the design of the class/method, but at no point can the developer wash their hands of the need to provide context through standard practices, just because the standard library provided more descriptive methods.
And I'm not saying this is what you're saying, but the original comment seemed to believe that more expansive standard libraries are somehow inherently better. I'd argue they're just different, and more factors to consider when choosing a language. Sometimes they help. Sometimes they're not worth the price of admission.
I think the discussion could be generalized as: how far should standard libraries go? No matter how good that standard library is, at some point, I must apply the fundamental skill of adding context and structure to the code I write.
However, some variation on n%2==0 is idiomatic in a large number of languages, both modern and old [0].
This includes Lua, Java, Python, Scala, Perl, PHP, Swift, and many others.
It is true that there are numerous languages with built in convenience methods in their standard libraries, but that does not by itself mean that n%2==0 is problematic and certainly this idiom is deeply entrenched.
There is nothing preventing a developer using those languages from wrapping this in a function, and indeed this is probably a good idea if it’s something you find yourself using repeatedly throughout your code base, just like any other often-repeated code snippets.
Because I firmly believe that crap like this happens because the language itself is very, very dumb.
It’s a much better choice to just copy the source code and possibly license into your own code to just eliminate the overhead.
It’s not like any of these packages depend on being updated for security reasons or anything.
He’s a prolific open source author that traditionally had a lot of modularization in his packages—though I think he has started to move away from it recently. He talks about it here: https://blog.sindresorhus.com/small-focused-modules-9238d977...
Most projects will include at least one package by him in the dependency tree.
I vastly prefer to use libraries with 0 dependencies but I never quite manage it, so I end up with the same problems as everyone else.
It's an unpopular opinion for some silly reason. It's like when Guava or Apache Commons was included on every single Java project in the 2010s. "It's just so much easier to use a well-tested library". That line of thinking is what got us here.
The individual functions are installable separately from NPM, and lodash-es should tree-shake quite well, but I do know what you mean about it dragging in its internal dependencies so you end up with a 15kb of lodash code for a single thing. I probably wouldn't love to use it client-side on an ordinary website, were I to make one.
This might be controversial, but I think it has to do with less experienced developers, beginners, who doesn't know how to find out via "pure code" if a number is even or not, or even if something is a number (IIRC, there is an isNumber package out there as well).
I see this in other languages as well, but not in a "JavaScript scale", but that could probably just mean that the language's availability and popularity decides the number of "stupid packages."
I haven't worked with anyone for years that doesn't start troubleshooting an issue by reaching for a random dependency. They just don't even bother learning Javascript or the Web APIs or CSS anymore. Without a solid base of fundamentals every trivial problem seems insurmountable so they just don't even try.
This minority makes money out of their popularity in the NPM ecosystem. Having 1000 NPM packages is better for your reputation than having 100. And having all 1000 packages with lots of weekly downloads is better than having just a portion of that.
But how do they achieve that? Well, they have 1000 NPM packages, so each one depends on 5 to 10, that then depends on a few handful more. You have packages for checking if an HTTP status is a certain number, you have packages that have colors as constants, you have is-even, is-odd and so on. All that exists to maintain that closed ecosystem.
So out of the 1000 they basically have 20 useful tools and 980 garbage packages that exist only to maintain their own ecosystem.
Most people isn't using is-even or is-odd directly. They imported some other packages that are quite useful, but often need 10-20 sub-dependencies. Another interesting thing is that those shitty packages aren't really that important in applications. They're often used in build tools, CI and testing, tools for making CLI tools, and the sort.
The crazy thing is that a lot of people using is-even/is-odd aren't really "noobs": they're probably experienced developers that said "fuck it, I'll use some random tool from the web" when facing some random a problem.
That said, it still leaves a sour taste as this effectively implies that a certain set of JS developers is very happy to abuse their (maybe initially rightfully earned) prestige to gain even more prestige while leaving behind a mess for the whole ecosystem. I don't understand why this is tolerated. The Node community needs to have a serious discussion about why certain packages are allowed to spread garbage, create forks of the relevant packages that rip out "is-even" etc. and then eventually converge to these forks. But to this day, I don't see the community taking this problem seriously enough.
Now, supply chain attacks and "too many dependencies" are a potential issue for every language with dependency management (see also log4j, etc.), but no other ecosystem seems to be have such a high frequency of issues and (widely used) "is-even" packages are simply not a thing in any other mainstream language (some languages, like Swift, include similar functionality in the standard library, which is totally fair).
For every person calling it out like we're doing here, there are ten others praising maintainers able to whip ten semi-useless packages per week.
It's not just random maintainers making small packages. The core infrastructure of Javascript is in it. Babel is made of hundreds of packages, which all live on the same repository (because of course the maintainers don't want the hassle of maintaining multiple things). Some of those packages don't even have anything of importance in it, just metadata, a couple flags and some boilerplate [1]. The package is just a way of organizing code. Webpack, ESLint and others aren't exactly better.
EDIT: And of course I got downvoted :)
[1] https://github.com/babel/babel/blob/main/packages/babel-plug...
> The core infrastructure of Javascript is in it. Babel[...]
Babel is not core JS infrastructure. It may be close to fundamental to the modern NodeJS development experience, but JS exists happily (and capably) without any of that stuff (including package.json, for that matter).
I don't really despise anyone in Babel, though, I'm only criticising their packaging method. Babel isn't doing the million-packages thing to gain popularity.
IMO, that's even worse, because that means that a lot of people are using stupid and vulnerable packages without knowing it.
I wish there was some sort of better control over the NPM directory, where someone could block/downvote (or whatever) packages that doesn't deserve to live. How this would - or should - work in practice, I have no idea, but it's just getting scarier by the day to import a package in your application.
If we ban those, we'd have to ban Babel and Webpack too... Oh now, wait a minute, now that actually sounds interesting...
What if we focused on fixing all the problems, and not just retreating thinking that "we can't solve this problem, because there are so many other problems related to it"?
Your thinking is literally the definition of the problem.
Should we treat them as spammers and polluters, then? Because if what you describe is true, that deserves to be called out and mowed to the ground.
If incentives were aligned differently, different results might have resulted. Probably with different externalities (or unintended consequences).
[0]: https://en.wikipedia.org/wiki/Tragedy_of_the_commons?wprov=s...
It is more akin to SEO spamming than black market spamming, though. They're polluting NPM in the same way SEO farms spam Google. It makes life difficult for everyone, but it's still a gray area in terms of legitimacy. Which is why nobody really talks about it.
Maybe you could argue other for other reasons to use these packages.
1. search for a package that does X
2. scan the search result and find the package that really does X
3. learn to use the package's API
4. import the package in the code
5. use it
Isn't it much easier to just copy-paste code from stackoverflow? There's also a good chance that the you can get some very good explanation and interesting discussion around the implementation there.
There's also Github Copilot, which pretty much replaced almost my entire usage of Stack Overflow.
But the thing is that with a rando package is that one doesn't have to review the code. Sure, the code is also coming from somewhere else. Sure, it might never get updated. Sure it might be full of bugs. Sure, it might be more dangerous than copying from StackOverflow. But out of sight, out of mind.
In the end the overuse of packages isn't about saving time or "doing the best for business" or "ensuring that the code is maintained by someone else". It's purely about covering our asses.
Most times I get the urge to pull in a rando package I find I really only need a few things it does. I check it out. Read it. Write my own.
I almost never need "all the things" outside the situations where I am using a big framework. So reading it, getting inspired / ideas from someone who did the thing and then I write a much more narrow focused version for myself.
The thinking was as follows: Of course you could just copy the code, but then that increases LOC in my codebase that I'm responsible for. More code is more work. Lines in a dependency are the responsibility of someone else. If there's a bug, even in a small function, the community can identify it and fix it. I can get new features I might not have known I needed. I can benefit from all of these fixes indefinitely into the future without ever having to have any mental overhead about that code. So can everyone else; it's good to maximize code reuse.
I don't think I've ever used something that could be an obvious one-liner like `isOdd` but for lots of only slightly more complex stuff like left-pad, email format validation, GPS coordinate math functions -- all stuff that's really less than 30 lines -- it was really nice to just not have to think about the implementation details of that and get back to solving your problem. I could have reviewed the code or written it myself but it's just more work when remaining at a high level `leftPad()` call let's me stay focused on my original task.
That said, I've since realized I was wrong of course. Trying to maintain projects that haven't been touched in more than a year led to hours of fixing dependency issues. We switched to using dependabot, which is better, but just makes it obvious how much work it actually is to keep dependencies up to date week-to-week. Then there's all of the security issues. These days, for small packages, I advocate for reviewing the code from these packages, ensuring we understand it, and then copying it in directly with a comment for attribution. We generally try to keep dependencies low; still more than in other languages but at least some thoughtfulness about whether it's "worth it". I think a lot of the community has shifted similarly, but there's still a lot of older projects with older dependencies.
Once your company is owned by a supply chain attack or by an RCE in one of your dependencies, you will learn that you are in-fact very much responsible for the code in your external dependencies.
This can be complicated. Depending on another package is usually very safe, at least as safe as "dynamic linking", but including code in your own source tree needs licenses to be compatible. Even then, you might have to change your license to "BSD-3-Clause + ISC" or similar composite and you will get complaints from users.
It actually works like this: Author X develops `iseven`, `isodd`, etc. No one really downloads such packages. Author X then develops `importantPackage` which does do something useful developers out here download. The thing is `importantPackage` relies on `iseven` and `isodd`. Now `iseven`, `isodd` are downloaded alongside `importantPackage`. Profit.
My point is, we should recognize certain NPM authors as toxic, but I guess "freedom of speech/code" stops us from doing so. Example of such an author: https://github.com/jonschlinkert/
The layout is actually not that much dissimilar from what twitter already has, just more info from non-main threads and better visibility for orphan leafs.
The chart in all forms (there's more than one) shows each tweet as a node, and arrows showing which tweets are replies, and which tweets are "Quote-Tweets".
If A->B, then B is a reply to A.
In this particular version I've taken the longest thread and laid it out on the left, allowing the other descendants to flow to the right. Another layout is to have the initial tweet at the top and lay it out as a simple tree (technically DiGraph), but that often leaves the top left corner empty, leading people who initially open it to think it's empty and close it without scrolling. Yes, they do, that's happened before here on HN, including in this discussion.
Does that help? Are you still confused? It's just a chart showing tweets, quotes, and replies.
I have several tools that do this sort of thing, including a 'bot that can be invoked automatically on mastodon. I don't do that on Twitter because popular discussions get seriously out of hand ... there was one that I stopped tracing after if got to 3500 tweets. I have some heuristics in mind to help control that, but it's all still very experimental.
I also have a pig-ugly, pre-alpha, bug-ridden discussion system that works directly on a DiGraph, but when it was submitted to HN sometime ago reactions were ... (significant pause) ... mixed.
I'm quite sure that the market for this is really big. Does something like this already exist?
EDIT: Actually, there are package managers + package collections that intentionally can be installed on top of an existing OS; nix/nixpkgs and pkgsrc are probably the big ones. I'm not 100% sure that those are what you want today, but it's likely that if you could get maintainers to add your desired packages they'd be accepted there.
Perhaps something akin to the UK OFCOM amateur radio license requirement would work, requiring people to update/confirm their details at least once every five years, or in this case, yearly?
How do people manage selling or changing a domain they've been using long-term as their primary email address for sign-ins like this? To stop a new owner of that domain from taking over accounts sign-ins that use that email address, you'd need to reliably update all your sign-in details with anything you've signed up to?
You could use a password manager to track everything you've signed into before to help with this but that's error-prone.
Alternatively, this forces you to keep paying for the old domain indefinitely?
What dev thinks oh I can’t upgrade because of this error, stackoverflow says use this flag —disable-signature-verification so I do and now I can develop again
At least that way upgrades, malicious or otherwise are opt in for production
The other thing would be ideally a crowd funded resource to vet particular versions of popular packages.
Try scrolling. Or searching. Or scaling.
There's a chart there, and you can scan the entire thread, or just click on a node to go directly the the tweet you want to see.
When ‘node-ipc’ overwrote all files on disk, NPM just waited while the author himself published an amendment and posted that the package ‘only created a text file, no biggie’.
I don't know what the solution is. I don't think there is a solution. We can't have 5000 dependencies from 2000 random individuals on the internet and still be safe. But if you want to avoid that situation, you're locking yourself out of the vast majority of the NPM ecosystem.
Losing access to an account at a major mail provider seems more probable than losing control of a domain you paid for.
If I couldn’t afford that anymore, I’d probably go back to Gmail. I’d probably think to switch over many of my email addresses (for services that I use), but some may slip the cracks when the domain expires.
< 10% had useful 2FA enabled. Most were just password reset questions or had a backup email to some old expired and re-registerable forgotten earthlink accounts etc.
I do this stuff all the time. I looked up the password reset questions controlling the zoom.us domain 2 years ago and collected an insulting $200 bug bounty for it. Zoom, like NPM, didn't do signed binaries, so this would have been brutal combined with other issues I found.
The solution is what all sane OS package managers do: code signing.
NPM has rejected this at least as far back as 2013 when they refused a PR by someone that implemented it for them: https://github.com/npm/npm/pull/4016
I don't know what else to do but keep publicly trolling them at this point until code signing is implemented. Unsigned code is a free pass for remote code execution when an account gets taken over.
Meanwhile Debian maintains hundreds of signed nodejs packages proving it is very doable by a low budget team. NPM seems to just want to reduce deveopler friction at all costs :/
"Once you cross 100,000 weekly downloads, all new publishes must be signed" would keep the easy on-ramp for small packages but dramatically reduce the risk of major attacks against all the big targets (like the example here).
Not as good as signing everything, but a good start that can be iterated on later.
Most people would have no problem marking a couple of big firms and orgs as trusted reviewers and only showing packages to be used if signed by them.
For example: "ah yes, a core contributor, a frequent code reviewer, and two downstream consumers of this library have signed off on the changes in this release. even if one of them is having a distracted day/week, that's good enough for our team to be comfortable upgrading it today since it's a minor version upgrade"
You could be correct, maybe this would work in practice, trading on the reputations of those companies. It doesn't feel particularly open or community-oriented, though. Why not present a trust graph built from a broad set of worldwide users instead?
(one of the benefits to a trust graph would be the volume of signers; perhaps you wouldn't want to weight each signing equally -- again referring back to something like the distance metric mentioned previously -- but for even mid-popularity packages, the detailed review possible by careful users of a package could, I expect, be more reliable than automated-and-manual review by one or two large companies)
I don't think your trolling is doing you any favours anymore (sigh, that guy again). Kicking a dead horse etc. Maybe just let it go?
In parallel I am writing specs and proposals for widely applicable tooling and improvements.
For example: I go to release a new version and I've lost my private key, so I roll a new one -- this will happen often across npm's 1.3 million packages. Do I then ... log in with my email and update the private key on my account and go about my business? What process does npm use to make sure my new key is valid? Can a person with control over my email address fake that process? How are key rotations communicated to people updating packages -- as an almost-always-false-positive red flag, or not at all, or some useful amount in between? If you don't get this part of the design right -- and no one suggests how to in those threads -- then you're just doing hashes with worse UX. And the more you look at it, the more you might start to think (as the npm devs seem to) that npm account security is the linchpin of the whole thing rather than signing.
It's not just npm; that thread includes a PyPI core dev chipping in with the same view: "Lots of language repositories have implemented (a) [signing] and punted on (b) and (c) [some way to know which keys to trust] and essentially gained nothing. It's my belief that if npm does (a) without a solution for (b) and (c) they'll have gained nothing as well." It also has a link from a Homebrew issue thread deciding not to do signatures for the same reason -- they'd convey a false expectation without a solution for key verification.[2]
[1] https://github.com/node-forward/discussions/issues/29 [2] https://github.com/Homebrew/brew/pull/4120#issuecomment-4068...
It seems there should be some multi-factor process.
Developers need to register a password, an email address, and a YubiKey/TOTP token. If they lose access to the email address, they can log in to their account with the password and token. If they lose the token, they can be issued a new one with the email and password (or recovery codes).
As long as the account stays secure (i.e. an attacker doesn't manage to keylog the developer's NPM password and email password) then the NPM account can be trusted to add new package signing keys. The npm client then needs to trust metadata from NPM which vouches for these new keys.
The web of trust PGP signing approach works reasonably well to protect most linux servers in the world since the 90s. You can complain about it and say there should be a better UX toolchain for it, and I would agree with you. Thankfully the sequioa-pgp team has made huge progress here and it is a shame they are not getting due support for their heroic and near thankless efforts to make this better.
Still, even with todays GnuPG tools, abandoning pgp for supply chain integrity and replacing it with nothing is crazy. Imagine if we abandonded TLS because early implementations sucked. Use the best tools we have then fight to make them better. That's just good engineering.
The software eng commuity at large basically said "Look we just stopped signing code and nothing bad happened... oh wait bad things are happening. Too late to change now!"
This was a reasonably well solved problem, but entities like NPM will need to have the humility to admit that rejecting best effort cryptographic authorship attestation was a mistake.
I expect this to change. NPM will roll out mandatory MFA for the most-downloaded packages[0] (RubyGems as well[1]). I expect this will rise to a 100% requirement at some point because Github's decision to require MFA by the end of 2023 will massively raise the waterline of folks who have the capability to MFA and experience with MFA.
[0] https://github.blog/2021-11-15-githubs-commitment-to-npm-eco...
If you're trying to protect people, why release an exploit straight to the public instead of responsibly disclosing?
Edit: Looks like it's already known, hard to know that from the tweet context.
The more focus we can get the better - hopefully NPM will add MFA at some point...