Social engineering campaign targeting tech employees spreads through NPM malware
socket.dev
socket.dev
1. https://blog.phylum.io/sophisticated-ongoing-attack-discover...
2. https://github.blog/2023-07-18-security-alert-social-enginee...
1. Can you say how you find things like this? Just bulk scanning packages doing pattern matching heuristics?
2. I know it gets talked about every time, but I'd be interested in your perspective on how we (as an industry) could stop this kind of thing from happening. Is it as easy as blocking hooks from executing, or is there a better way forward? EDIT: Oh, just now seeing https://news.ycombinator.com/item?id=36869587 - so I'm assuming your answer is better sandboxing:)
NPM exposes that info in the _npmUser field: https://github.com/npm/registry/blob/master/docs/REGISTRY-AP.... That gives "name" (NPM username) and email.
While there are thousands of packages, I bet there's a much smaller number of publishers to worry about.
Like this guy pretty much has a separate NPM package for every single function that he writes: https://www.npmjs.com/~sindresorhus (1162 packages under his name alone, maybe even more in other namespaces).
Doesn't matter if you trust them; can you trust the 100+ dependencies they pull in? If it's a large package, you can be certain that they aren't verifying their deps.
Always.
It's a poor ecosystem, dominated by untrustworthy participants. There's no digital identity verification of participants or their packages. anyone can claim to be whoever they want to.
Pointing this out is not pearl clutching, and if it was easy to solve it would have been done so by now.
The only way to trust a package is to verify the authors identity and verify that that identity generated that package.
But tying packages to domain names means that all these low effort two line package authors have to put in non-minimal effort (acquire a domain, maintain the cert) which is not going to happen.
It works fine in practice; I've seen at least one big company do it. It doesn't happen in the public ecosystem only because people don't care enough.
> It's a poor ecosystem, dominated by untrustworthy participants. There's no digital identity verification of participants or their packages. anyone can claim to be whoever they want to.
Fixing that isn't hard. The Java/Maven ecosystem already enforces OpenPGP signing of all package uploads; NPM could do the same if they wanted.
I basically act as a middleman whose job is 'THIS IS A ONE LINE PIPE!'
If you're still using Lodash or Ramda or whatever it's because the standard library isn't quite as expressive as they are sometimes.
(Except Java....maybe.)
Does the library you pulled in to manipulate an array polyfill newer array prototype functions? If so it likely fixes issues in older browsers.
Does the library get tested by a much larger QA / QE user base than your ad hoc code? Importing a library might benefit you from not having to reinvent the wheel.
Does the library have unit tests and documentation? That is likely less work for your team to deal with bugs and maintenance.
In practical terms, I doubt I have ever written a line of code that nobody else in the world has written. In that sense, I know that my value to my company is to judiciously choose when to write code in-house and when to import. And the heuristic is way different than what you describe.
Secure
Performant
Non breaking
etc.
I would hesitate to import a package for a one liner (a left pad for example) but happier to import a decent bunch of battle tested utils (date-fns, underscore, etc.)
Since each dep is a political org you need to trust, so I wanna get a decent amount of problems solved with that dep not just save a line of code.
I’m very interested in ways to identify high quality+security libraries and orgs.
Simply saying “I am willing to rewrite everything from scratch” doesn’t work. Finding the best way to work with the existing ecosystem is more or less a requirement for most companies/projects.
My focus was on the types of poorly thought out package choices without the very care you are explaining. We would never reinvent rxjs, most tailwind plugins, or a key example of what you are explaining: Angulartics2.
But on the other hand, my example here is closer to some packages that have zero user base or are, in themselves, a reimplementation. The angular Material example is that we already have a design system with buttons that is heavily documented and tested, it would have been less work than dealing with the package to just read the documentation or ask questions.
The people in charge of each larger package can review contributions made to it and have the authority to reject code that includes attacks. If a larger package doesn't do a good job of preventing attacks, its reputation will decline and because there are fewer packages involved it is easy to remember these reputations.
I've had this debate before and, as far as dependencies go, sometimes you have to rip the bandaid off and audit your shit, and understand why exactly you outsourced something to a package.
I agree in the sense that smaller supply chain surface area is easier to secure/verify.
But it also makes the developers/maintainers of the few packages much brighter targets.
I’m more of a fan of standardizing how libraries can be rated, evaluated, trusted, marketed, etc. we need more efforts analogous to Consumer Reports / Underwriters Laboratories. So far there are a few for-profit companies in this space, including Google.
jQuery and Lodash could be considered "secondary stdlib" packages. Lodash for various data functions, and jQuery for dom manipulation functions. Now that other frameworks are more common and the basic stdlib is larger, many projects have migrated away from them to improve bundle size. (As an example, pulling in one lodash function tends to pull in lots of other functions from within lodash, unnecessarily inflating the amount of code sent to the browser.)
For the front end, ideally tree shaking is used, and even fancier build systems create code that doesn't pull down dependencies until they are actually needed at runtime.
I created a small Node+Express project, added a few dependencies to handle things like CORS, API Authentication, ORM, and data validation, and my node_modules folder has 300 packages already. Frontend wise, it's worse.
I only used a few libraries when I was doing Android, most of them utilities and vendor SDKs. And the contrast is even higher in another project I'm working on. The Java Code takes less than 10 minutes to build, while the frontend code turns the laptop into a jet.
The vast majority of node dependencies are for testing frameworks and other tools that are never shipped to the customer.
For example, if you install any tool that does pretty command line output, odds are it installs some console UI libraries as a dependency, and odds are they also install a lower level curses like library as a dependency.
(Also I've yet to meet a simple ORM system!)
Jest is another common offender, much like every other popular testing framework, it is large and complex, and due to the nature of Javascript, it is also more powerful than a lot of the other testing frameworks I have seen (e.g. Jest can do stuff at runtime that Java frameworks need to do at compile time).
Those Java frameworks that go in and mock classes are complicated, but they are all hidden behind a "single" dependency.
Regarding front end build times, if a frontend website is taking any amount of time to build, something else is wrong (e.g. TypeScript is rebuilding everything one very build). Typical frontend website dev stacks have hot reloading enabled that allows for the site to be refreshed and updated the instant a file is saved. Nicer setups live change the DOM and try to avoid reloading the page unless something touching state management has been altered.
(My experience is with React and Svelte alongside TS, Angular may be a different story, not sure!)
I'd take complicated and sane over the literal software illusions that many npm packages are.
> Regarding front end build times, if a frontend website is taking any amount of time to build, something else is wrong;
Not necessarily. Any production build should start from scratch (the dependencies may be cached).
oof... way to snatch defeat from the jaws of victory :(.
I generally prefer less magic over more magic.
So just to test I installed Django in a new virtualenv. I get a testing framework, and ORM, CORS, authentication, data validation, hot reloading and much more. It's three dependencies: django, asgiref and sqlparse.
Part of it is that Django just builds a ton of stuff on their own, but they also have the Python stdlib to build on.
As others mentioned you have things like Boost, the Go standard library, or even libc (of which there are multiple implementations). There's no reason why you couldn't have a JavaScript standard library, or multiple. They don't have to be shipped with the browsers, in fact I think that would be a bad idea. So when you start a new JavaScript project, you pull in your favorite standard library and when you ship your code, the bundler or whatever it's called could analyze your code pull out the relevant bits of the standard library and only include the bits you need.
That is kinda-sorta how things work, except JS devs like to split code up into smaller modules.
https://www.npmjs.com/package/jest?activeTab=dependencies has 4 dependencies, if you look at @jest/core a lot of its dependencies are in fact other parts of the same project.
The lack of batteries included means there is a lot more room for diversity, if someone wants to write a new server to compete with express (and plenty of people have) they can follow the same middleware API standard and you can swap our your existing web server for another one but your ORM will still work!
Also a minimum viable JS Rest endpoint is literally 4 or 5 lines of code. Getting cors working is 2 or 3 lines more (based on if you count the import statement as a LoC).
Adding schema validation to an endpoint is another 2 LoC.
The number of concepts needed to be understood to setup an endpoint in Node land is minimal, and you can then build out whatever tools you want on top.
If you want a batteries included framework, there are also plenty of those, from fullstack systems Sveltekit and Next.js to multiple starter templates for backend services that give you whatever tools you wanted.
People go "Python has Django" and there are huge advantages to having a single golden path, but Node has a richness in the ecosystem that lets people start off small with and choose what tooling is best as their projects grow.
And at the end of the day, an ORM is going to require N lines of code, it doesn't matter if it is bundled with a framework of it is an external dependency.
I wish more languages would invest in their stdlib.
With Go there is an owner at least.
But who is going to do it?
I guess it’s a tradeoff
- The standard base class libraries with .NET are quite comprehensive; I end up rarely adding "utility" libraries
- For non-standard libraries, Microsoft is a first party publisher of many packages on nuget; many of their first party frameworks and tools are published this way
- For third-party libraries, generally speaking, I end up using maybe a dozen or so very well-known and well-vetted libraries which have low dependencies of their own (again, because of the rich standard libraries)
- You end up with far less surface area for dependency chain attacks
According to the 2020 State of the Octoverse report [0]: - Node projects had nearly an order of magnitude more transitive dependencies than the next highest (683 to 70 for PHP) (p.11)
- npm had the highest percentage of critical and high severity advisories (p.14)
- 72.8% of JavaScript repos had dependabot alerts (well above the weighted average of 59%) compared to 6.5% for .NET and 40.2% for Java
Some would say that Microsoft's stewardship of .NET and C# is a negative; it can never be truly OSS. The flip side of that is that there is a professional team of paid software engineers -- some of the best in the world -- maintaining, patching, and securing the platform. While the third party ecosystem doesn't have as many options for doing X, the few options for doing X on .NET tend to be more well-rounded by the tighter community focus on a few, well-known libraries.Maybe some variation of an Apache Foundation like approach to curating a set of OSS NPM libraries and some mechanism of establishing trust would be beneficial to the NPM ecosystem?
Just my take.
Things are better now, but for all its (many!) positives, .NET is not an example of how to create a good ecosystem.
Of course back then Microsoft considered an ecosystem to be 3rd party partners selling libraries, of which quite a few successful companies did.
.NET itself is only 2 decades old (first release in 2001).
Examples: are you likely to write unit tests for your GetFirstChar? A specialized open source library that built this function is far more likely to have unit tests and documentation. Those unit tests and documentation come in handy during edge cases and when your coworker decides to extend the function with an `options` parameter.
Also, general developers rarely understand the full breadth of character sets, for example. So if your function just grabs the 1st byte from a string and expects it to be a character, your end user might be angry when only the first byte of a multi byte character is returned. Similarly the difference between a character and a glyph and a rune, etc. are important and lesser understood.
What is the difference?
"rune" is Go's name for a unicode code point.
It's not surprising that this is continuing to happen in JavaScript, the language whose userbase could actually be considered largely computer-illiterate.
That's not to say other languages don't have the same problems, but in my experience those who have experience with lower-level languages are more likely to have come from being power users and thus have a reasonable amount of computer literacy.
In contrast, the JS world is full of people who, lured by the promise of a well-paying easy job, "teach" themselves via the most shortcut-filled way and churn out a torrent of barely-working code that relies on others' barely-working code. It's pure quantity over quality and no one really understands what they're doing.
And academia has gone down the drain just as well, just look at how ridiculously insecure and underpaid most academic employment is, outside of the holy grail aka tenure track of course. Meanwhile, universities blow billions on fucking sports teams.
It's not a matter of capitalism, it's a matter of priority.
We at this forum are a tiny minority of people who have the time and mental capacity to learn science. Most people in the world are struggling to have quality food and a roof.
Although not as essential as housing, the software industry also faces such challenges. Sure, every IT shop would love to have computer scientists, but they need some things done and there wouldn't be enough scientists if every developer was waiting to accomplish a full CS degree and get 5 years of high quality CS experience.
Besides most people in the world are not struggling. Where do wealthy Americans get this idea that out there most people are poor and ignorant? Global poverty and vulnerability stands at around 15% of population.
but this app needs to ship now, and we need 3 more bodies to make that happen. in a perfect world they'd be from an actual tech school, but bootcampers who can spell C-R-U-D might be good enough.
I expected npm packages to somehow carry the social engineering, like carry a message somehow, but that’s not the case.
Anyone who understands the auditing role and provable chain of accountability FOSS OS packagers volunteer for... knows why it was necessary.
Malware will always show up, but it is arguably more important to be able to hold sources accountable. =)
Sure you can use a VM, a container, firejail (for an app), etc. but it all has quite a high barrier to entry & burden to bother with, so the vast majority don't, and the majority of the rest only do occasionally for something they already trust less.
https://github.com/phylum-dev/birdcage
It's baked into our CLI and supports limiting access to network, disk, etc. during package installation. For example, running something like
phylum npm install react
Will perform the installation in a way that disallows access to disk and network that aren't explicitly approved. This prevents malicious packages from grabbing and exfiltrating things like developer SSH keys during package install.It's probably manageable through direnv & some shell magic, but it'd be nice to have a built-in way of saying 'everything in this directory gets run through phylum [with these defaults] and has access to only this directory [by default]'.
If you have any questions/issues, feel free to shoot me a message. My email should be in my profile!
I really want to see those WhatsApp chats...
The argument can be made that the _important_ work was already done by GitHub and Phylum. This just appears to be a derivative work without any additional value add. That is, it functions mostly as an advertisement.
Doesn’t help that they clearly don’t even have all the malware packages in their system from this campaign, three weeks after the attacks…
At a total cost of $1B that would put scanning a single package at $16/package? Nonsense. Even at $1 in compute time on aws that feels outlandishly high…
You can’t retroactively scan the package Socket missed. It was removed from npm. Scanning on demand means you’re going to miss critical relationships across threat groups, or across disparate but related packages.
As for the numbers: see 39 mins into https://youtu.be/jWujI7Hk8O4
I used to work at npm as the SRE manager right before the acquisition and the ballpark figures make sense to me.
Had a second to sit down, and I _way_ overestimated my counts.
Per NPM there are 2,480,373 packages. According to Whitesource [1] there are an average of 12.3 versions per package. This gives us a total of 30,508,587.9 archives to scan. At a cost of $1B to scan that means each package would cost $33.959…
There’s no way that’s reality. That seems like several orders of magnitude too high.
1. https://threatpost.com/malicious-npm-packages-web-apps/17813...
Yeah I’d get it if their solution was only tangentially applicable (it’s directly applicable). I certainly don’t want YC:HN to turn into GPT generated ad-city either.
If you met Feross, you’d agree he and his team are deserving of the credit for finding a timely and constructive solution to supply chain problems without filling OpenAI with your IP. He lives the kind of values we respect on YC:HN.
I had the pleasure at Refactor Conf in Toronto a couple of weeks ago. You can watch his talk here: https://youtu.be/jWujI7Hk8O4
(I have no pecuniary interest in Refactor conf. I just think it was an awesome time.)
1. An article is posted with informative, useful, original content.
2. An important motivation of the article author is marketing to get people to use or become aware of their product.
I see no problem with this. I think it's always important to keep the motivation of the author in mind, regardless of whether it's an ad for a product or just the author themself.