NPM repository flooded with 15,000 phishing packages
scmagazine.com
scmagazine.com
> A large number of corrupted packages use names related to game cheats, free resources, and social media platforms, such as "free-tiktok-followers" and "free-xbox-codes," to entice users to click the links and direct them to multiple well-designed phishing webpages.
So just because the spammers are using the TikTok brand/name in their spam (also mentioned is "xbox" but they didn't use the Microsoft logo), they used THAT as the image? Not the NPM logo or something much more relevant? I know it's minor, but that rubs me the wrong way as very "sus" and clickbait for when the article is used on social sharing sites and they most likely pick that image as the article open graph image. You also know it's most likely for clickbait because the avg person is much more likely to click some link with the TikTok logo and "phishing" in the title, vs some NPM logo that they may not even know.
> "While being flooded with spam is never good, it gets immediately noticed and mitigated. It's harder for open source projects to spot and stop rare one-offs"
This is the real problem that NPM and other ecosystems face. A determined attacker that is trying to "poison" a popular Open Source package just has to feign as a maintainer long enough to succeed[0].
Defeating these types of attacks will require rethinking how we think about trust of packages. Projects like Deno are one approach (fork the ecosystem) while projects like Packj (mentioned elsewhere here), Socket.dev, and LunaTrace[1] are taking the other angle (make it harder to install malware).
It's hard to say which approach is better right away. (Probably a hybrid of both, realistically) It's just non-trivial to fix this in one clean swoop. It's messy.
0: https://www.trendmicro.com/vinfo/us/security/news/cybercrime...
https://www.merriam-webster.com/words-at-play/bury-the-lede-...
There are still issues with things like history being re-written if you're pinned to a tag instead of a commit... since potentially the hash might change and you can't necessarily understand what changed unless you have a copy of the original code still.
That said, I feel like the bigger challenge here is with transitive dependencies (and a large total number of dependencies) because they create more places for malware to "hide", so to speak.
If every time I bump a package, it bumps 50 others... am I really going to notice/review the changes in all of those? Odds are I'm just trying to ship some feature or do something else! The libraries are rarely the center of my focus. (Maybe you get a PR with a security fix that you land it without thinking. Surprise, malware!)
It's an insidious problem without a very clear cut solution. Do we stop using Open Source code? Probably not. So how do we adapt? TBD!
> If every time I bump a package, it bumps 50 others... am I really going to notice/review the changes in all of those?
Is there a difference between registry and codehost here?
> That said, I feel like the bigger challenge here is with transitive dependencies
This seems to be more a result of the stdlib and ecosystem, Node & NPM being the poster child for this
> There are still issues with things like history being re-written if you're pinned to a tag instead of a commit...
The Go team runs a sumdb server to catch this, and NPM does something to prevent tag drift, though we don't do this with our container images. It is trivial to move an image tag on docker hub or any image registry.
You should have a go.sums committed next to the go.mods file which also checks that the code at the tag is the same as you expect, so the sumdb is not necessarily needed if you trust the hash you originally downloaded and committed to git.
> It's hard for the enterprise world because they need to mirror everything internally.
We explored this, NPM and PyPi present the same issues as Go, the shear volume and maintenance makes it practically not worth it.
---
What about having additional accounts with a registry/hub in the software supply chain, that have been demonstrated to be vulnerable to account take over? Adding this extra middle layer in the supply chain seems like an inefficiency and increases risk.
This happened before during the glory days of SourceForge where project names were global, and there was all the resulting control and misdirection issues. Thankfully Github made it a two level namespace OWNER/PROJECT which vastly reduced the problems. IMHO the repositories will need to do something similar.
Namespace partitioning has advantages, but it also doesn't solve squatting problems on its own: anybody (or any process) that's fooled by `requests` vs `pyrequests` is also going to be fooled by `some-owner/requests` vs. `requests-org/requests` (the latter being the fake one).
I think the most "complete" solution here needs to involve code signing.
My point was about that in this case the problem is the users, or education, not the available features.
Say you have a package that only triggers the evil activity one out of 10,000 evaluations? It'd have a 99.99% of passing validation but could still wreak havoc on a heavily used client side application.
Ditto for a package that simply delays it's malicious payload until some point in the future, after the scan has occurred? I guess this is a bit easier to test for by future dating the system clock, but again what if it's a random window? (e.g. every other day for five minutes, at 4:59pm UTC).
Having said that, obfuscated code (e.g., base64 encoded) or runtime payloads (e.g., downloaded from pastebin) will not be analyzed with static analysis, but it will tell you if the recent package version uses base64 or exec APIs in the code (file/line) for you to take a deeper look. We've some ideas on runtime lightweight sandboxing, and will eventually implement those.
...You mean, in the same manner everyone should have been auditing code all along, but hasn't, thereby creating the problem you're trying to solve?
It's alright. You're trying, and there's nothing wrong with that answer, but the people you're trying to sell a shovel to don't want to dig. That's the chief problem.
That's exactly where a tool like Packj can help. It can cut the digging time down significantly. Auditing hundreds of direct/transitive dependencies manually is impractical, but Packj can quickly point out a "risky" behavior (file and line number). Using obfuscation to hide malicious code itself is a suspicious behavior, which is flagged. After all, we have to start somewhere.
It’s not a bad idea to stake out a product that’s mature and ready to go when that happens. It doesn’t even have to be foolproof. It just needs to be earnest enough to check off some boxes and show that “best practice” effort was made.
Is this tool unix-only or just the sandbox thing?
Nevermind, found the answer regarding PyPI install:
> Packj only works on Linux.
But won't this require your target users turning off SIP [1] on Macs to enable dtrace? A lot of dtrace tools require SIP to be disabled before they can produce any meaningful output.
1: https://developer.apple.com/documentation/security/disabling...
The world would have to be aware of these attack vectors and attack tactics to be able to mitigate them. That's not to say they haven't existed in concept for the last 10 years (which many have), just that very few utilized them maliciously until the last few years across the industry. (Also see the exponential increase of supply chain attacks in the last couple years alone and the vector/tactics used)
Security tends to be a cat and mouse game in my experience. If you wanted to even go down the Sun Tzu perspective, the best defense is to subdue the enemy without fighting. That allows one to understand the strengths and weaknesses of one and their enemy and to develop a defense that is flexible and adaptable to the moment and harden it for the future.
We download code
From the internet
Written by unknown individuals
That we haven't read
Which we execute
With full permissions
On our trusted devices
Where we keep our most important data
We're working on making it similar to email spam. Will post more soon. Meanwhile, we welcome changes that can make it less noisy for the community.
[Edit:] one can specify Github API token to get around rate limiting. We'd deeply appreciate if could create an issue for this on the repo.
We trust that there has been a testing process that validates that things do what they say they will do. Admittedly when dealing with non-OSS vendors you can sue them if something goes badly wrong.
The issue is rather that this is only a largely collaborative one. Which makes the thing kind of worse, because you can _often_ get away with simply trusting everyone and there are economic and practical pressures to do so.
A system like that cannot give any modicum of control, because typically a real world project will involve hundreds of package publishers, any one of which can break that trust with a poisoned patch version of a transitive dependency (or can be hacked), and it is effectively impossible to review all those dependencies every time they do a patch version bump. So developing with npm is always an act of faith that everyone will do the right thing. The remarkable part is how rarely it breaks down.
Most JS OSS is MIT licensed. There are no guarantees. It’s on _you_ to ensure the software works.
Secondly people don’t really “trust” packages. They trust the usage, weekly downloads and quality of docs. They just assume this implies that it’s going to be OK, or that when deps break it’s nobody’s fault. Shit happens.
You could already do what you're describing today, but odds are, you don't, because you value your time more than you value "having small amount of packages" for whatever reason.
> A large number of corrupted packages use names related to game cheats, free resources, and social media platforms, such as "free-tiktok-followers" and "free-xbox-codes," to entice users to click the links and direct them to multiple well-designed phishing webpages.
The same problem exists on Docker Hub, e.g.
https://hub.docker.com/r/neyprivedpu1972/windows-7-7264-down...
What makes me wonder is why it’s always NPM that’s hit with these kind of things, and not, say, Maven? Arguably the attack surface of Maven is much larger and intrusive (lots of enterprises), but the mechanisms of package distribution are completely different. There are a shitload of scans and sanity checks happening when you push a package to a Maven repo such as Sonatype, which I don’t see as much with NPM (and also Pypi for that matter)
(fwiw, I am responsible for package distribution of all APIs of a reasonably popular timeseries database)
Along with all the scanning and what not, I think that’s the biggest reason you see attacks primarily on npm, PyPi, and to an extent Ruby Gems.
People all over the world are getting desperate.
Otherwise, it is way too easy to automate the submission.
It makes a lot more sense to use cryptography to verify that releases are not malicious directly. Tools like crev [1], vouch [2], and cargo-vet [3] allow you to trust your colleagues or specific people to review packages before you install them. That way you don't have to trust their authors or package repositories at all.
That seems like a much more viable path forward than expecting package repositories to audit packages or trying to assign trust onto random developers.
[1]: https://github.com/crev-dev/crev [2]: https://github.com/vouch-dev/vouch [3]: https://github.com/mozilla/cargo-vet
[1]: https://learn.microsoft.com/en-us/dotnet/core/tools/dotnet-n...
Then again, if someone wants to pay me to do nothing but read and vouchsafe code, hell, why not. As QA I basically do that anyway. Will have to think on this.
Yes ultimately someone you trust has to vouch for the library for you to trust it. That's unavoidable either way. Those systems only allow for that person to be anyone instead of just the author.
Hypothetically, this could be done: put up a public key at a known domain, sign packages via public key, and if the package manager extends to include a "fetch from URL provided by the package and check signature" tool, we end up rooting trust in packages in trust in domain name ownership; not too shabby. We still need devs to actually bother to do the check and, like, read the domain name to make sure they haven't accepted authorization from "micr0soft.com", but the end result is a distributed signing solution that doesn't require any centralized authority (except, indirectly, the SSL cert issuers and DNS providers of the world).
You could have alerts when an old account is using a new key that nobody has seen before, for example.
... we do it anyway because the alternative is to just get 0wnz3d by bad actors.
Automating GPG signatures is trivial. GPG is not an anti-spam mechanism.