NPM security update: Attack campaign using stolen OAuth tokens
github.blog
github.blog
> Using one of these AWS access keys, the actor was able to gain access to npm’s AWS infrastructure.
How many individual best practices were not followed to result in this nightmare? Sigh.
Keep those keys out of source control, folks. There are a lot of options for secrets management these days, and making it harder for attackers to totally own you if they only manage to crack one piece of your infrastructure is key to limiting damage from this sort of attack.
I have this setup pretty well in my code now but getting there wasn't simple or easy from my perspective and keeping the list of secrets your IAM user can access up to date can be a pain as well.
I'm working in a lambda environment so my options might be more limited but I'm interested to see how other people are solving this issue (maybe specifically for small/side projects). As it stands my lambdas all get a role applied to them that gives them access to the secrets but something not AWS-specific would need a "bootstrap secret" to be injected before the code could call out to the third-party to get the other secrets. For Lambdas I suppose I could inject that "bootstrap secret" in via environmental variables but now I've got a new issue to deal with. Injecting at build time via something like GitHub Actions Secrets is an option I guess.
All that to say, while I agree secrets should never be in source, in practice it's not super easy (I'd love to be proven wrong, maybe I'm not doing it right).
There is a middle-step between "lets have API tokens committed in SCM" and "lets deploy a full-authentication system/use this costly solution", and that is using environment variables. In your code, do `process.env.MY_SECRET_KEY` instead of `myGitHubPersonalToken` and then when you run the program, run it with ` MY_SECRET_KEY=myGitHubPersonalToken npm start`. Magically, you can commit your code without exposing any secrets, and share the secret where you need it out-of-band.
Zero-cost, actually easier to configure your software when you need it, and as a bonus, people won't get access to your infrastructure in case someone gets a hold of your source code.
That npm inc isn't aware (or failed to uphold the code quality) of environment variables for secrets is embarrassing.
But where does this live? Or do you literally mean that Jane The Sys Admin is supposed to type this into her terminal every time the service restarts in the middle of the night?
What if I need to replace a node? Or scale a service? How do these secrets get there?
Depends on how the service is deployed. If you're just running it on a Digital Ocean instance by manually SSHing into the instance and running systemd services, define it in the .service file (it supports defining environment variables).
If you're doing instances via automation (like Terraform), most of them (including Terraform) supports loading things from environment variables. So you run `MY_SECRET_KEY=myGitHubPersonalToken terraform apply` when you create the instance, and use the environment variable in your hcl definitions.
Software security is rarely free (even with an OSS tool you've got infra and management costs), but the cost is almost always cheaper than a major breach that could stem from something like this incident, which fortunately was pretty contained.
Our goal is to simplify this as much as possible. We use client-side end-to-end encryption so you don't have to trust a third party, and pricing on our paid plans is based on number of users, not number of secrets.
Just like the rest of Airbnb started before VPCs were GA and thus required a large engineering investment to move everything to VPCs, we started on the secret management stuff before there were a lot of good other options available (though arguably Hashicorp Vault was around and mature enough at the time, and would have been the best alternative). I haven't looked at envkey for production use but I've definitely considered it for home use since it's just so deliciously simple.
[0] https://medium.com/airbnb-engineering/production-secret-mana...
I know there's a free "community" hosted version, but I'm not sure what the differences is outside of the limits and support, and I'd prefer to see the pricing scale up a bit more gently than 0 -> $150 as soon as I reach the limits of the free offering.
Very clearly cut down to the point where it's not feasible for anything outside of toy projects.
The open source version is fully functional and can definitely scale beyond toy projects. You don't get high availability, multi-instance clustering, auto-scaling, multi-region failover etc. built in, but if you put it on a beefy host it can easily handle a large number of users and a very high request rate.
The way I think of it is we give you the fully functional server, but charge for advanced infrastructure and a few advanced features (SSO/Teams).
It's comparable to the open source version of a tool like Vault where you get the server, but need to implement advanced stuff like HA, auto-scaling, networking, etc. yourself, or else use a paid version.
As far as I can see, this is fully available in the open-source version of Vault: https://learn.hashicorp.com/tutorials/vault/ha-with-consul
This is only true if I can be confident that I can replicate the hosted setup with the open-source version if I invested the necessary resources. Otherwise, the existence of an open-source option adds little value. In fact it can turn me off from a product since it'd seem like they're using open source as a marketing hook with no real intention of empowering users to be able to actually move off their hosted platforms.
This is why Vault requires another piece like Consul (plus a whole lot of tricky infra/networking work) to achieve HA.
That said, we could allow users of the Open Source version to specify a url via an env var to look up a host's internal IP so that clustering would work.
Auto-scaling is provider-specific though, so I don't see how that could be baked in. Same with secure networking.
I'll also just say that while we do want the open source version to be fully functional (if a bit more DIY), another motivation for us that I see as equally important for a security product is transparency.
While it's inarguably crucial for any clients implementing end-to-end encryption to be open source, I think there's a lot of value in open sourcing the server as well (regardless of how practical it is to actually run) so that users can know what's happening on the server-side, see that the code is high quality and tested, and so on.
> EnvKey Business Self-Hosted runs in an AWS account you control. You can use it with any host or cloud provider.
Is a bit confusing, that section should be clarified a bit. Does it mean that you can run the systems using it somewhere else? Does it mean other variants of EnvKey can be run everywhere? ...
2FA is already effectively built-in to EnvKey through device-based authorization. A user can only sign in to EnvKey from an authorized device, so an email account compromise won't be enough for an attacker to gain access--they would also need access to an authorized device.
A passphrase can optionally be supplied on top of this for an additional factor (though it's unnecessary if you're already using OS-level disk encryption).
It's basically the same model as SSH. And imo it's superior to SMS or app-based 2FA (perhaps not token-based, which I'm open to adding). It handles the main threat models (phishing/email account compromise) with far better UX and convenience.
"Is a bit confusing, that section should be clarified a bit. Does it mean that you can run the systems using it somewhere else? Does it mean other variants of EnvKey can be run everywhere? ..."
I agree this could be less confusing.
The EnvKey host server runs in your AWS account, but that doesn't mean that apps you integrate EnvKey with are in any way limited to AWS. You could have your apps running in Heroku, GCP, Azure, or whatever, and integrate with your self-hosted EnvKey installation for configuration and secrets management with no problem.
Is that clearer?
Anybody ever do something like that? How effective it would be probably depends on unit test coverage.
You could also probably just do the same thing in prod with a dummy user.
I've always wanted to apply strong type systems to this problem – wrapping sensitive data in types that do not have the ability to be printed to logs would theoretically allow you to know after type-checking that passwords can't be output to logs. However again I think this is wishful thinking as a password needs to be sent somewhere at some point, and that creates places where issues can occur.
There are alternatives to avoid this, on the same model of SSH key authentication where the secret stays on the client-side.
Nothing prevents them from using a password to derive a private key using PBKDF2 in client-side and answer to a specific challenge.
Our tests would set up the app's full context, get a hook into the logging framework to watch for log statements, then make requests to the service containing a set of dummy credentials, like { username: "foo", password: "bar" }. If a log statement containing "foo" or "bar" was detected the test failed.
It's not going to catch every type of issue, but at least some potential footguns can be preventing this way.
This way it would blow up on the test that is leaking the credential so you could track it right down and it would transparently apply to all current and future unit tests without any more effort.
They did actually patch it before I got there though.. but they didn't get rid of the years-old log files with the passwords. Found them while trying to find the root password (unsuccessfully) for a host that we couldn't reboot. The ones I tested still worked.
I wouldn't be surprised if something similar happened here. Old log files in backups and such.
> Using their initial foothold of OAuth user tokens for GitHub.com, the actor was able to exfiltrate a set of private npm repositories, some of which included secrets such as AWS access keys.
So NPM was storing AWS secrets in their (private) git repos. IMHO that was an accident waiting to happen.
This isn't the first time GitHub has found logging of plaintext credentials [0]; it's not a good look, for a company with the resources that GitHub has, to have to disclose it again almost exactly 4 years later.
[0] https://www.zdnet.com/article/github-says-bug-exposed-accoun...
PII, sure, even login usernames or email, also sure, but not credentials (plain text passwords, tokens, etc).
It happens. You fix the issue, purge the logs, and learn from it.
APIs are extremely fraught, because users like to build integrations that jam credentials into the wrong places until they get a 200. You haven't lived until you've added a regexp for your API keys appearing in the wrong HTTP headers
It was fixed much later. You can look over the git logs.
I'm not sure this is so much of a "how good this happen," but a "thank you for being transparent." Most organizations cover this sort of thing up, like MIT/Harvard (and now 2U).
Good of github to announce this openly!
Meh. Shit happens. If we all Pikachu face every time an exploit happens we are lying to ourselves. We'll never reach perfect security. It's a pipe dream.
What matters more is the disclosure and response. I'm not a huge advocate of npm personally, but I respect their response to this thus far. From what I gather (the email was a bit long-winded) nothing vastly detrimental occurred, they automatically invalidated passwords and going to publish again next time will require a couple minutes tops of extra work. I'll take it.
Let's all stop acting like products need to be perfectly and eternally secure. That's not how threat modelling works, any security professional knows that's impossible, and it's unfair to expect that from anyone, including big corporations.
Npm has done a lot of relevant and good work toward their security efforts over the years, in some cases going a bit far even in my own opinion. The comments I've seen so far have been a bit unfair.
But for gods sake, having secrets hardcoded in VCS??
You seem to understand threat modelling. What's the threat towards one of the biggest and most used package registry?
It's not like npm Inc just started running the registry. They have been doing this for years. To let such a beginner mistake risk the supply chain of basically the entire JS ecosystem is not only sloppy, it's completely unprofessional.
What can we do in the short-term? I'm not sure, but I hope smarter people than me comes up with some solutions ASAP before a compromise like this starts actually impact developers using npm.
0: https://docs.github.com/en/actions/deployment/security-harde...
https://blog.tedivm.com/guides/2021/10/github-actions-push-t...
I know the available libraries are a fraction of the ecosystem, but very often it's a good enough fraction if you are willing to be flexible in your choices.
The extra friction also makes Linux package managers useless. Just about 0% of the things people are installing with npm exist in distro repos. Distro repos are also extremely poor at keeping multiple versions around at once.
Surprisingly people don’t like having their react upgrade tied to the Debian major version update.
It varies widely depending on the distro.
Debian is the the most strict on review process but also on screening the volunteers doing the work.
Yet, most popular distros has been very effective at blocking a large number of supply chain attacks over the last 20 years.
I remember that somebody posted a reply to a comment of mine here on HN years ago saying that s/he was rebuilding every single gem as deb package before deploying it in production and that was the only sensible way to do. I don't think it adds much to security unless they also read all the code, but it's a lot of work that none of my customer is going to pay me for. I also probably don't want to start a profession of deb builder for Ruby gems.
That's almost certainly true, but the case of having to choose, I'd almost always select slightly older but trusted code to less trustworthy bleeding-edge-newness.
I guess that's why I'm content with dozens of instances running Debian.
Vulnerabilities can be discovered in released software, but new vulnerabilities are only introduced in new releases.
Security fixes are backported by the distro. Over time, stable distributions become and remain more secure than cutting edge software.
> I don't think it adds much to security unless they also read all the code
Some distros ensure that hundreds of thousands of users deploy and use the same packages consistently before making a distro release.
This creates plenty of accountability, making it very difficult for a supply chain attack to go undetected.
> A fraction and based on my old Ruby gems memories, fairly out of date.
Just like buying a car, you have a tradeoff between well-known and little-known, well tested or bleeding edge.
What's safer for production use?
This one?
EDIT: it's a bit old, probably due to the usual mess of dependency management in JS. Still perfectly usable tho.
Having seen that, I would stay away from react.
Stop using thousands of packages? Start vetting packages, as if security was important?
There are already dozens of us saying it is possible to not have too many dependencies, and vet packages before installing. But every time we open our mouths we are treated as if we just escaped some sort of insane asylum.
At a place I worked in the past we used to have a 40-line microservice using plain-node without any dependencies. That was by design. One junior dev took it upon himself, in their spare time, to convert the whole thing to use some js MVC framework, complete with a full-blown build process, transpilers, and all the nine yards. There was a big discussion in the PR and a lot of juniors complained that we should migrate because they "didn't learn plain node.js in college".
We can't have nice things anymore.
Ensuring that the core infrastructure of their software systems doesn’t have the same security standards as teenagers wordpress site would be an awesome start.
It's Java or Python for programming basics (arithmetics, ifs, print), then straight into frameworks.
I find however this will become harder and harder with newbies being constantly bombarded online with messages stating that creating anything from scratch is a futile exercise, together with companies influencing college curriculums.
What we need is a set of policies around dependencies that we agree on as an industry and tools to help make securing our systems easier.
We're working on it: https://slsa.dev/spec/v0.1/requirements
To vet dependencies properly you have to stop using thousands of them willy-nilly.
To stop using thousands of them you mostly have to change only your development dependencies. Runtimes dependencies, even in large Javascript projects, are often well behaved and are rarely a big issue. Except maybe for a few large backend frameworks.
No, that won't destroy your company.
Anything that helps vetting will be more than welcome, though.
> Will this be met with a shrug from the JS community?
Go read the comment that begins "Top 10 maintainer here". Not even a shrug.The JS ecosystem sucks, but anyway, not particularly their fault in this case.
Supply chain security is immensely important, and I encourage you not to learn about it the hard way like I did. Which somewhat ironically happened in the .Net ecosystem when one of our trusted Nuget packages got hacked many years ago. Now, I could be mistaken and I hope I am, but I suspect that if you ask a Java, a JS and a C# developer if they trust their ecosystem, then only one of them is likely to say yes.
So no, there won’t be some great revelation in the JS community. The best you can hope for with stories like these is that fewer developers feel like imposters when they realise that GitHub stores plaintext security assets in their logs.
The implication of a successful attack on NPM, with huge unvetted dependency graphs currently in fashion, would be that any of the thousands dependencies of a modern small JavaScript app could suddenly include malicious code that runs your dev machine or your production systems.
(That's why the key part of the announcement is "GitHub is currently confident that the actor did not modify any published packages in the registry or publish any new versions to existing packages".)
https://github.com/cloudsecurityalliance/CSA-IT-Operations/t...
https://github.com/cloudsecurityalliance/CSA-IT-Operations/b...
As Heroku did, in its own hamfisted and still too slow way.
So, there was a blog post on April 15, at least.
[0] https://github.blog/2022-04-15-security-alert-stolen-oauth-u...
[0] https://github.com/ssrathi/go-scrub [1] https://www.nutanix.dev/2022/04/22/golang-the-art-of-reflect...
Even if you screw up, the impact is so much less severe.
But yes, client side hashing is just one part of a security story, but it sure does make log leaks less damaging.
Too vague to be useful. Stolen from whom? Simultaneously from two different third parties? The whole thing began April 12? Based on what signal and from whom?
Here's what this vagueness tells me: GitHub OAuth token management for integrators is badly operated and poorly monitored by the entity issuing tokens.
My conclusion is that Microsoft (GitHub), SalesForce (Heroku), and Idera (TA Associate, Travis CI) have not learned the SolarWinds lessons of Supply Chain Security yet. Perhaps this will be the incident that helps.
Some of the posts in this thread are a little too c'est la vie for my tastes.
> Too vague to be useful. Stolen from whom?
Heroku and Travis CI. That information is literally in the end of the sentence you cut off:
"On April 15, we published a blog detailing an attack campaign utilizing stolen OAuth user tokens issued to two third-party GitHub.com integrators, Heroku and Travis CI."
Which is so frustrating. When you upgrade your hashing algorithm, always always _always_ immediately remediate the weak hash mess by hashing your weak hashes with your new stronger hash, and turn the login check into
bcrypt(sha1(user-entered-password)) == stored-bcrypted-sha1
you can then upgrade them to a straight bcrypt if the check succeeds, but keeping the weak hashes on disk indefinitely until the user logs in (if ever!) is such a risk.“The password hashes in this archived data were generated using PBKDF2 or salted SHA1 algorithms previously used by the npm registry. These weak hashing algorithms have not been used to store npm user passwords since the npm registry began using bcrypt in 2017. ”
It means some combination of the following: people don't make the mistake in the first place, the mistake is caught in code review, the mistake is caught in audits, the mistake is caught by automated tooling...
This is important for this forum, imo. Heroku was cast a lot of shade during this.
> This is probably a good time to remind people to check their authorized OAuth applications on Github[1] and make sure that any unused apps have their access revoked.
>
> [1]: https://github.com/settings/applicationsThis is awful advice. I’d never trust my authentication needs to external providers who can hold my user information hostage.
The same can be said for GCP (Identity Platform) and Azure.
To your point about availability, AWS, GCP and Azure all have managed authentication services, that are fundamental to their platforms. I highly doubt they are going anywhere. I have yet to have an Cognito outage in three years.