Why has software supply chain security exploded?
opensourcesecurity.io
opensourcesecurity.io
• There used to be a focus on attacking endpoints directly, but eventually OSes, browsers, and servers stopped being easy targets, so attackers looked for another way in.
• OS (especially Linux) distributors acted as gatekeepers. Getting into a distro requires having some reputation, and the risk of getting caught with a backdoor is a deterrent (although I don't think distros are as much of security barrier as people assume, since maintainers already have too much work to do and their source code review is at best at "LGTM" level).
• New tooling gave more visibility into dependencies. Transitive and build-time dependencies in precompiled libraries typically aren't visible, so it's easy to think you have one or two deps, but transitively there are still tons of them: https://wiki.alopex.li/LetsBeRealAboutDependencies
• Exploits became easier to monetize. Thanks to crypto miners and crypto wallets any exploit is almost a direct payout. Way easier and safer than trying to sell a botnet, adware or even stolen CC numbers, so it lowered barriers to entry to this "business" as well as made it more attractive.
The example of 1000s of packages having a transitive dep on a package which left-pads strings looks bad, however the two obvious alternatives to this are:
1. everyone writes their own function to left-pad. As well as being wasteful, most of these implementations are incorrect or perform badly or have their own security issues.
2. the function to left-pad is part of a huge package which does thousands of things, eg. the programming language itself, or some standard library. In order to get a new version of left-pad, you have to upgrade the whole thing, and are potentially left with irreconcilable incompatibilities.
You sure? You can write it, with no security issues, with acceptable performance for most cases, in the same time than it takes to find the dependency. In fact, you'll save time since you don't have to vet the dependency.
Uh, you do like, vet your dependencies right, whenever you add them, right? As well as every time you update to a new version, right?
If you're not doing any of that and just adding dependencies willy nilly then I could see how it could seem easier, but this is an illusion. You need a dependency to do more than just left padding to make it worth adding, unless you don't give a crap about stability or security.
> 2. the function to left-pad is part of a huge package which does thousands of things, eg. the programming language itself, or some standard library. In order to get a new version of left-pad, you have to upgrade the whole thing, and are potentially left with irreconcilable incompatibilities.
Yeah this is how languages have traditionally done it, and when done well it is very nice. As you say, it is hard to change the contract because it gets depended on, and you know what? This is a feature, not a bug. Things that are really simple and settled like leftpad don't need to constantly change. This is also why emerging languages try to stay "alpha" without guarantees about breaking changes as long as possible early on, so by the time they have to really freeze the standard library, a lot of experience has helped polish it.
A similar model is something like Java, which has done a fantastic job of evolving its standard library over the years while staying very (but, mostly, not excessively... mostly...) backwards compatible. Mostly through the trick of letting users create their own popular libraries and then eventually integrating it with the JEP process once one gets really big. Kind of like a lot of the stuff from Guava, or Jodatime.
I do think better tooling can exist (and is being developed), but it needs to track the whole system, not just a single language.
If you depend on libcurl, you count that as 1 (one) dependency. Nice, and restrained.
But curl has everything from gopher, FTP to SMTP and IMAP! It has a half a dozen TLS libraries, SSH, auth methods, multiple compression algorithms, socks proxies, etc, It's a lot of code, and a lot of transitive dependencies that you don't even think about.
OTOH when you add `reqwest` in Cargo, Cargo will visibly fetch everything needed to build it from scratch, and it will look like 140 dependencies, not 1. But in total it's still only a subset of what libcurl contains. There's no Samba or POP3. It's just a very detailed list showing you it contains mundane things like a CLI arg parser and base64 encoder, and a wrapper for ca-certificates too.
Please don't interpret it as a criticism of curl. It's just to contrast that you may see libcurl as 1 dependency, because it's one precompiled package, but you may see a smaller amount of code as extravagantly bloated shocking 140 dependencies, because it's just been split into tiny pieces and rubbed in your face.
In contrast, when I tried to add a token-replacement library to mdbook, Cargo installed 187 crates.
There is a real and legitimate difference in the size and functionality scope of older C and C++ libraries with the type of single-function packages you see from NPM and Cargo that has made full supply-chain audits far more difficult. To get functionality equivalent to something like Boost, you might need to vet a thousand separate vendors, the majority of which are pseudonymous Github accounts with no physical address or backing organization.
I'm not just speculating about this, either. I used to be a platform lead on fully-disconnected, classified information systems where we needed to bring in and compile every dependency from source via DVDs hand-carried to a classified workstation. Somewhere around 2015 or so, this stopped being even a remotely feasible job and created multi-year backlogs of teams that had something fully working in their unclassified development environment with Internet access suddenly finding themselves staring at a minimum 6-month, more likely indefinite wait to get to production when they discovered security teams needed to individually approve every package and every developer they wanted to bring onto a classified system.
Cargo/crates.io has a publicly available index of every crate with every dependency. `cargo tree` lists everything in your project. `Cargo.lock` contains inventory of every package used in the build, including their checksums.
Package manages made it easy to have dependencies. Every added dependency adds risk. That doesn't mean you should never add dependencies - there are good reasons to do so - but you should weight their pros and cons. Adding dozends, hundreds or thousands of dependencies (as it is common in the npm ecosystem) adds lots of risk. None of that is particularly surprising.
Regardless, I was very surprised that it could be 7 figures!
Even for my side project - building a custom 3D rendering engine from scratch - I use libraries like System.Numerics to calculate the camera perspective transform and perform all the linear algebra for me. I was wasting a lot of time trying to reinvent this, which is a typical but lethal risk to any project.
We have to contend with not only nuget and the sheer abundance of packages available (with their own transitive dependencies similar to the problems with npm - yes, even Microsoft packages have this risk) posing problems, but also the tooling itself.
VS/Code is a fucking minefield of potential threats. Their extensions market place(s) are automatic RCE risks, and a developer/admin/whomever box getting pwned is a monumental risk.
Zero-Trust applies to everyone and everything. Microsoft don't get a pass just for being Microsoft.
That doesn't mean it's more vulnerable, but it means it has to be less vulnerable to offer the same security. To offer a bad analogy: if a bank vault was as vulnerable as my fridge, people would constantly break into the bank vault but keep ignoring my fridge.
The GP never said Microsoft "gets a pass." They were pointing out that there are stacks like .NET where reliance on third-party unvetted libraries is greatly reduced if not eliminated outright. Additionally, nearly all businesses run Windows anyway, so using Visual Studio and Visual Studio Code are not the risks you make them out to be. Centralized Active Directory policies already help guard against rogue extensions and such, assuming IT even cares (they nearly never do in my experience).
Yes, the surface area threat of .NET and NuGet is far, far lower than that of NPM for nearly all apps.
Too few attackers understood it well enough? Or were other attack vectors easier?
EDIT: It's much like building a tower. Upper floors depends on lower floors. Taller the construction more unstable it gets.
It has resulted in a major thinking shift for me. Mostly initiated by a malicious nodejs package being quietly pulled in.
I am fortunate that my own contributions only require that.
The tradeoff isn't simple, and will be different for each company. I think everyone can point at a couple of components that are complex and massively benefit from the work poured into them and the large install base finding all the edge cases. Equally everyone can name a couple of dependencies that anyone could rewrite in a couple hours, and maybe even improve on the way.
Note that I’m not saying “never use libraries”, which your points seem mainly aimed at. Code reuse is great! I’m just not sold on automatic, unmonitored updates to libraries.
1. one or more major package hubs
2. one package manager, or at least, a package standard
3. easy to consume/build (source code to binary)
4. easy to publish
5. easy to download (one liner in your build script, e.g. package.json, vcproj, cargo.toml, build.gradle etc.)
At the moment C++ only really has 5. - easy to download: cmake, conan, vcpkg, build2, can all download easily.
1-4 are a work in progress for C++, but are far better than 10 years ago.
Perhaps the reason is package managers for C are called distro maintainers (for Linux/BSD)
Eventually, due to a lot of factors, threat landscapes start shifting and attackers start moving their targets and tooling. One of the bigger shifts would be the huge, rapid improvement to both OS and browser security that occurred within a few years, radically increasing the cost of attacking desktop users via malicious websites. Another would be crypto, where 'account takeover' attacks that could lead to wallet access are now easier to monetize + crypto itself as a tool for transfers.
With regards to supply chain, enough of these changes occurred that some attackers took the leap and have started looking at this area. There are probably a lot of reasons why - prevalence of dependencies, increased interest in tech companies, etc.
If it continues to prove viable (it's obviously viable from an attack perspective, unclear if it's something attackers will rally around to monetize) we'll see it escalate and get better tooling around the attacks.
I know that automating stuff to check for CVEs and vulnerabilities of sorts is getting a lot of traction recently, but I don't think this is really sufficient. Who do you trust? Are we going to see some companies creating a "Security Seal of Approval" for packages?
I'm not sure what's the answer here, but certainly it isn't what we have been doing for the last couple years.
You can trust who you pay properly.
Who entrust the survival of theire family to your continued existence.
Whos bread i take, whos song i sing.
Thats the supply chain.
Turns out there is no free lunch to be had even in open source.
Tidelift is one company that has a bunch of "catalogs"[0] of packages. I'm not sure how their package metadata is generated though -- maybe semi-manually?
There is also Bytesafe[1] which is supposed to help give you a way to "firewall" yourself from unapproved dependencies. I don't think they sell data though. Just tools.
A company based on the Open Source project, packj[2], is Ossilate[3] which is trying to use automated analysis for analyzing 3rd party packages.
And another that's in this space is Socket.dev[4] which tries to warn you in your PRs about bad packages.
I know about this space because I work on a project[5] that's also related to supply chain security. It's a bit different from all of the above since we're focused on patching known vulns, but the idea of "vetted packages" has crossed my mind before.
Are there any other services in this space that I missed?
0: https://tidelift.com/solutions/catalogs
Even if you face a supply chain attack, the only value lost may be money for most businesses, and not enough that worrying about supply chain attacks is worth it.
It's still C and C++ world out there, and everything gets shrink wrapped. It's more likely that you would find a logic bug or a well known security issue in a single library than somehow infest the supply chain of these particular applications.
Every piece of off the shelf software is labelled and audited in that world.
1. A conscious shift in focus from triaging risks when they occur to stopping threats before they arise.
This is a natural next step from contemporary security and disaster response. Threat response and continuity planning which both incorporate plans that respond to threats were once the primary objective of organizations. They are still valid, but a more modern and proactive approach includes mitigating the risk at the source.
2. A forced increase in material spend toward securing dev and devops ecosystems in a time where they are one of the most targeted parts of the organization.
One only has to watch the news to see this play out... unfortunately after decades of deployment and intranet security emphasis, hackers have recognized that IP and source code are the best way to get money out of a company, and that both are ironically some of the least protected assets.
Looking at integrated vulnerability reporting as found e.g. in npm (and soon Golang, I think) it seems that 99 % of alerts are just false positives, as they affect build dependencies (e.g. the tar library) somewhere down the chain that will never run in production or even touch anything sensitive, and that cannot be easily exploited even by sophisticated actors. In my opinion the "high severity" threat category is thrown around way too lightly by these tools, which leads to everyone panicking.
If anything good could come out of this whole trend it might be the realization that open-source software needs better funding to address supply chain issue, my expectation though is that most of the money will go to tool vendors that simple tape over the underlying problems by creating an artificial "security layer" between the maintainers and users.
I do think that organizations should have a process for identifying and triaging bugs in dependencies. But "there are vulns in dependencies" is not itself a compelling narrative because there are vulns in your business' code too. And people sure as hell are getting updates more frequently when using npm or whatever then in the alternative world of "link in some open source code and never update ever for all time."
To me, the more interesting problem is the "malicious actor takes over an npm account and distributes malware" that is more unique to the npm ecosystem. This one is hard. On the one hand developers are expensive and a robust library ecosystem is basically a superpower. On the other hand you really need a trust relationship with your software vendors if you want to be confident you are delivering safe code.
Boom, new industry of snake oil.
Not suggesting that the investors have stakes in any companies doing supply chain audits or building automated solutions for it, of course.
If you're looking for a more precise metaphor, think of it as a large pot of slightly overcooked spaghetti and meatballs.
Where anyone can throw stuff into the pot.
But people have automated the throwing in of stuff, and even when the original stuff was excellent, stuff sources are sometimes bought out by bad actors.
And sometimes the stuff has ingredients, where the ingredients are also so managed, so I would say that there is still a chain.
Cf. https://discourse.nixos.org/t/generating-software-bill-of-ma...
Then npm came, along with some ideology about code-reuse, and the friction went away. But the friction was serving an important function: if adding a library is annoying, you'll add a small number of large libraries rather than a large number of small ones, and you avoid getting an ecosystem where major libraries have hundreds of tiny dependencies. This is important because the friction of adding a dependency is only a small portion of its true cost: the main cost is that you're trusting too many different developers and developers' computers.
Can we keep the benefits of automatic dependency resolution while adding some other, artificial deterrent? What if NPM hosted packages for free if they have less than five (or whatever number) of transitive dependencies, but above that charged some number per additional one? Even if it were a small amount, this might put significant downward pressure on the size of dependency trees.
I've been delivering systems where we had to deliver the source environment (ie what is used to build production) into escrow with a lawyer as part of the contract. It effectively required us to prove that what we delivered (both hardware and software) into escrow, would be able to reproduce what was in production, reliably and provably.
Reproducible builds are really hard, especially with people encoding CVS/SVN/git tags into the build, compilers and linkers not producing equivalent results between runs (eg caches warm/cold can make a difference) etc etc.
It's a real problem and a real issue.
Well, companies used to have employees write code, rather than stitch together random garbage written by random dipshits who could be tricked into using loose licenses. That's one cause for concern.
> Open source won because everyone worked together.
No, ``open source'' has ``won'' because it helped corporations defang Free Software and get gratis labor.
> Supply chain security will happen because everyone works together. If you try to do this alone, you will fail.
This isn't food, but perfectly uniform applied mathematics. There's no ``working together'' with corporations. Doing it alone is the only reasonable option, which means aggressively reducing the size of this foul shit is necessary.
> Thank goodness others had the foresight to create our current SBOM formats.
Yeah, MicroSoft totally *<3* ``open source'' guys.
> Long ago all the software supply chain work was done by hand.
Oh, right, all of this new security theatre is always about trust and reputation, and not trusting those disgusting lone programmers such as me or other silly things; it's always really about doing anything but truly auditing that yucky code.
It was up to the dev to "fix" this "defect" before code could be merged.
There was no pinning policy. There was no actual security review (where do we use this? are we vulnerable? why do we use this? are there alternatives?).
Meanwhile it's running in production. (have we discussed this with our MSP? can I discuss this with our MSP? has our MSP notified us of any issues?)
Although we say ignorance is no excuse, having spent a decade in cybersecurity and 20 years before that as a quality-focused consultant, I'd say ignorance lets people get away with a lot. But I know better and I would never be allowed to claim "I vass only followeeeng zee orders!" if the shit hit the fan.
https://blog.conan.io/2019/09/02/Deterministic-builds-with-C...
https://reproducible-builds.org/ https://bootstrappable.org/