Making Rust supply chain attacks harder with Cackle
davidlattimore.github.io
davidlattimore.github.io
Now of course you could disallow unsafe in most crates, but any that do are now much higher value for compromise.
> That said, using unsafe to say perform network access is harder than just using Rust’s std::net APIs, so we’re at least making it harder for a would-be attacker
No, it’s actually quite easy. Done every day by C and C++ developers and in a few hours you have built up some high level APIs. Heck you can copy paste the std or nix implementations since that’s what they’re doing under the hood anyway.
- Require audit if there is new unsafe code
- Otherwise, rely on cackle to enforce no use of fs/net etc in safe Rust
This could provide the best of both worlds, automating most of the audit burden while still providing strong guarantees.
In general, sandboxes don’t work well at the language level. You really need to go at the system level.
Even at the language VM level, it doesn't seem to be tenable.
Microsoft tried to go all out on this back in the day with Code Access Security. I remember three things about it:
1. Engineers/sysadmins would easily get frustrated, and just let the app run under full trust
2. Perf issues, since security demands would result in checking the call stack up
3. When they changed things in .NET 4 a lot of web code would break unless you added a magical attribute.
Needless to say, microsoft more or less gave up on it in .NET core
Unfortunately supply chain security is incompatible with developer convenience. At least not without a lot of work to make it bearable.
We will have to suffer through a lot worse attacks than now before people will take it serious (most developers likely never but governments will at some point intervene - see EU's CSA).
Back then, you often just ran VS as local admin, supply chain attacks weren't a 'real' thing most of the time, so NBD.
So then you try to deploy your app, and discover the joys of signed assemblies.
And you -make absolutely sure- when you leave, you give instructions to rebuild the whole pipeline if need be.
TBH at least we knew there was the polite illusion of a sandbox...
However, there is some appeal to the syntax introduced by the author if we use it for a proper and portable sandboxing mechanism. Maybe WASI, with capabilities?
To be more specific, there’s no reason for rust to know that writing to a specific file will allow modifying the program’s memory. It’s also not a security problem from the system, it’s just how it works. It really only makes sense for the system to enforce that kind of sandboxing, because it has enough context to enforce things sensibly.
There you can more exactly specify what a function may do, instead of relying on blunt categories like "filesystem".
I think this is what OP is talking about when he calls using the std::net API’s easier.
Driveby hackers are (probably) not going to spend a few hours re-implementing a network call.
Also, there are too many things to keep track of for the authors of such approaches. Eg: How would this tool handle changes in the standard library API?
Languages and repositories need to be funded appropriately to solve a problem that is systemic.
I've played with several methods, built tooling around this problem (github.com/R9295). Hell, my thesis is going to be about supply chain security. Solutions that inform the developer and allow them to make decisions don't work.
I get that there is no absolute defence and security is layered etc. But these approaches feel hacky
Not using Rust though (yet?), so can't vouch for Cackle. Great idea though.
Hence why many companies force internal repos on CI/CD, and dependencies only uploaded after legal clearance.
Even if true, the uptake fraction (e.g. usage %) isn’t the key point: to provide an additional (rather novel to boot) layer of analysis and therefore security. Remember: security is layered.
> Developers just don't care enough.
Even if narrowly true, this neglects to recognize that developers exist in varied situational contexts. When properly motivated, some can use these kinds of tools.
Crates.io, NPM, Maven repository, etc are a wild west. That's fine for OSS developers and hobbyists. I think it's crazy for a business.
Honestly, if I were starting a commercial project from scratch, I'd dispense with it all and "vendor" packages like Google does: they go in your (probably mono) repository as tagged, versioned, maintained-by-you, directories. Perhaps as a git submodule, perhaps as a fork. But not pulled from the Internet.
Yes, you can set up your own mirror or private repository, but this only solves half the problem.
Because the other half is cultural: systems like cargo or npm encourage very deep dependency chains, lots and lots of packages. It becomes exhausting and difficult to track what is doing what and who is maintaining it. It leads to bloat in build times, but also to brittleness and complexity.
Tracking dozens of semvers and dependency trees I think becomes seductive to people as a kind of exciting busywork. But it's not usually solving the core business problem. It very often creates new ones unrelated to the thing you're supposed to be thinking about: shipping a stable, working product that solves customer needs.
I'd disagree with this one being characterised along with the others. When i want to publish on Maven Central, i have to:
1. Prove i own the domain i'm about to upload a package under, e.g. if i claim com.myname - then i'm going to need to prove that to Maven Central by creating a DNS TXT record on com.myname
2. Sign my release - every jar i publish needs to be signed by my (or my org's) GPG key
3. On top of the automated mechanical controls, there's an actual human sign off in the loop for the registration process at least
This might still leave attacks like typo-squatting potentially open but not an easy thing to do and it effectively stops most of the other horrors like replacing an already published artifact with a malicious version, or "brand-jacking" my library's name and pushing up a new malicious release.Crates.io doesn't even have the concept of an organizational namespace. It's ridiculous.
But Maven still has the broader cultural problem of automatic dep resolution: it's just too easy to go adding deps for every little utility and nifty function or feature or framework.
- Trust it blindly
- Audit the code (and do that again for each update)
- Write it yourself instead
The last two can be time and resource consuming so you sometime have to choose the first solution.
Cackle can be a useful tool to (occasionally) raise alarms for when dependencies you trust blindly start using different APIs (so the trust isn't completely blind anymore). But it doesn't really solve the problem.
This sort of capability-based approach to security would make untrusted code relatively safe to execute because the worst it could do without the explicit cooperation of the developer is an infinite loop.
I wonder if anyone tried to use it to limit dependency risk in that way.
The whole paradigm is to avoid needing to check permissions by making it impossible in principle to do anything you’re not allowed to do.
Plug: we've been building Packj [1] to detect malicious Python/NPM/Ruby/Rust/Java/PHP packages. It carries out static/dynamic/metadata analysis to look for "suspicious” attributes such as spawning of shell, use of files, network communication, use of decode+eval, mismatch of GitHub code vs packaged code, and several more.
But you can get around source scanning too by compiling your code, `include_bytes` the executable, mmap'ing and then calling the entrypoint.
Unsafe is not really a red flag, except for people who don't understand what unsafe means. And like you point out, if you see it, you need to audit the code regardless.
(For non-rust folks, using mmap requires using unsafe, so you can't map your own machine code in without being explicit about unsafe)
Forbidding unsafe is a big gun. You need unsafe code for any FFI, so attackers will just look for crates that link to system libraries and add their malicious code there. If you want to do any kind of fast iteration over slices/vectors you need to use unsafe to explicitly elide bounds checks.
I think the real nugget of truth is that the vulnerability is `cargo update` and that tooling needs to look at changes that happen between versions, and alerting to new unsafe or extern bindings is a good way to do that. But I still see a few obvious ways around it - what you really want is a linker shim that can detect when symbols are being relocated into the final executable that shouldn't be there.
As I said, it helps. It's not perfect. But there are a lot of crates out there that don't use unsafe, and having a way to focus your attention is useful. Looking at some of my own projects, for example, I notice that I often use crates like "env_logger" with 3.9M downloads. What an absolutely magical target for a supply chain attack -- but it doesn't use unsafe, or net, or tokio, or a lot of other things, and pinning that restriction with Cackle would let me update it with much less worry. I kinda like that.
I'm really not arguing that this is a silver bullet. It's just .. nice.
Edited to add, as a counterpoint: The "log" crate, which is used by env_logger, _does_ use unsafe. Great target for a supply chain attack -- and perhaps reason to nudge the developers to find ways to get rid of it!
Or perhaps nudge users to understand what `unsafe` means and why its necessary rather than evangelize at maintainers.
I'm quite familiar with what unsafe means. And in the context of supply chain attacks, there's a clear benefit from going from "a few uses of unsafe as a possible performance optimization" to "no unsafe". Obviously, that doesn't work for all crates and we shouldn't demand it. But I made the comment about log after looking at the code to understand its use of unsafe.
> For crates where you don’t need or want new features, bug fixes etc, you could consider pinning their versions.
It seems to me that you should _always_ pin your dependency versions. I'd go so far as to say that build tools shouldn't even have an option to automatically pull the latest - "version" should be a required field for all dependencies (and, ideally, that gets decorated with a hash after the first download).
Yes this means you might miss out on security fixes that you'd otherwise get automatically. But having the code that you run & ship change absent your intention just feels like a bizarre default approach.
This is exactly why cargo uses a lock file that does pin your dependencies as described for binary crates, but does not use the same mechanism for libraries. https://doc.rust-lang.org/cargo/faq.html#why-do-binaries-hav...
Crate C depends on them both. It now can't bring in updates to A until B does, and when B updates that's a breaking change, so it better bump its major version.
Take a look at this trick, for example, for foundational crates updating their major version: https://github.com/dtolnay/semver-trick
Now imagine that being an issue every single patch update.
Granted that the supply chain attack defense would fall on the hands of the users of the library. Unless the library writer aggressively pins of dependencies on the Cargo.toml (not the default semver action), problematic, but at least Cargo allows multiple versions of dependencies on a program, on Python for instance this is much complicated scenario (but still the smallest of the dependencies problems there).
For the new guidance: https://doc.rust-lang.org/nightly/cargo/faq.html#why-have-ca...
See also https://blog.rust-lang.org/2023/08/29/committing-lockfiles.h...
If you pin your dependencies, when something you depend on fixes a security bug you don't get the fix unless you notice the new release and realize you need to apply it.
If you do not pin your dependencies, then your code will randomly break because you are depending on something that changed (Hyrums law). You push to CI and now you have build failures to fix unrelated to your change in the best case. In the worst case everything builds and you ship to customers who discover the issue (even if you have a great manual test plan that would have discovered it: if it is the last build before release you probably don't run the full plan)
I don't have a good answer here. I'm in the pin everything camp, and can report it is hard to notice all changes upstream. It was just a few years ago we finally threw away a computer with Windows XP with no service packs - it sat in a closet for years untouched just in case we had to rebuild some old software.
When I'm developing, I frequently grab latest versions. Until i share, I want to keep up to date.
- `Cargo.lock`: handled implicitly so long as you check it in which, as of rust 1.74 (next release), `cargo new` will scaffold all projects this way by default.
- `Cargo.toml`: requires using a `=` version requirement. We generally recommend against pinning this way; see https://doc.rust-lang.org/cargo/reference/specifying-depende...
Zig might be a better fit. The community tends to like minimal dependencies and bootstrapping. The downside is that it is a relatively new language with all the grief that brings. But, if you're going Rust with minimal dependencies, you're already in "Here Be Dragons" territory.
And, this might especially apply to you as you mentioned C. I see Zig as a better C while Rust is a much better C++.
Rust makes it quite easy to forgo a large portion of the standard library. rustc and cargo are separate tools, if you don't want cargo's features, you can use rustc as normal.
These paths may not be popular but that doesn't mean they aren't well supported, and depending on your niche, may not even be that unpopular. embedded on things that are small enough you're not running embedded Linux, for example, rarely use the standard library.
https://davidlattimore.github.io/making-supply-chain-attacks...
Blocking macros might be, in my opinion, one of the best defense that you can have today.
https://github.com/rust-lang/rfcs/pull/1133 Yeah definitely I haven't been wanting this for years...
Very impressive stuff
There will still be the possibility of confused deputy attacks, but it still looks like a significant improvement to supply chain security.
hardening supply chain attacks is probably not the goal here