A study of malicious code in PyPI ecosystem
about.honywen.com
about.honywen.com
The https://en.wikipedia.org/wiki/Principle_of_least_privilege as constructed by https://en.wikipedia.org/wiki/Capability-based_security is the solution for this, yet Python will probably never be capable of this kind of internal encapsulation, it's too much of a fundamental change - and even if some sort of sandboxing ability is accomplished, creating separate/recursive sandboxes (needed when importing more, separate libraries which in turn import libraries whose authority they want to limit) will probably require another interpreter instance for each sandbox (as with WebAssembly).
I hope current and future language designers will take this into account, and construct their compilers, virtual machines and interpreters accordingly. Python was created before the internet as we know it now existed, so perhaps its lack of security mechanisms shouldn't be surprising. But it and any new developments that fail to consider this aspect of computation will be fundamentally flawed from the beginning.
It's 2023. I never want to run untrusted code on any of my devices outside of a sandbox ever again.
WebAssembly feels like it should be that sandbox. Every few months I check to see if it's easy enough to run regular Python/JavaScript/etc code inside a WebAssembly sandbox yet... we're not quite there but it feels really close.
More notes here: https://til.simonwillison.net/webassembly/python-in-a-wasm-s...
I have Docker on my Mac. I run things in it if I have to, but it's pretty resource heavy and I have to look at my notes every time I want to run "docker run".
Plus Docker isn't actually a fully secure sandbox!
There's Firecracker, which I believe is secure (AWS built it for Lambda after all) but good luck figuring out how to run that on a Mac.
I want something that's tiny, fast, ridiculously easy to use and that I can run dozens if not hundreds of instances.
WebAssembly is almost that thing.
Then, how about using AWS Lambda for your use case?
I'm still hoping for a solution to the run-on-personal-devices problem though.
From an isolation perspective, that's giving you hypervisor level isolation which is a step above most container-based solutions, which you point out rightly isn't a full sandbox. A hypervisor with CPU support is probably the gold standard for isolation, asides from running it on a separate physically isolated system.
Personally, I use Google Colab for my convenient, lightweight Python sandbox. But I’d rather not depend on Google. It feels weird that this is my best option for now.
`bwrap` is great for this on Linux. A one line alias should let you run any program in a lightweight and highly customizable sandbox safely-ish. See my reply to the parent comment for an example.
```
$ alias SafeRun="bwrap --ro-bind / / --dev /dev --proc /proc --unshare-all --tmpfs ~ --bind ~/Untrusted ~/Untrusted"
$ SafeRun python3 ~/Untrusted/RandomScript.py
$ SafeRun ~/Untrusted/RandomExecutable
```
This is basically manually invoking what Flatpak does:
https://github.com/containers/bubblewrap
This is also useful for more than just security. E.G., you can test how your app would behave on a fresh install by masking your user configuration files. I personally also have a tool that uses it to basically bundle all dependencies from an entire Linux distribution in order to make highly portable AppImages— Been meaning to post that, will get around to it eventually maybe.
The flags above should hide your user data (`--tmpfs`), disable network access (`--unshare-all`), hide/virtualize devices and OS state (`--dev` and `--proc`), and make the rest of the root filesystem read-only (`--ro-bind`— Including the insecure X11 socket in `/tmp`, which you might want to expose for GUI apps).
Check them against `bwrap --help`; I might have omitted one or two more things you'd need.
Though, this is likely to be a cat and mouse game for the foreseeable future. Detection will get better, and attackers will change tactics.
In the meantime, we've been open-sourcing some tooling to help protect developers from these sorts of attacks. Namely, a sandbox that locks down network/disk/env [1] and our CLI [2] that allows you to perform a `pip` install in the sandbox, after checking our API for behaviors/issues with the package. For example:
phylum pip install <pkgName>
Really glad to see software supply chain security getting some academic, rigorous study. Backstabbers Knife was one of the first I came across, and it's been a consistent stream of papers since.How do you see detection ever getting better to the point where it isn't a game of cat and mouse? Isn't the whole point of these security threats that go undetected is that they are novel attacks? Or even popularized attacks with different parameters?
In these trust but verify models of a central registry, it will require some heavy lifting like AI security detection to even keep up with. That seems like a new problem in itself if introduced.
[1]: https://about.honywen.com/publication/2023ase/2023ASE.pdf
(I believe most Chinese mirrors use bandersnatch[1], like most non-Chinese mirrors.)
Anything touching the internet where user provided content will be consumed should be hardened and preferably sandboxed in some way, stuck in an unprivileged DMZ, and monitored.
As for the pypi modules themselves, a web of trust system might actually work. For scientific packages especially. There's always the risk of an author getting their machine compromised and someone pushing an update with either a subtle bug which can pass unnoticed or a loud and quickly detected attack. Making a change in account 2FA tokens visible would be a great start, as would integrating signing keys.
(The long file name in this case was the digits of pi.)
That seems exceptional. The registry maintainers must be working extremely hard on these problems.
(At least I think it's to accompany this paper, I found it in a footnote and it's not entirely clear if it's officially related or not.)
https://github.com/ossillate-inc/packj/blob/main/.packj.yaml
Secondly, what about impersonation where attackers imitate a popular package and its respective metadata?
Packj detects typo-squatting (impersonation) as well.