Dozens of malicious PyPI packages discovered targeting developers
blog.phylum.io
blog.phylum.io
Example of defining project dependencies:
{
"apollo-client": {
"version": "...",
"access": ["fetch"] // only fetch allowed
},
"stringutils": {
"version": "...",
"access": [] // no system resources allowed for this dependency, own or transitive
},
...
}
It would probably require the language to limit monkey-patching core primitives (such as Object.prototype in javascript),
and it would be more cumbersome for the developer to define the permissions it gives to each dependency.
These required permissions could be listed on the package site (eg npm or PyPI) and the developer would just copy paste the permissions when adding the dependency.
But if you upgrade a dependency version and it now requires a permission that seems suspicious (eg "stringutils" needing "filesystem"), it would prompt the developer to stop and investigate, or if it seems justified add the permission to "access" list.[0]: https://deno.land/manual@v1.27.0/getting_started/permissions
Good luck with the project. I hope it delivers, because I’d love to use it. Signed up for the mailing list.
Deno is a really cool project, imo.
It does exactly that (although on a per-process basis).
I don't think this kind of permission system can be retrofitted into an existing language without direct OS support, and probably not at the library level (you'd need something like per-page permissions which would get hairy real fast).
But it gets better. If you build Python in the Cosmopolitan Libc repository:
git clone https://github.com/jart/cosmopolitan
cd cosmopolitan
build/bootstrap/make.com -j8 o//third_party/python/python.com
Then you can use cosmo.pledge() directly from Python. $ o//third_party/python/python.com
Python 3.6.14+ (Actually Portable Python) [GCC 9.2.0] on cosmo
Type "help", "copyright", "credits" or "license" for more information.
>>: import cosmo, socket
>>: cosmo.pledge('stdio rpath wpath tty', None)
>>: print('hi')
hi
>>: socket.socket()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
File "/zip/.python/socket.py", line 144, in __init__
_socket.socket.__init__(self, family, type, proto, fileno)
PermissionError: [Errno 1] EPERM/1/Operation not permitted
Since we didn't put "inet" in the pledge, you can now be certain your PyPi deps aren't spying on you or uploading your bitcoin wallet to the cloud. You can even use os.fork() to rapidly put each dependency in its own process, then call cosmo.pledge() afterwards to grant each component of your app its own maximally restrictive policy.Cosmopolitan Python also ports OpenBSD's unveil() system call to Linux too. For example, to disallow all file system access, just call cosmo.unveil(None, None). You need a very recent version of Linux though. For instance, I use unveil() in production on GCE but I had to apt install linux-image-5.18.0-0.deb11.4-cloud-amd64 in order for Landlock LSM to be available to use unveil().
That is, you can protect your own program from doing network stuff because of incorrect input, but you can't use it to sandbox another program.
See this thread: https://marc.info/?t=162367803300003&r=1&w=2 and this mail about sandboxing: https://marc.info/?l=openbsd-tech&m=162367954705721&w=2
You've described OpenBSD in general. I recommend a deeper dive - it's fantastically refreshing, how simple yet functional an OS can be.
I often use the FreeBSD Handbook as an example of first-party documentation done right, and that's only possible because of deliberately limited "churn for churn's sake". The kinds of regular code rot and attrition that Linux suffers just does not take place on BSD systems because if you contribute something new you're expected to make your thing mesh with the ecosystem, they don't tolerate the "I'm gonna churn everyone's environments because my version is 10% better than the existing one" that tends to take place on Linux.
As an example, the amount of init systems that Ubuntu has gone through over the last 20 years is completely insane by BSD standards. They've gone from sysvinit to upstart to systemd. You don't have people ripping out the graphics and audio subsystems and doing total rewrites, either. It's nuts that a lot of Linux people don't even realize that it doesn't have to be that way - that's not "just how maintenance is supposed to work", that's a bunch of incredibly bad behavior on the parts of distros and Poettering specifically that has become highly normalized and overlooked. Maintenance shouldn't be breaking things like that. "we never break userland" doesn't have to be a suicide pact either - BSD has actually maintained a much more stable ABI than the linux kernel without crystallizing a bunch of shitty bug-compatibility stuff like Linux is doing either. It's all insanely stable and competent and professional by linux standards.
If I was developing a product for true long-term support, with the minimum possible engineering effort devoted to fighting churn, BSD would be at the top of the list.
If you read "rolling stones gather no moss" to mean "keep moving or risk becoming obsolete", choose Linux. If you understand it as "change for change's sake prevents the achievement of mastery", go with a BSD.
A if a language wants to deal in library's security it should strive to make static analysis possible. Eg: the language guarantees that network and filesystem calls can only be done with a single function, statically so I can audit that leftpad indeed doesn't make network calls .
Currently allows you to specify allowed resources during the package installation in a way very similar to what you've outlined [1].
The sandbox itself lives here [2] and can be integrated into other projects.
1. https://github.com/phylum-dev/cli/blob/main/extensions/npm/P...
const apollo = require('apollo-client', {fetch});
const stringutils = require('stringutils');
That way, you can pass in at a very fine-grained level any object you want to, including things like `fetch` and `fs`. But then, you could just as easily pass in a custom `loggedFetch` module like so, that wraps `fetch` and logs all network requests made: const apollo = require('apollo-client', {fetch: loggedFetch});
This is the object-capability security model, and as a sibling commenter pointed out, LavaMoat basically does this for JavaScript dependencies. (This is also basically just lexical scope, by the way, if your language is strict enough - though most unfortunately aren't.)Also- tangential but I really wish people would stop using require and use ESM whenever possible. I cringe when I see `require` just as much as when I see `var` being used in 2022, but I’m not certain there’s never a good reason not to use ESM.
var bar = require("bar", {baz: loggingBaz})
var foo = require("foo", {bar})
This is true even if you never otherwise use bar.This is certainly how peer dependencies work in nix flakes, for instance, where you can say something like
bar = //....
bar.inputs.baz.follows = loggingBaz;
foo = //....
foo.inputs.bar.follows = bar;
Since nix flakes is a package/dependency manager, it generates a lockfile that you could inspect too see what all the dependencies are, and make sure you got it right (oh, I still have two bar's, I must have forgotten to override some other dep). I suppose any compiler that involves a linking stage would in principle be able to generate some comparable output at the language level.(Apologies for using require and var, but I'm convinced the functional syntax is semantically clearer in this context.)
If you own a package at version 1.X.X and you want to add a permission requirement you have to bump the version to 2.0.0. If you also allow people to opt in to a less strict "auto allow all the currently required permissions for these dependencies" mode, they would at least know for sure nothing can touch anything new unless they explicitly bump the major version.
If you're extra concerned about security you can explicitly specify them so it's really obvious when a major version bump adds new ones, but it removes some of the friction.
(a) Discourage any future use of ">=" in version dependencies. Specify an exact version. That way a future compromised version doesn't get pulled
(b) Every build system needs better ways of having multiple versions of a same dependency coexist. I should be able to have one of my project's dependencies depend on "numpy==1.15" and another dependency depend on "numpy==1.16" and they should be able to coexist in the SAME environment and "see" exactly the numpy versions they requested.
For python we should think about how to support something like this in the future:
import numpy==1.15
and have it just work.That way if a hacker compromises PyPI and releases a malicious numpy 1.19 it won't get pulled in accidentally.
Here's a bit of a joke I made before that might be an interesting starting point, though since it uses virtualenv behind the hood it doesn't have a way for multiple versions of one package to exist. I don't think it's impossible to do though with some additional work.
https://github.com/dheera/magicimport.py
Sample code:
from magicimport import magicimport
tornado = magicimport("tornado", version = "4.5")Now I see what people at one respectable, big project were thinking when they allowed 7 different versions of OpenSSL to be statically linked to the same executable…
Seriously, this idea may save you from some not very interesting work, but it will create the need of much bigger amount of work which while potentially interesting is not very productive. You are toying with exponential growth here, like a chain reaction - like a bomb.
foo() becomes v1_15_4_foo() automatically
Anyway I'll leave this from 2013 here:
> After having work during 3 years on a pysandbox project to sandbox untrusted code, I now reached a point where I am convinced that pysandbox is broken by design. Different developers tried to convinced me before that pysandbox design is unsafe, but I had to experience it myself to be convineced.
> It would also be nice to help developers looking for a sandbox for their application. Please tell me if you know sandbox projects for Python so I can redirect users of pysandbox to a safer solution. I already know PyPy sandbox.
-- https://mail.python.org/pipermail/python-dev/2013-November/1...
1. `autobox` (to be renamed lol) [0]. It's basically a Rust interpreter that performs taint and effect analysis, reporting on both, allowing you to use that information to generate sandboxes. ie: "autobox sees you used the string '~/.config' to read a file, and that is all the IO performed, so that is all the IO you get".
2. I'm working on a container based `cargo` with `riff` built in that aims to work for the vast majority of projects and sandbox your build with a defined threat model.
The goal is to be able to basically `alias cargo=cargo-sandboxed` and have the same experience but with a restricted container environment + better auditing of things happening in the container.
3. I previously built a POC of a `Sandbox.toml` and `Sandbox.lock` with a policy language that allowed you to specify a policy for a given build step. Unfortunately, I couldn't decide on how I wanted it to work in terms of "do I generate a single sandbox for the entire build, or do I run each build stage in its own sandbox" - there are tradeoffs for both.
Here's a lil snippet:
[build-permissions.file-system]
// All paths are relative to the project directory unless they start with `/`
"../" = {permissions = ["read"]}
// "$target" being a special path
"$target" = {permissions = ["read", "write"]}
// Source this path from the environment at build time, `optional` means it's
// ok if it isn't available
"$env::PROTOC_PATH" = {permissions = ["read", "execute"], optional=true}
// Default protobuf installation paths, via regex
"^(/usr)?/bin/protoc" = {permissions = ["read", "execute"], regex=true}
Once I'm done with (2) though I think I'll tackle (3).`autobox` is fun but I think it may be impractical without more language level support and no matter what I'd end up having to implement it in the compiler at some point, which means it would be unusable without nightly or a fork.
I'm going to try to wrap up an autobox POC that handles branching and loops, publish it, and see if someone who does more compilery things is willing to pick it up. As for (2) and (3) I believe I can build practical implementations for both.
Though, like scopes, I think many times packages would need broad access, but maybe not?
I suspect you may run into similar sandbox escapes once things are complicated enough. So it seems like a good idea if they can be made bug free, but good luck with that?
But, yes, we've tried this idea many times now, and it never held up for long.
In the end they became yet another attack vector, and now everyone should use OS security services instead.
I think that's the better approach - just assume all packages are malicious by default. Can't rely on scanners because of the large number of packages and attacks.
1. An attacker who can access your development environment and your production environment is worse than one who can only access your production environment. You might say "but the end goal is prod", but it's not that simple because of (2).
2. We already have very good tooling for isolating services at runtime. Separating them onto different instances, firewall/security groups, limited API keys, docker/ containers, apparmor, selinux, etc. We have a lot of tooling for "a service in production is owned". What we lack is "a library in dev environment is owned".
3. Devs often have more privileges than your services. It's unfortunate but at a lot of companies, perhaps given some lateral movement around dev envs, you'll find SSH keys to production, browser session cookies that give you console access, source code, chat sessions, internal documents, git keys, gpg keys, etc.
So I'm actually fine with a tool that sandboxes the build process but leaves open the hole of "but the attacker can patch the binary and execute code in production". That's a huge win.
If we can spot odd behavior during development and eliminate it from our stacks, the product will be more secure for end users too.
There was a time when convenience overrode any security doubts in my mind. But now I routinely use these tools to restrict access, monitor, and review runtime behavior.
The only allowed package (gem) server was one ran by the project. This package-server scanned, vetted and manually checked any version of a lib before "publishing it".
If you wanted to e.g. upgrade a package, you'd have to do this on this server first. It would then go through some steps, -automatic scanning, risk analysis, sometimes even needing the eyes of someone from a security team. After that the package was published on this server, and you could pull it onto your dev machine and use it in CI/staging/test/prod etc. Similar steps, to get a new package listed.
IMO this is better, because it stops supply-chain attacks before they hit your code, not after they've (potentially) infected the system.
Edit: for clarity "only allowed package" wasn't enforced very strictly. A linter and CI would catch any changes to code that would want to fetch packages from elsewhere. It wasn't to protect against rogue developers, but against "stupid me, accidentally upgrading to a version that is infected" and such.
If you run everything from the new UID, it will mostly be contained to it's own $HOME directory and be unable to modify your user's files or system files. Some distros do not protect home directories from being read so it might be worth setting your actual user's $HOME to umask 0700 or whatever.
If you are using bwrap while running X11 and not running the sandbox with a new UID, the sandboxed processes may be able to escape via the X11 socket! This can happen even when you don't mount the X11 socket into the sandbox (see abstract sockets)! I think unsharing the network namespace fixes this specific issue (not 100% sure), but there are probably more subtle footguns like this.
I really suggest running Wayland with XWayland disabled, and the Wayland socket protected from the sandbox if you want to use bwrap for security purposes!
Is there no way to safely run graphical applications in a bwrap sandbox? I thought Wayland was supposed to be better about this.
On Wayland, assuming you don't have XWayland enabled and running, it depends on the specific compositor you are using and what Wayland protocols it supports.
Sandboxing GUI stuff on Wayland requires at the very least not having XWayland running, and also requires understanding what the compositor allows clients to do by default. Some compositors may have permission dialogues that prevent clients from doing stuff that you didn't expect.
Unless your code is never going to touch important data or resources, like for example (but not limited to) being used commercially in any vein then you can’t keep it in a padded cell forever.
Where possible, persuade end users too to be equally careful. The "linux is safe" cliche blinds both us and end users to its obvious security problems, like running every script as the logged-in user with the same level of access. These malware developers know it and rely on it. That's why we need to move everybody towards a restrict-monitor-verify mindset by default.
It's absolutely not a foolproof approach, but it is a lightweight layer that can be used in a "defense in depth" approach.
I'm very frustrated with firejail since I can't for example block execution in my home directory, with the exception of one subdirectory.
It just can't be done.
There have been to many contaminations of major package repos lately. Only one typo in an import statement up the dependency chain and you’d be compromised.
We are actively working on a solution that will fully sandbox package installations for npm, yarn, poetry and others.
It's rolled up as part of our core CLI [1], but is totally open source [2]:
[1] https://github.com/phylum-dev/cli [2] https://github.com/phylum-dev/birdcage
Though I’m not sure of the solution really is / should be increased sandboxing.
The alternative may be a rethinking of the increasingly smaller packages. Maybe it’s better to have few large packages maintained by reputable organisations or personalities?
We really need a defense in depth approach here. Sandbox where it makes sense, perform analysis of code being published, consider author reputation, etc.
If you put code in the package itself, this would side step the "installation" sandbox. However we're also doing analysis of all packages introduced to the ecosystem to uncover things that are hiding in the packages themselves.
So you're right, we need a defense in depth approach here.
If you log into your email from the virtual machine, you are at risk.
With local development environment it is a bit different, because unless you are running build/test etc. in a container/vm/sandbox, then attacker has access to all of your files, especially web browser data.
Not in Qubes OS:
https://github.com/Qubes-Community/Contents/blob/master/docs...
[0] https://www.qubes-os.org/doc/how-to-use-disposables/
[1] https://www.qubes-os.org/doc/how-to-copy-and-move-files/
(If it makes any difference, I would probably be using VMWare Workstation Pro)
How do you move data in/out of the guests? I always found that part of interacting with VMs to be annoyingly painful.
Yes. Yubikey. ecdsa-sk key requires you to tap yubikey to have a working key. It consists of 2 parts - a private key file, but which is useless without yubikey. https://developers.yubico.com/SSH/
https://developers.yubico.com/SSH/Securing_SSH_with_FIDO2.ht...
Azure DevOps does it too
1. https://github.com/ossillate-inc/packj/blob/main/packj/sandb...
It DOES NOT require a VM/Container; uses strace. It shows you a preview of file system changes that installation will make and can also block arbitrary network communication during installation (uses an allow-list).
Disclaimer: I've been building Packj for over a year now.
Doesn’t even have to be a typo if the actual project is compromised. Like one of the 100s of NPM modules without 2FA for publishing.
With modern virtio interfaces for network, disk and graphics practically giving near metal performance, there's no reason to not utilize VMs for development.
I think it's an interesting cultural phenomenon that different language communities have different levels of dependency fan-out in typical projects. There's no technical reason golang folks couldn't end up in this same situation, but for whatever reason they don't as much. And why is nodejs so much more dependency-happy than python? The languages themselves didn't cause that.
Part of it—but I'm sure not all—is that the core language was really, really bad for decades. Between people importing (competing! So you could end up with several in the same project, via other imports! And then multiples of the same package at different versions!) packages to try to make the language tolerable and polyfills to try to make targeting the browser non-crazy-making, package counts were bound to bloat just from these factors.
Relatedly, there wasn't much of a stdlib. You couldn't have as pleasant a time using only 1st-party libraries as you can with something like Go. Even really fundamental stuff like dealing with time for very simple use cases is basically hell without a 3rd party library.
Javascript has also been, for whatever reason, a magnet for people who want to turn it into some other language entirely, so they'll import libraries to do things Javascript can already do just fine, but with different syntax. Underscore, rambda, that kind of thing. So projects often end up with a bunch of those kinds of libraries as transitive dependencies, even if they don't use them directly.
Could it be that nodejs has implemented package management more consistently and conveniently than other languages/platforms?
npm pulls dependencies into node_modules as a subdirectory of your own project as default.
Python really should consider doing something similar. Dependencies shouldn't live outside your project folder. We are no longer in an era of hard drive space scarcity.
That’s not to say that there’s not still some libs out there that haven’t updated docs to get with the times.
I don't think you really gain much either; vendoring was useful before modules, but now we have modules and go.sum I don't really see the advantage. If you have "github.com/foo/bar" specified at version 1.0.4 the go.sum will ensure you have EXACTLY that version or it will issue an error in case of any tomfoolery.
Going on a trip somewhere without an Internet connection? Checkout the repo on your laptop and go. Without vendoring: oh shoot, I forgot to download the deps, I guess I’m going to be forced into a work-life balance. With vendoring: no additional step needed after checking out the repo. The repo has everything you need to work.
Another case: repo of your dependency is removed, or force-pushed to overwriting history. You’ve lost the ability to build your project, and need to either find another source for your dependency, or rewrite it. With vendoring: everything still works, you don’t even notice the dep repo went under.
Generally, with vendoring your code is in just one place instead of being a distributed being which crumbles when any part of it gets sick.
Moreover, relying on checksums to me seems a bit overcomplicated. It’s like going to a pub and giving each drink from a stranger to a chemist for verification to make sure they didn’t slip any pills, when you could just carry your own drink around and cover the top with your hand.
> Another case: repo of your dependency is removed, or force-pushed to overwriting history. You’ve lost the ability to build your project, and need to either find another source for your dependency, or rewrite it.
The GOPROXY (https://proxy.golang.org/) still contains that removed repo, and since everything is summed people can't just force overwrite it. Plus, you still have it in the module cache locally.
You can of course always come up with "but what if...?" scenarios where any of the above fails, and all sort of things can happen, but they're also not especially likely to happen. So the question isn't "is it useful in some scenario?" but rather "is it worth the extra effort?"
> Moreover, relying on checksums to me seems a bit overcomplicated.
It's built-in, so no extra complications needed.
That’s assuming I’ve built the thing previously on that same computer. I’m talking about the common case of working on a normal desktop day-to-day and then switching to a laptop, when travelling to a place without internet (or internet of such a poor quality you might as well not bother). With vendoring I don’t need to think about any other steps than copy/checkout the repo. The repo is self-contained. Without it, I’m making the quantum leap to a checklist.
But like I said, it's not about "is it useful in some scenarios?" but "is it worth the extra effort?" I'm a big fan of having things be self-contained as possible but for this kind of thing modules "just work" without any effort. Very occasionally you might go "gosh, I wish I had vendored things!", but I think that's an acceptable trade-off.
pip install google/tensorflow
It would significantly reduce the attack spacequite why python refuses to learn from anything that went before it I really dont know
“Namespaces are one honking great idea — let's do more of those!”
Actually it should be not just a key but a whole TLS certificate, with references to a CA, activity dates, etc.
So "_google/tensorflow" would be official.
"google/tensorflow" would not be (plus it would be reserved by default to avoid confusion).
not perfect, but better.
Haha, simple & effective...
I've personally been hacked by a supply chain attack via a GitHub wiki link. I contacted GitHub support and didn't hear back from them for 3 months. They are completely useless.
GitHub announced a few years ago that they would crack down on malware and were about to introduce some very strict T&C. After a huge backlash from the pentesters (justified in my opinion), they backpedaled a little bit. Hosting pentesting tools is fine, using GitHub as your C2 server or to to deliver malware in actual attacks is not.
I was a skid once, I get it, probably a lot of us were.
I contacted them, showing the plainly obvious malicious account that was distributing malware. Two months later, they send me a generic message saying that they've "taken appropriate action", but the account and their payload was STILL THERE, they hadn't done anything. The attacker was rapidly changing their username, and honestly I'm not sure their support staff has a way of even dealing with that. I tried to explain the situation as best I could, but they were not helpful in the slightest.
I despise the idea of GitHub removing any code just because YOU (anyone) think they are criminals.
Of course, this doesn't protect against compiling malicious code, and then running the code. But at least I try to shut off all attempts at simply compiling the code being a vector.
I haven't heard of anyone creating a malicious source file that would take advantage of a compiler bug to insert malware, but there have been a lot of such attacks on other unsuspecting programs, like those zip bomb files.
I'm not familiar with D, so I'll use the example of Rust. My usual workflow looks something like this
1. Make some changes
2. Either use `cargo test` to run my tests or `cargo run` to run my binary.
In both those cases the code is first compiled and subsequently run. I care if running that command gives me malware. I don't care at what step it happens.
Walter's solution here allows the compiler to be used by the editor without the editor being susceptible. Which at the very least negates the need for a pop-up in your editor asking for permission.
Say there is a Github repo with a crate called "totally_safe_crate", and I want to use this crate in my project, but I am not sure whether I can trust it or not. What do I do?
What I would likely do is clone the repo, then open my editor and look through the source and whatnot and make my decision.
In this case I never intend to run "cargo run" at all, but I may want to run an LSP server to help inspect the code, or I may accidentally enable LSP in the editor out of habit or something.
In this case, it would be nice if I could be certain that simply inspecting and reading code was safe, but as it stands now, in Rust, this is not the case. We can't even inspect code to make sure it's safe unless we open the source files in a "dumb" editor.
In your example you are just adding a crate to cargo.toml and firing off cargo, so of course it's not going to be useful there. Some of us may want to be more cautious than that and actually read code before putting it in our projects.
I'm not experienced but I thought it was normal to have some way to try out what you've written on your dev machine. Is everyone else stepping through code in their head only, and their code is run for the first time when it's deployed to production?
This way you don't need to sandbox the compiler, and it can freely use system resources and access source trees. You only need to sandbox the execution.
(As some people point out in this thread, editors are starting to use compilers to get overall meta-information, too-- if you can't even -view the code- to tell if it's malicious without getting exploited, that's bad).
Note that CTFE runs as an interpreter, not native. Although there are calls to JIT it, an interpreter makes it pretty difficult to corrupt.
I've basically converted to doing all node development in docker containers with volume mounts for source, now it looks like python is going to need to be there as well, at least for stuff that pulls in any remote dependencies.
I don't disagree at all. We're building an open source sandbox for devs right now for this exact reason. Linked it in another comment.
One way to address this is to move to a traditional "debian" style system, where packages are people affiliated with / known by Debian/Mozilla/Google, and specifically aren't the developers of the software themselves. The software is written by Developer X, but is then packaged and distributed by Packager Y, who ideally has no commercial affiliation with Developer X. If Developer X sells out to Malware Corp Z, end users can hope that Packager Y isn't part of that deal and prevents the malware from being packaged and distributed. This still isn't bullet-proof, but it's a lot better.
Also devs should get into the habit of providing sha256 hashes on offical channels (i.e., github readme) so users can validate (if its possible to validate a pkg before executing malicious code in the python ecosystem, I'm not sure how that'd work).
If we are talking typos or other human errors, guess we could only warn people that there are other package with similar name available. Can't predict what people have in mind when they make a typo.
[1] https://www.cisa.gov/uscert/ncas/current-activity/2021/10/22...
I would think the easy solution is to publish a public signing key per-person or per-project, and then sign individual files with that. So, GPG.
This is an excellent idea. Authors are something we are digging into heavily as part of an ongoing effort to improve trust in the open source ecosystem.
Afaik there are two types of supply chain attacks strategy - you either compromise the a legitimate package by somehow getting a PR approved with malicious code, which is very hard to do, or you "typosquat". The latter is way easier and probably the dominant strategy, so package repositories such as PiPy need to invest into preventing it.
Edit: formatting, grammar
I haven't encountered this in the wild yet, but IMO it's reasonably easy. More work than typosquatting, for sure, but probably less work than getting a malicious PR accepted in an existing project.
[1] It's just an example, doesn't exist, and if (by now) it does, I never meant to hint it is malicious. DO is just an example, and I have nothing but good to say about their services (and libraries) where they exist.
[2] The example here would be to just copy the S3 library, because DO block storage is fully compatible with S3 API. Which would also be the reason a dedicated library doesn't need to exist.
Pip install ansible --from curated.pypi.org
Pip install anssible --from wildwest.pypi.org
Let them typosquat all they want in the wildwest. I don't worry about this nonsense with Apt. Heck maybe curation becomes a revenue stream to pay to get in to support such activites.
This isn't a case of 'oh, python (javascript, etc) programmers are dumb and lazy'. If C++ had a packaging system and it wasn't so horrible to add third party dependencies we'd be seeing the same thing there.
But malicious actors can get value from polluting the sharing network, and that costs effort to defend against, which means someone(s) has to pay to secure the network, or be open to attack.
// help my name is ###
// i am being held at #### (address in china)
// please contact my family ###
(this was in chinese, we had to translate it)
Scary!
It's actually quite devious, like a reverse honeypot that the bad guys use against the good guys, exploiting their empathy.
You get some vetting, and in addition, standard practice for many distros is to build everything that goes into the repos in sandboxes or VMs which have restricted or no network access. Additionally, some package managers incorporate that kind of sandboxing into their builds categorically, like Nix and Guix. (For Nix, this may only be on Linux— there are issues with sandboxing on macOS.) So if you build your project's dependencies via Nix or Guix, you're also protected.
This only protects you from `setup.py`-type (build time) attacks, of course. If the distro packages get compromised in some other way so that malicious code ends up in your installed programs (this attack has elements of that, IIRC), you're still in trouble.
https://medium.com/agoric/pola-would-have-prevented-the-even...
Hah. Is this true? I find it funny since IRC has/had this reputation for being a means of communication with malware and it's often blocked on this grounds.
Nice to know that malware is going on with the times and is using Discord for that now.
https://sanctiontrace.com/malware-hosting-providers-sentence...
I'd be very surprised if something shady isn't there as of now.
Is this also happening for repos for other languages (e.g. CPAN, RubyGems)?
RubyGem: https://www.bleepingcomputer.com/news/security/malicious-rub...
Perl CPAN https://news.perlfoundation.org/post/malicious-code-found-in...
Many languages since then decided that causes too much overhead.
It can run code during installation phase
It’s very easy to obfuscate due to its dynamic nature. An import of urllib or os.system isn’t immediately visible at the top of a file as it would be in Java, it can be hidden in eval() or basically anywhere.
Finally, even legitimate packages have very, and I mean it, very bad names. Usually with a trailing number for no good reason. Lacking organizational namespaces. Together with a culture of using a lot of such dependencies and lacking a culture of freezing transitive versions. Blending in among those is too easy. Just take a common name and suffix it with 3.
The certificate could contain information about the owner and the consumer could check if he wants to deal with the owner or not. Developers could add a desired whitelist to pip (or use a curated one) to continue using automation.
On the other side, we have to contend with the fact that malware can be slipped into otherwise legitimate packages. This has happened numerous times over the years. In this case, the hash would serve as a way to say "yup, you definitely got malware". Useful for incident response, but I think we can do better and try and prevent these attacks from being viable in the first place.
Somehow validating packages before inclusion?
IIRC mirrors for NPM, Packagist and others is not impossible, can be done for PyPY and others too?
Maybe it's a stop-gap before all the fancy permissions feature build out (which seems hard)
I know this is also possible for Python because we did it at Uber. I don't remember the specific details anymore though.
In either case though, a lot of people have written proxies for this use case (I helped write one for NPM at Uber). Companies like Bytesafe and Artifactory also exist in this space.
We're working on something similar that's on GitHub here: https://github.com/lunasec-io/lunasec
Proxy support isn't built out yet but the data is all there already.
You can see the rules we tried here[1].
[1]: https://github.com/pypi/warehouse/blob/main/warehouse/malwar...
The attack surface area is too big when random python code is executed, which is the case for `setup.py`, but even if there wasn't code executed there, as soon as you import the package and use it, you'd have the same issue.
Manual auditing is impractical, but Packj can quickly point out if a package accesses sensitive files (e.g., SSH keys), spawns shell, exfiltrates data, is abandoned, lacks 2FA, etc. Alerts could be commented out if don't apply.
1. https://github.com/ossillate-inc/packj
Disclaimer: I developed this.
Feeling a bit left out, guys! How will my code get compromised randomly?
Shamless plug: Wasmer [1] and WAPM [2] could help a lot on this quest!
[1]: https://wasmer.io/
[2]: https://wapm.io/
Explicit allow-lists and policies such as requiring "greater than X downloads per week" go a pretty long way to filtering out malicious packages.
Disclaimer: I currently work for Sonatype, but in a different area of the company.
Or about 190 downloads per package, and who knows how many of those are actually real downloads by victims.
This is basically a non-issue; a few fools downloaded random code from the internet and ran it on their computer and, hopefully, learned a lesson. Shock and horror!
I think it makes more sense to verify individual releases. There are tools in that space like crev [1], vouch [2], and cargo-vet [3] that facilitate this, allowing you to trust your colleagues or specific people rather than the package authors. This seems like a much more viable solution to scale trust.
[1]: https://github.com/crev-dev/crev [2]: https://github.com/vouch-dev/vouch [3]: https://github.com/mozilla/cargo-vet
Package Dependency land is a crazy place
the same argument can be applied to, "when you park on the street in manhattan, close your windows and lock your doors". Well that won't save your car from being compromised. But it will certainly make it way less likely that someone will steal something out of your car.
Blue check would gatekeep a lot of noble, new developers.
but IMO they'd never charge for such a thing, that's not at all in the spirit of OSS / python dev.