PyPI: Python packets steal AWS keys from users
blog.sonatype.com
blog.sonatype.com
I think certain things would be picked up pretty easily e.g. obfuscated code would be a pretty loud feature, but subtle stuff might be undetected and generally I can't see the model being super accurate.
1. https://github.com/IQTLabs/software-supply-chain-compromises 2. https://github.com/rsc-dev/pypi_malware 3. https://github.com/osssanitizer/maloss/blob/master/malware/R...
The correct approach would seem to be to not run untrusted code in an environment where it can read your AWS credentials.
I imagine you'd want some kind of fuzzer with security-oriented tracing on top. I've never heard of such a tool, but I'd bet it exists somewhere.
For example, instead of writing
open('/etc/password')
One can write method = calculateMethodName()
name = calculateEtcPassword()
geattr(__builtins__, method)(name)Note, it's been around since JDK 1.0, so 1996, 28 years ago (!), though it's not widely used.
But it allows sandboxing of libraries you use. I was using it to make a faulty library that was calling System.exit(0) - yes, a library that was shutting down the entire process, closing the main app... - not do that and instead wrap the exit in a regular exception.
Python and especially Javascript are reinventing a lot of "enterprise" ideas that Java invented 20+ years ago. Maven groupIds are another example from the Java ecosystem (they're used to avoid top level package name squatting).
IMHO it's like the difference between basic research (mostly not very useful on its own but extremely important) and public health in medicine (the thing that actually makes a massive impact on the world that leverages the former but at a slower and more controlled pace). To stretch the analogy even more any doctor worth their salt will tell you that a treatment plan is only as good as the likelihood that the patient will stick to it. Other things like let's encrypt (https everywhere) and signal/whatsapp/telegram (strong privacy enabled by default) come to mind as implementing ideas on things that experts had been handwringing for decades that people "should" be doing that are now done so well that it feels almost silly not to do them the good way.
It would be cool if there was a universal access control platform, but I guess the best we're getting is Docker, because such a platform would commoditize OSes, which OS creators for sure don't want :-(
https://wiki.python.org/moin/SandboxedPython has some details
1. https://github.com/ossillate-inc/packj/blob/main/main.py#L47...
The point of that project is that you can create or use an existing repository proxies and attach to it what I called "audit policies" those are basically a list of packages/versions you want to block or allow. The default ones include for example malicious, vulnerable, yanked packages etc... (the blacklist repository) to which you point pip, poetry etc... and it will block installation of the packages listed in the audit policies attached to the repository. You can also create ad-hoc repositories or repository per project to keep it separate and operate in whitelist mode where you allow only whitelisted&audited packages.
On top of that there is also "monitor" mode where you can allow installation of any package or subsset of packages and it will capture all depedencies for purpose of tracking the software supply chain accross the company or project and those packages would be automatically scanned and audit using integration with another project of mine called Aura that is a static analysis scanner designed for the python supply chain.
As mentioned this is currently in open alpha mode so access is limited and user registration is not open (I am currently working on users&permissions for making their own repositories and audit policies) but if someone is interested in testing or this project in general or an early access to features behind the curtain feel free to shoot me an email at admin @ sourcecode.ai . The license is open source so it can be also self-hosted.
Altough, I guess in that case, the shell call itself is suspicious.
What if the shell call itself is obfuscated?
This is a very common malicious behavior. Packj detects obfuscation [1] as well as spawning of shell commands (exec system call) [2]. I've updated threats.csv to flag code obfuscation.
1. https://github.com/ossillate-inc/packj/blob/main/main.py#L48... 2. https://github.com/ossillate-inc/packj/blob/main/main.py#L48...
- esprima==4.0.0 requires Python 3.6 (EOL) or lower because the package is really old and uses the async keyword (promoted in py37) as an attribute name, which is a SyntaxError on py37+.
- GitPython==3.1.27 requires Python 3.7 or later (requires-python:>=3.7).
PIP is designed to treat all indexes as mirrors, rather than to specify the source of a package: So a higher package version of an internal package name would be chosen whether it is on public or private PyPi. Equally a certain percentage of the time it would choose public over private with the same package version.
How do you mitigate this, other than just totally blocking access to the public PyPi? Shouldn’t everyone be blocking access to PyPi, then, and only be relying on their own private package server?
So one solution I see in the wild a lot is to prefix your company packages with a company identifier, then to set proxy rules to prevent fetching these from public PyPi (often these are set directly on the private PyPi and all traffic is directed through it).
Having been an agency pentester, I can tell you that given access to one company package (Open Sourced?), you can correctly guess the package structure, naming conventions and internal build infra of the company 90+% of the time.
[1] https://arstechnica.com/information-technology/2021/02/suppl...
I like Android's system of per-app uid/gid. But AFAIK it's not implemented by any mainstream Linux kernel or distro.
There's AppArmor, but the last time I tried it, I came away with the opinion that it's not very convenient or user-friendly. Perhaps some kind of friendlier CLI or GUI frontend may help.
I'm assuming SELinux can achieve this but I don't have first-hand experience with it and from what I've read online, it seems to be less user-friendly than even AppArmor.
Any other approach you know about? Do secure distros like Qubes OS or Tails implement this systematically?
I already use Dockerized aliases for some CLI apps (e.g.: ffmpeg) but I didn't find the approach as convenient as I'd like.
I've found docker mounts difficult to secure. I'd like to mount $HOME but exclude even read-only access to "$HOME/.ssh/", "$HOME/passwordsafe.pwsafe3" and a dozen other sensitive file patterns. Some kind of predefined "access profiles" to create FS access rules and assign processes to them ("assign all python processes to python profile") is probably what I'd like.
Containerization is probably the best approach but I'd prefer if it's more opaque and less effort than Docker. For example, if I create a pyenv environment and run its 'python' command, I want that python process to not have full access to the filesystem without having to create container images, command aliases, or volume mounts.
https://github.com/paskozdilar/dockerify
I might try to optimize it a little bit later, perhaps bind-mount dynamic libraries instead of creating a new image for each command.
You have to assume that any code running inside a container has broken outside of its mount namespace & can interact with anything running on the host. Only Linux's traditional mechanisms (credentials; capabilities; SELinux policy; others are available) are able to defend against this.
You can create users manually for each app.
For GUI apps, https://firejail.wordpress.com/
I’d like to see something like OpenBSD’s pledge/unveil.
These all work at the process level, though, not individual portions of code in a process.
If you want something low level look at bubblewrap (brwap). Firejail is also another tool to do this, a bit higher level and with more features (maybe too many IMHO).
SELinux, if it was easier to use!
Apparmor
Bubblewrap?
If anything the problem is that there are too many mechanisms and most users are familiar with none of them...
Wonder if something like pledge and unveil around library code could be helpful; perhaps library code needs to be separated out into a separate process that would not have reason to access AWS keys.
Also, looking at the screenshot, could a simple programming searching for URLs in the library code help in this case?
Looks like they removed the modules, so one can't examine them any more.
Python is tough though, a very dynamic language so probably kind of hard to lock it down.
https://canarytokens.org/generate#
Like any other tools though, i recommend to have a script to trigger it every now and then to make sure it works (and alert you about it so you dont go into panic mode)... for personal stuff, I usually have a specific day in the month i expect to see some canary tokens fire :)
With supply chain attacks much more common, I'm starting to think that egress control is essential for all.
Even if the creds are stolen they’d need access to an instance in your account to use them. Also you can be alerted if someone attempts to use them anywhere else.
Of course, just because it takes the credentials doesn’t mean it does anything else with them, but it could have done anything.
This is not a new idea at all, but it's still a very good one.
An automated tool to check all depency changes for suspicious code and flag that to a human could be valuable. Whether or not that could be done in a way where wading through the false positives is worth it, I'm not sure, but it's a reasonable idea.
They have a lot of expertise in this, it feels like they could branch out in finding malware in source code, instead of in binaries.
It would be better to run the code in a sandbox and log all the syscalls it makes, similar to how Cuckoo's malware sandbox works.
In real-life it is not of much relevance. I haven't seen a practical case in my life where I would conclude that it is undecidable to say what a function does (happy to see some practical examples if someone has some). For a theoretical program that "uploads credentials iff a sub-program halts" you would probably block it and live with a possible false-positive.
Any and all functions that involve templating/plugins/metaprogramming/etc (common in various web frameworks) that can effectively run arbitrary code and thus are undecidable. Also any function that does deserialization and thus will instantiate novel objects with their code, e.g. Python code that uses pickle will often be undecidable as its execution will depend on what exactly is unpickled.
env -i <command>
And only set the ones you absolutely need.https://0bin.net/paste/TTOdW9+B#olQ+FmoCmluokC8UdWllx8GXQtos...
TL;DR
At program initialization, clear the `environ` object, and stash the environment variables in some random location in memory. You can then restore the `environ` object when needed via a contextmanager - meaning that code must explicitly be granted permission to access env vs. being able to snoop regardless.
Curious to hear your thoughts!
EDIT: I wonder if this approach could be extended to allow the use of certain "restricted" libraries (I'm thinking stuff to do with network calls/file system) only within specific scopes - as this would defend against publishing the env vars to a public endpoint...