Altough, I guess in that case, the shell call itself is suspicious.
What if the shell call itself is obfuscated?
This is a very common malicious behavior. Packj detects obfuscation [1] as well as spawning of shell commands (exec system call) [2]. I've updated threats.csv to flag code obfuscation.
1. https://github.com/ossillate-inc/packj/blob/main/main.py#L48... 2. https://github.com/ossillate-inc/packj/blob/main/main.py#L48...
The correct approach would seem to be to not run untrusted code in an environment where it can read your AWS credentials.
I imagine you'd want some kind of fuzzer with security-oriented tracing on top. I've never heard of such a tool, but I'd bet it exists somewhere.
For example, instead of writing
open('/etc/password')
One can write method = calculateMethodName()
name = calculateEtcPassword()
geattr(__builtins__, method)(name)Note, it's been around since JDK 1.0, so 1996, 28 years ago (!), though it's not widely used.
But it allows sandboxing of libraries you use. I was using it to make a faulty library that was calling System.exit(0) - yes, a library that was shutting down the entire process, closing the main app... - not do that and instead wrap the exit in a regular exception.
Python and especially Javascript are reinventing a lot of "enterprise" ideas that Java invented 20+ years ago. Maven groupIds are another example from the Java ecosystem (they're used to avoid top level package name squatting).
IMHO it's like the difference between basic research (mostly not very useful on its own but extremely important) and public health in medicine (the thing that actually makes a massive impact on the world that leverages the former but at a slower and more controlled pace). To stretch the analogy even more any doctor worth their salt will tell you that a treatment plan is only as good as the likelihood that the patient will stick to it. Other things like let's encrypt (https everywhere) and signal/whatsapp/telegram (strong privacy enabled by default) come to mind as implementing ideas on things that experts had been handwringing for decades that people "should" be doing that are now done so well that it feels almost silly not to do them the good way.
It would be cool if there was a universal access control platform, but I guess the best we're getting is Docker, because such a platform would commoditize OSes, which OS creators for sure don't want :-(
https://wiki.python.org/moin/SandboxedPython has some details
1. https://github.com/ossillate-inc/packj/blob/main/main.py#L47...
- esprima==4.0.0 requires Python 3.6 (EOL) or lower because the package is really old and uses the async keyword (promoted in py37) as an attribute name, which is a SyntaxError on py37+.
- GitPython==3.1.27 requires Python 3.7 or later (requires-python:>=3.7).
The point of that project is that you can create or use an existing repository proxies and attach to it what I called "audit policies" those are basically a list of packages/versions you want to block or allow. The default ones include for example malicious, vulnerable, yanked packages etc... (the blacklist repository) to which you point pip, poetry etc... and it will block installation of the packages listed in the audit policies attached to the repository. You can also create ad-hoc repositories or repository per project to keep it separate and operate in whitelist mode where you allow only whitelisted&audited packages.
On top of that there is also "monitor" mode where you can allow installation of any package or subsset of packages and it will capture all depedencies for purpose of tracking the software supply chain accross the company or project and those packages would be automatically scanned and audit using integration with another project of mine called Aura that is a static analysis scanner designed for the python supply chain.
As mentioned this is currently in open alpha mode so access is limited and user registration is not open (I am currently working on users&permissions for making their own repositories and audit policies) but if someone is interested in testing or this project in general or an early access to features behind the curtain feel free to shoot me an email at admin @ sourcecode.ai . The license is open source so it can be also self-hosted.
I think certain things would be picked up pretty easily e.g. obfuscated code would be a pretty loud feature, but subtle stuff might be undetected and generally I can't see the model being super accurate.
1. https://github.com/IQTLabs/software-supply-chain-compromises 2. https://github.com/rsc-dev/pypi_malware 3. https://github.com/osssanitizer/maloss/blob/master/malware/R...