3500 packages uploaded to PyPI, pointing to a malicious URL
twitter.com
twitter.com
Not sure how we could fix it without slowing way down and doing a lot more work.
I'm sure it's insufficient but you asked who audits all that (re: npm) well, apparently somebody.
https://docs.npmjs.com/cli/v7/commands/npm-audit
The audit command submits a description of the dependencies configured in your project to your default registry and asks for a report of known vulnerabilities. If any vulnerabilities are found, then the impact and appropriate remediation will be calculated. If the fix argument is provided, then remediations will be applied to the package tree.
https://libraries.io/ is a project of theirs I use quite often when vetting third party dependencies for our organization.
Presumably all of those will need to be verified as well, and the problem recurses.
In the past, there wasn't much software or much memory, and software was simple. Now, memory is plentiful, there's too much software, and software is too complex. A central authority doesn't scale. Fixing our naivete about how much to trust software does.
Increasing the trust we are able to place in the tools and libraries we leverage isn't a complete solution either, but it will have to be an element. Currently we've got nothing but blind faith and crossed fingers.
Unfortunately, this is pretty much the only way, besides the "walled garden" approach.
I am often derided as being a "grumpy old control freak." They def have a point, but I am very, very leery of bringing code into my projects that I don't understand, and stepping through source takes almost as long as writing from scratch.
Most dependencies cover a wider range of functionality than the narrow application that we seek in our individual projects, so, to truly understand the dependency, we need to examine all of its faces.
I recently had to yank a dependency out of a project that I'm developing, because it had a regex bug that caused memory leaks and random crashes in worker threads. It would have been way too difficult for me to figure out what was happening, so I removed it, and wrote my own functionality that is not as good, but doesn't crash my application. This is the kind of behavior that gets me scorn. I'm supposed to keep the crash, and shrug.
So, yeah. I guess I am a grumpy old control freak...
Abstractions are leaky. You should strive to understand where your black box is transparent.
I agree, I've never gone as far as reviewing/understanding source code for all the libraries I use. Might be nice to have the time ;-)
>it is _good_ to black box things
I trust that you are talking about using other people's black boxes, as opposed to reinventing your own?
I would say that there is some element of risk in trusting a library that someone else made, particularly over time as systems need to be upgraded. Will there be problems during or after an upgrade? Is the whole dependency chain upgradable? Will the upgrades work on all operating systems in use? Are there relevant corporate policies and/or whitelists that might interfere with or slow down upgrades?
Personally I pay close attention to the dependencies I'm using in a project, especially long-lived projects. I find it super aggravating when something I've tested and deployed breaks a couple years later after routine upgrades due to some kind of drift problem to do with a dependency.
How much benefit am I getting by using the library? Best is using base library that comes with the language. Next best would be a well maintained third-party library with a good reputation in the community, one that doesn't have too many dependencies itself. Am I happy with the api of the library, and its documentation? Is is reasonably easy to approach given its functionality? Is the library available as a package in the operating systems I'm working with? How long has it existed? How stable has it been over time? Is the library a good fit for what I am doing, or do I have to bend my code to use it as intended? Do I have multiple libraries to choose from? Can I use the library in such a way that I can change my mind later?
I think there is value in trying to minimize the use of external libraries, up to a point.
It's not about having the time, but -making- the time. For the most part we simply do not prioritize looking into the source code of the libraries we use, to any meaningful degree. (I'm guilty of this for sure). Hence things like leftpad and the many other instances of dependency controversies.
In an ideal world we [and our employers] would put emphasis on inspecting the huge amounts of code we pull in and allocate time for it.
(Cost to dependencies doesn't have to be just review time. E.g. in my embedded software projects, the calculus for "will we need to update this" looks very differently than it would for your typical web app. And extracting interesting bits from libraries and including only that into the main project is maybe more common, reducing the surface of a dependency - which would be a bad idea if you expect frequent changes, but these projects typically don't)
I don't mind dependencies, but I won't use them for anything mission-critical, unless I spend a great deal of time vetting the dependency. That might mean auditing the code, but more often, it is auditing the coder.
The main thing is, is that we need to really think about anything we add to our precious project, from outside. In my experience, people seem to be alarmingly casual about this.
Good results can be building very big systems, in a very short period of time, using very few resources, and everything works great.
Bad results can be building very big systems, in a very short period of time, using very few resources, and everything works great.
Until it doesn't.
For example, an upgrade introduces things like renamed CSS hooks, or introduces thread safety issues. A PR isn't vetted properly, and license-problematic code, or malicious code is introduced into the system. Maybe the author has a car accident, and can't manage the library anymore. Maybe a dependency down the chain gets taken over by someone that uses...let's say "solarwinds123" as the password for their CI server, and they get pwn3d, but, since they are buried three levels down, no one in the main dependency realizes what happened, and so on. Maybe you get to a point where you outgrow your backend, and need to change to a more robust or secure backend, but the SDK you are using has you addicted to the initial backend that you selected, when the project was still in the crib, etc.
It should not be a casual decision.
In my original post, I talk about having to yank a dependency.
It was a phone-number parser. If you know anything about phone numbers, parsing them is non-trivial. Apple has a built-in utility (a scanner for URLs, phone numbers, addresses, etc.), but that is not as effective as the parser I imported, which had a fairly good pedigree. I looked at the code, and checked out the author, and it looked good.
The issue that I encountered was most likely (can't be sure) a thread-safety issue. I ran parses in non-main threads, and noticed memory leaks, and random crashes.
I do not suffer memory leaks or crashes of any kind, in my apps, so it had to go.
Phone number parsing wasn't a dealbreaker. Because it was a dependency, its application was fairly encapsulated (I always encapsulate external dependencies -just plain old horse sense), so I can go back and maybe reinstate it (if I can fix the issue, or use it only in main threads), or add a different one in the future. In the meantime, the Apple tool is fine for my development.
The backend for the project, on the other hand, is mission critical, so I wrote it myself. I am not a backend expert, but there are some things that I needed assurance on, that no backend engine can give me. It's highly encapsulated and layered, so future efforts can replace it. I can tell you that whatever it gets replaced with, will need to meet my exacting criteria.
My current project has one external dependency (it has quite a few, that I wrote): A keychain abstraction. It's very simple, and is actively maintained. I've used it in other projects for years.
The dependencies that I wrote are pretty intense; even the small ones. The testing code dwarfs the implementation code in all of them. One of them is an SDK for a pretty vast system, where I was the original architect, but has since passed on to a new team that has my complete confidence; so I guess you could also call that an "external" dependency, but one that I know very, very well.
This is a very unsafe assumption, at least until we invent a programming language that disallows bugs of any class.
It's a choice to be ignorant of your dependencies, not a law of nature.
The fact that most people don't audit code they depend upon is a problem to be fixed, unless the underlying code is fixed first.
I do! In fact if we exclude the language implementation itself[0], a plurality of my projects have NaN% of their dependencies fully audited.
0: ie, compiler, hardware, etc; if you object to this, please do explain how to eliminate these dependencies; I'm quite interested.
... to teams you can 100% trust and whom are in the same sphere of responsibility. So yes, take a binary library blindly from the team down the hall who reports to the same VP.
From an unknown rando on the internet, not so much.
"But why would you do that?" and "But thats the point of using those libraries!" right?
when you're the only one that can find the heisenbug and fix it; your effort and the need for it aren't evidence of over-reliance on dependencies. They are an argument for re-implementation from scratch in brand new, clean, safe technologies!
you may have started a howl
This is seldom actually done. More frequently, customers will find bugs, and fix or report them, thus, improving the tool for all users.
In this one case, it was not reasonable for me to do this. The application that I'm developing is closed-source (for now), and I was in the middle of a pretty massive refactoring job. There was no way that I was going to take a couple of days to try to nail down a random memory leak/thread safety bug in someone else's code, when I could simply use the Apple OS tool to do something similar, and I knew it would be safe.
I definitely could have figured it out (given time). I'm a very good debugger, but it would have been an unacceptable branch in my workflow. I have a pretty full dance card, and this was not on it. Thread safety issues can be a pain to track down, and they are like cockroaches; for each one you see, there are a dozen more, behind the drywall.
I wasn't about to raise a ruckus about it. I suspect that the tool works fine in standard main thread apps, and my application uses worker threads that originate from things like network callbacks. The author won't shut down the tool, because one random person with a specific workflow has an issue. I have managed open-source projects, and know what it's like to be on the other end of that crap.
The biggest issue with dependencies is trust. That's why walled gardens are attractive. There's a chain of responsibility, and some kind of accountability baked into them. Even then, some of these libraries can get compromised, or sold off to nefarious actors. Also, I suspect that many of these SDKs and dependencies are really "first hit is free." They can make us dependent upon a person, company, or whatever, that does not have the best interests of our user at heart.
I tend to rely on manufacturer API/toolsets, and even then, I'm really careful. Some of the SDKs that ship with devices and services can be...questionable. But if I find myself blaming the compiler for an issue, I can be reasonably sure that the actual problem is mine. If I am using a blackbox tool that sits between me and the iron, it can be difficult to figure out where the problem lies.
Just a couple of days ago, I had to change the hosting provider that runs the backend for the application that I'm developing. I wrote that backend, and there was definitely some kind of bug, but the hosting provider decided to introduce a new "can't be switched off" inline cache that completely broke interactions with the backend, so I couldn't find my bug. If the backend had been a server that I don't control, it could have taken many more days to even figure out that the bug was in the host, report it (or diagnose it), and then wait for the inevitable back-and-forth before it was addressed. Since my primary work is the frontend app, that would have meant a huge delay. I figured out the cache issue, because I am intimately familiar with the backend, and could easily diagnose problems, using the classic "divide and conquer" methodology. Once the hosting provider was switched, it took me half an hour to find and fix the bug.
Even that assumption is being eroded
> The net effect of it seems to be that a full system for really acceptable programming will be at the same time a full system that will suffice for the description of constructive mathematics.
He said that in '85!? Wow.
[1] https://maniagnosis.crsr.net/2007/10/1985-dijkstra-interview...
That's basically the Curry-Howard correspondence[0], which was first explicitly stated in 1980.
0: https://en.wikipedia.org/wiki/Curry-Howard_correspondence
This is a circular argument. “We chose a language that has millions of tiny dependencies now we need a tool to manage millions of tiny dependencies”
Can create a market for library audits that the library's team would pay for and provide "vetted" tags for libraries that have audits on minor updates.
Things can continue as they are for smaller libraries, but any library that's used in more critical areas should have no issue finding audit sponsors.
This would also create in incentive for somebody to develop automated auditing tools.
You restrict what your dependencies can do in the first place, so that if they're malicious (or just buggy) the scope of what they can do is limited. This doesn't eliminate the risk entirely - after all, it's possible for a library to introduce vulnerabilities just by doing its job incorrectly - but it massively limits the scope of what you'd need to audit. Right now, any dependency can do anything, so you'd need to audit all of them.
See POLA Would Have Prevented the Event-Stream Incident[0] for more explanation and LavaMoat[1] for an example of tooling that's trying to tackle this problem.
Now, this might count as "doing a lot more work" since it's admittedly not quite as simple as just typing in a project name. It's much less work than rewriting everything from scratch though. :)
[0] https://medium.com/agoric/pola-would-have-prevented-the-even...
That's the point of software distribution maintainers. They are actual humans who must essentially sponsor a project before it becomes a package in a repository. They take responsibility for the packages they maintain. Users trust these humans since they usually don't make mistakes like letting literal malware into the software repositories.
Most modern languages are deliberately designed to be incompatible with this model. They optimize for developer ease of use and developers don't like getting approval from other people in order to publish their code. So they make their own little isolated worlds where everyone can upload anything they want.
Honestly it's amazing it took this long for people's trust to be exploited.
For example, with `cargo`, let’s say a library I require passes a security audit. However, because a library it requires, or one of its dependencies requires doesn’t pin down a specific version but instead just requires “latest” or “2.1.”, my security audit is for naught given that malware can slip in any time for 2.1..
That goes for testing too. I’ve tested my software etc, but sometime down the line a transitive dependency updates and adds a bug. Now my software that i haven’t changed, is broken.
Software requirements without pinning is a code smell.
This might work in environments like Node, where each library having their own private versions of their dependencies is acceptable, and Rust/Cargo in some situations, but doesn't work in environments where only one version of a package can be present (e.g. Python) or you care about the total size of the dependencies.
Also if you're auditing, wouldn't you audit all dependencies? What does it matter whether they are pinned by your direct dependencies, wouldn't you yourself pin down the transitive dependencies after auditing them?
What’s the point of a transitive dependency can pull the rug out from under your audio with a single update
I understand the desire for these free-for-all package distributions where every tool has it's own library downloader and anyone can publish anything... but the model is inherently incompatible with security.
On the provider side, there needs to be a strict process of validation and code reviews to get any code changes into a public library. On the consumer side, there needs to be the will to take personal responsibility to curate everything that is imported and minimize dependencies.
There's no shortcut to secure code.
* Not allowing packages with similar names to popular ones
* Not allowing packages creation to be anonymous (in the extreme case you would require to validate your passport or similar)
* Automatic detection of malicious code
* Central auditing organization ...
This is just on top of my head, there must be many more ideas.
"The package only contains __init__.py file, that says:
# the purpose is to make everyone pay attention to software supply chain attacks, because the risks are too great."
> The project contains a setup.py file that sends a request to a malicious URL during installation.
I can't comment on whether the URL is actually malicious or, perhaps, just logging requests for statistics / tracking.
The code reads:
url = "http://101.32.99.28/name?<package_name>"
requests.get(url, timeout=30)
So it looks like the author wants to track the number of installs. Nothing is done with the response value (at least for the setup.py that I saw).Edit: seems to be a Tokyo based IP.
Unrelated tidbit: tencent.com returns an empty reply for me at the moment.
That's because the correct address is https://www.tencent.com/ (prefixed with www.)
To test "correctly", you'd likely want to use the requests library or, at the least, ussend the same User-Agent header that it does.
(And, for all we know, they could just be targeting certain IP addresses... or only responding the "malicious" response the first time the URL is requested per IP... or some other weird conditions that we aren't aware of.)
Salient points being that cupy releases a new named package for each cuda version, so future package names are of course predictable. Since PyPi doesn’t allow namespacing, cupy’s plan is to register new names ASAP when cuda releases a new version and monitor and report other packages purporting to be cupy that get uploaded.
I feel like domain name disputes tilt disproportionately to trademark holders, so I wouldn't like to see, e.g., cupy block a cupyd package or vice versa (or, for that matter, NVIDIA somehow strongarm cupy). On the other hand, you'd like a mechanism by which you can trust that a package comes from the cupy maintainers.
Yet. I suppose it might wind up vaporware, but the feature request is under discussion/development.
BTW, it's PyPI.
1. Host thousands of typo packages that phone home every time they are installed. Store IPs.
2. Find which companies own those IPs, filter out residential
3. Report vulnerabilities
Pypi is not a white list, it’s just an index.
People shouldn’t randomly download packages without understanding and verifying EVERY one.
I think this is lazy devs and users who mixed up the App Store with pypi. These “exploits” aren’t very useful other than helping people understand these problems.
It seems like an easy solution that some company could set up to test and verify packages and create a white list service that I could subscribe to.
I currently use dependency checkers like safety and snyk that will alert me to real vulnerabilities like when someone takes over a package.
That I care about. We’re someone to take over pandas and put malicious code, that would be damaging.
That’s also why I pin all packages to specific versions and don’t update unless necessary.
I think this can be solved by better training.
I worry that these articles are some new form of FUD against Python. So far, I’ve gotten some non-tech security folks ping me about “Python vulnerabilities” because they read these articles. But it’s pretty easy to evaluate if my org is actually at risk by reviewing the dependency graph and seeing that no one is using these 3500 bogus packages.
I also use dependency checkers that monitor al those various packages for CVEs related to particular versions. I think I typically catch them during dev and GitHub does a pretty good job of detecting and even suggesting updates.
There is a risk that a package gets corrupted at a used version, but I think that’s fairly small and would be detected quickly. But I think that’s similar to what would happen if Microsoft or Apple put out a bad update. And pypi is policed better than commercial setups that don’t have checks and balances (eg, SolarWinds)
yes but people do and all the training in the worldn't isn't going to help - the package manager needs safe defaults that prevent this sort of thing.
There are some package managers like R’e CRAN that are really thorough on what gets in, so maybe if pypi got more volunteers they could start a “tested” tier where updated require test suites and whatnot.
Further analysis of PyPI typosquatting (2020) https://lwn.net/Articles/834078/
C/Unix way avoids centralization of power/trust and also avoids implicit assumption that every code should be in one canonical repository.
Packages are namespaced and "tier 2" by default.
Official packages, packages from trusted vendors, might-as-well-be-in-the-stdlib packages, otherwise vetted packages, are upgraded to "tier 1", might drop the namespace.
Project manifests can specify: stdlib only, "tier 1" packages only. The latter sounds like a sane default.
> Get the most from your use of Perl, Python or Tcl and reduce your compliance, legal, and security risks. We offer custom managed and self-serve distributions, including support and maintenance, on Windows, Linux, AIX and more – even for 32-bit and older releases.
Now Linux programmers do that intentionally via package managers.
Outlook users probably didn't, most of the time.
Looks like unneeded legal risks with no reward given people can already reveal their name if they want to.
If you want a curated Python there’s Anaconda and ActiveState.
https://geekfeminism.wikia.org/wiki/Who_is_harmed_by_a_%22Re...