Rye and Uv: August Is Harvest Season for Python Packaging
lucumr.pocoo.org
lucumr.pocoo.org
The linked post is the author of Rye's take on that.
I’m only a casual Python user. But wtf was it doing and why did it take so long? That’s bonkers.
Since it was posted a lot of work was done on areas which likely caused performance problems, and I would expect in the latest version of pip to see at least a doubling in performance, e.g. I created a scenario similar to OPs that dropped from 266 seconds to 48 seconds on my machine, and more improvements have been made since then. However OP has never followed up to let us know if it improved.
Now, that's not to say you shouldn't use uv, it's performance is great. But just a lot of volunteer work has been put in over the last year (well before uv was announced) to improve the default Python package install performance. And one last thing:
> for a non-compiler language?
Installing packages from PyPI can involve compiling C, C++, Rust, etc. Python's packaging is very very flexible, and in lots of cases it can take a lot of time.
Home Assistant is an absolute behemoth of a project, especially with regard to dependencies. Dependency resolution across a project of that size is nuts. There are probably few currently projects that’d see as big an improvement aa HA.
I have no idea how they keep version conflicts from breaking everything. Do integrations have isolation or something?
I wish python had a native way to have different libraries in the same project depend on different versions of a transitive dependency, seems like that would make a lot of stuff simpler with big projects.
One problem we have is that support for any repository features beyond PEP-503 (the 'simple' html index) is limited or entirely missing in every repo implementation except warehouse - the software that powers pypi. So if you use artifactory, AWS codeartifact, sonatype nexus, etc, because you are running an internal repository, PEP-658 & PEP-691 support will be missing, and uv runs slower; you may not even have accept-ranges support. (and if you want dependabot, you need to have your repository implement parts of the 'warehouse json api' - https://warehouse.pypa.io/api-reference/json.html - for it to understand your internal packages)
I've been playing with https://github.com/simple-repository/simple-repository-serve... as a proxy to try to fix our internal servers to suck less; it's very small codebase and easy to change. Its internal caching implementation isn't great so I wrapped nginx round it too, using cache agressively and use stale-while-revalidate to reduce round trips, it made our artifactory less painful to use, even with pip.
Further, if a package index supports PEP 658 metadata, pip will use that to resolve and not download the entire wheel.
uv does the same but adds extra optimizations, both clever ones that pip should probably adopt, and ones which involve assumptions that don't strictly comply to the standards, which pip should probably not adopt.
I was talking about the maintainers not wanting to include features a lot of the userbase would like (like monorepo stuff), because they are saying their target audience is package authors, while in fact most of the users aren't.
Would be curious if anyone thinks this is a useful direction, ultimately hope uv/hatch include something like this.
the uv support on workspaces (virtual and concrete) has me intrigued.
Why does it work for FAANG but not for you?
if some smart and dedicated engineers can do the work to build a tool that lets everyone trivially manage a monorepo, that is certainly the best possible situation to end up in.
Package A depends on C 1.0 and B depends on C 2.0. How much work it is to get down to one version of C in your dependencies is up to how different 1.0 & 2.0 is, and how A & B use it. But if you want them resolved, it's up to you to do the engineering to A or B.
The shell scripts I wrote are a painful, but less painful than dealing with Bazel.
One of Python’s great strengths is the belief there should be one, obvious right way to things. This lack of unity in the packing environment is ruining my zen.
But there is no way for you to define which Python version your project is built against.
If you’re building a package you probably need to test against multiple versions. If you are building a “project” that isn’t an installable, distributed package but a bunch of code that is shared with a few developers and does something (run a machine learning model, push itself to a cloud function, generate a report) you probably want to target 1 and only 1 Python version.
As for the mono repo approach are you suggesting to copy paste numpy and pandas into the repo?
It's in-scope for Rye and hatch.
And I agree with that choice. Choosing a python version is as integral a dependency as choosing a version of numpy.
Personally I haven't had many problems with package management in Python. While the ecosystem has some shortcomings (no namespaces!), pip generally works just fine for me.
What really annoys me about Python, is the fact that I cannot easily wrap my application in an executable and ship it somewhere. More often than not, I see git clones and virtualenv creation being done in production, often requiring more connectivity than needed on the target server, and dev dependencies being present on the OS. All in all, that's a horrible idea from a security viewpoint. Until that problem is fixed, I'll prefer different languages for anything that requires some sort of end user/production deployment.
You are not wrong, but let's unpack this. What you're saying is that there is a need to make it easy for another person to run your application. What is needed for that? Well you need a way for the application to make its way to the user and to find some Python there and for that process to be transparent to the user.
That's one of the reasons why I wanted Rye (and uv does the same) to be able to install Python and not in a way where it fucks up your system in the process.
The evolved version of this is to make the whole thing including uv be a thing you can do automatically. You can even today already (if you want to go nuts) have a fully curl to bash installer that installs uv/rye and your app into a temporary location just for your app and never break your user's system.
It would be nice to eventually make that process entire transparent and not require network access, to come with a .msi for windows etc. However the per-requisite for this is that a tool like uv can arbitrarily place a pre-compiled Python and all the dependencies that you need at the right location for the platform of your user.
The cherry on the top that uv could deliver at one point is that fully packaged thing and it will be very nice. But even prior to this, building a command line tool with Python today will no longer be an awful experience for your users which I think is a good first step. Either via uvx or if you want you can hide uv entirely away.
The more we invest into all of this and get it right in at least one way, the easier it will be for alternatives (like pyoxidizer!) to also “just work”. At least that’s my belief.
There are tactical reasons to focus on one strategy, and even if it’s not your favorite… everything being compatible with _something_ is good!
Prior to it, i had my own version of Rye for myself but I only had Python builds that I made for myself and computers I had. Critical mass is important and all these projects can feed into each other.
Before Rye, uv, and pyoxidizer there was little community demand for the standalone python builds or the concept behind it. That has dramatically changed in the last 24 months and I have seen even commercial companies switch over to these builds. This is important progress and it _will_ greatly help making self contained and redistributable Python programs a reality.
So fast lint, type checking, code scans, PR assistants, yes, we can swap these whenever. But install flow & package repo, no.
That is unfortunate given the state of pip and conda... But here we are.
This is how I came to believe this is the case:
Few years ago I wanted to write Python bindings to kubectl. I discovered that in order for that to work cross-platform, I need to make CGo use the same compiler on all platforms as does Python. Unfortunately, on MS Windows, CGO uses MINGW while Python uses MSVC. I wrote to Python dev. mailing list (which still existed at that time) and asked why did they choose to use a proprietary compiler for their "open-source" project. The answer I received in a round-about way was that MSVC was a historical choice, which cannot be presently changed because MS provides Python Foundation with free infrastructure to run CI and builds, and it also provides developers to work on Python (i.e. MS employees get paid by MS to work on Python interpreter). And that they are under orders not to drop MS tools from the toolchain.
Year after year the situation was getting worse. Like in a lot of similar projects, success created a lot of ground for mediocre nobodies to reach positions of power. Python foundation and satellite projects like PyPA started to be populated by people whose way into these positions was not through contributing any useful code, but rather writing pages of code of conduct. This code of conduct and never-ending skirmishes around controlling positions eventually led to some old-timers leaving or being outright kicked out (latest such event was the ban of Tim, the guy who, beside other things, wrote Tim sort, which is a somewhat famous feature of Python).
Year after year MS was pushing its usual agenda they do in every project they get their hands on: add crapload of useless features for the sake of advertising. Make the project swing every way possible, but mostly follow the fashion trends as hard as possible. This is how Python is now devoted to adding as much of ML-style types as possible (in the language with a completely different type system...), AoT compilation and JIT (in the language that's half of the time used to dynamically glue native libraries...) and so on. Essentially, making it a C#, but without curly braces.
MS is smart enough to understand that publicly announcing their ownership of Python will scare a lot of people away from the technology, so they don't advertise it much. But they keep working on ensuring developers' dependency on their tooling, and eventually they will come to collect on their investment.
They’re all so caught up in their internal systems and politics[1] that they don’t seem to know why they’re actually there anymore.
So if someone’s gonna go do an actually-good job and capture the market in the way that Astral is, then that’s exactly what we as a community deserve.
[1]: I mean internal politics. This is a weird alt right “DEI HIRE!” rant.
I'm for small teams doing big pushes, and the example shows such efforts can coexist with healthier long-term governance strategies. It's sad that the effort was needed, but nonetheless, a clear demonstration of separation of concerns working for small engineering team => bigger governance org.
It's also a bad look for the engineering priorities of anaconda inc on their flagship product. I wonder 'how we got there': Anaconda devs help build GPU python stacks, numba, etc, so they are quite good on what they do work on. Is every for-profit packaging-oriented company ultimately not paid enough to keep the packaging core fast? So VC $ just speeds up the schedule to mediocrity?
At the same time pypa wasn’t able to provide a comprehensive solution over the years, and python packaging and development tools multiplied - just 3-4 years ago poetry and pipenv seemed to solve python packaging problems in a way that pip+virtualenv couldn’t.
We need pypa to now jump on the astral.sh ship - but will they do that without a certain amount of control?
As a practical matter, I hugely appreciate the hard technical work PyPA does. However, I don't concern myself much with which toolsets they're recommending at the moment. Use what the community uses and don't worry about the "official" suggestions.
I think the real endorsement that could help would be the core Python project itself. In a perfect world the official Python tutorial would start with "here is how you install python" and it starts by installing uv, the same way as the official Rust docs point to rustup and cargo.
I hope strongly that the PSF will manage to establish some sort of relationship with Astral which would enable that to eventually be a reality.
* Python package installation, package format and loading of modules are defective. The design is bad. It means that no implementation blessed by PSF or not isn't going to solve the packaging problem. So, there's no point to ask PSF or PyPA to adopt any external tool. If the external tools are better than pip in some way, then it could be the speed or memory footprint etc. They will not solve the conceptual problems, because they don't have the authority to do that.
* PyPA and PSF, but maybe to a lesser extent, are populated today by delusional mediocre coders who have no idea where the ship is going or how to stir it. They completely lack vision and understanding of the problems they are to deal with. They add "features", but they don't know if those features are needed, and, in most cases it's just noise and bloat. From a perspective of someone who has to deal with fruits of their labor, they just ensure that my job of "someone who fixes Python packaging issues" will never go away.
So... as "someone who fixes Python packaging issues" I kind of welcome the new level of hell coming from Astral. From where I stand, it is pouring more gas into a big dumpster fire. Just one more tool written in a non-mainstream, non-standardized, quickly evolving language, impossible to debug without a ton of instrumentation, with the source code hard to write and hard to understand. It's just another deposit towards my job security.
Also please consider writing some articles about typical packaging problems in Python that you come across your work.
Please also consider setting up some way of sending you money to support this interesting work!
So, in order to work on the projects listed above I had to study the spec. I won't dwell on all the problems I've discovered, here are just some highlights: the first time my eyebrows began to rise was when I read that Wheel despite being a binary package discourages programmers from putting binary artifacts in it... Yes, you read this correctly.
Now, to elaborate on the matter: when Python interpreter loads Python source files installed in platlib (and possibly some other locations, I haven't researched this subject in-depth) it byte-compiles the sources for future "expedited" loading. At first, the byte-compiled Python code used to be stored alongside the sources, but later it migrated into \_\_pycache\_\_ directory. The unfortunate decision to byte-compile when loading rather than when installing is what back in the days created a lot of frustration for the dummy Python users who would try to install Python packages on their Linux system when running as sudo, but later unable to run their programs because Python interpreter would break trying to put byte-compiled files in a directory not owned by the user running the interpreter. And this is how --user option of pip was born. A history of kludges and bad fixes for a self-inflicted problem.
So... the advise to not put the byte-compiled Python code in the Wheels is what tripped the poor unwitting Python programmers: had they put the byte-compiled files in there, the problem would've gone away, and they could happily install and run their programs in a simple and straight-forward way. Today, the number of kludges around this problem grew by a lot, and simply undoing this advise will not work, but this isn't the point.
Anyways, the motivation for this advise? -- premature optimization. The authors of Wheel format decided to "save space" for programmers publishing their packages. In their mind, if someone published a binary packages, but with... sources in it instead of binaries, that package could be applicable to more than one combination of OS/architecture/Python version. A huge improvement, considering most Python packages are hundreds Kilobytes big! Not to mention that anyone who packages native libraries with Python packages has to packages them for all those combinations anyways. And that's like half of all the useful Python packages.
This is what should've been done instead: copy from Java JARs. Have packages with byte-compiled code, have them used for deployment, and have source packages for those who want editor intellisence etc.
This story is just a drop in a bucket of all the bad decisions made when designing the Wheel format, but for the lack of space and time I will not go further into details.
Similarly, I will only touch on some problems with installation, just to give an example. Python allows specifying arbitrary URL as a package dependency. This makes auditing or even ensuring stable builds a huge issue. I.e. you might think that all the packages you are installing are coming form the PyPI index, unless you configured pip to use something else... but it's possible that some dependency will specify to load from a URL that was convenient for the author submitting the package at that time.
Another problem is what happens when a package for the desired combination of OS/architecture/Python version doesn't exist in the index known to pip: in this case, instead of failing, pip will download the source archive and will try to build the package locally. This means that the users will get the version of the package the authors of the package are guaranteed to never even have run... And, unfortunately, quite often this process "succeeds" in the sense that some package is produced and installed. Infrequently, but still often enough for it to be a problem, such packages will have bugs related to API version mismatch. Some such problems may be sometimes swept under a rug. I've personally encountered a bug that resulted from this situation where some NumPy arrays were assumed to have 16 bit integers but in fact had 32 bit integers. Those arrays represented channels in ECG data (readings from electrodes attached to patient's scalp). The research was made and the paper was published before this hilarity came to light. (You may say that ECG is a borderline scam anyways, but still...)
Now, to the last part: the module loading. The whole reason why Python came up with the kludge of virtual environment is due to how modules are loaded. Python source doesn't have a way of specifying package version when requesting to import a module. Therefore, if more than one version of a package is found on sys.path, there's no telling what will be loaded.
What should've been done instead of virtual environments: Python modules should've only been loaded from platlib (not from the project source directory as it's often done during development). When loading modules, the package info directory should've been examined, and the dependencies from the META file parsed. These dependencies then would be kept in memory and refined every time new module loading request is made to narrow down the selection of versions that could be imported. This would allow Python to install and use multiple versions of the same package in the same Python installation without conflicts. This might not be a huge deal for developers working on their (single) projects, but it would be huge for developers packaging their projects for system use. Today many Python-based projects available on Linux are packaged each with their entire virtual environment and often even the Python interpreter and a bunch of accompanying libraries. I.e. in order to install a project that has single digits Megabytes of useful code, often hundreds of Megabytes of duplicated code are pulled into the system.
As for the money part: dealing with Python problems pays the bills! :) I even sometimes get my name on scientific publications because that's often the kind of projects I have to deal with. So, I have both fame and compensation in good order (at least for now). But thank you for suggestion!
How are they defective?
At this point I also don’t care about nuances like “is it execution issues with PyPA, or that their set remit is faulty?” I’m sick of getting drawn into that stuff, too.
Any tool hoping to dominate the Python packaging landscape must be community-driven and community-controlled, IMO.
(also, how many more decades does this imaginary community need to create a great dominant tool?)
For our projects we use pyenv-virtualenv (so we can have specific python versions per project) and then poetry "just works" (though can be slow, hence rye, uv and friends).
(still a happy Nix user)
Versioning Python isn't hard. pyenv, asdf, mise, now uv... I honestly don't see what Nix brings to the Python ecosystem. I can see using it to version Python if you already use it, but that's it.
On my team we use Nix for actually distributing our Python programs, both directly onto machines for local development, and to build containers for deployment in the cloud. We use Poetry for development and generate the Nix package from our pyproject.toml.
Actually plugging Python into a general-purpose package manager for native dependencies is admittedly a pretty clunky experience today because Python packages lack sufficient metadata and packaging formats within the ecosystem are so fragmented. But with a sane implementation of something like PEP-725, that could actually make that pretty smooth, including for system package managers other than Nix if that's not your cup of tea.
The biggest challenge was dealing with older pacakges that used non-standard packaging and ran arbitrary code; generally ones that didn't have wheels.
From the article:
> As of the most recent release, uv also gained a lot of functionality that previously required Rye such as manipulating pyproject.toml files, workspace support, local package references and script installation. It now also can manage Python installations for you so it's getting much closer.
These are all things that dead project I wrote could do.
Strong +1 on a one tool wins approach - I am so tired of burning time on local dev setup, everything from managing a monorepo (many packages that import into each other) to virtual environments and PYTHONPATH (I’ve been at it for like 8 years now and I still can’t grok how to “just avoid” issues like those across all pkg managers, woof!)
I am really excited to see what’s next. Especially looking for a mypy replacement and perhaps something that gives compiling python a “native” feeling thing we can do
Perhaps people are trying to solve real pain points, and by getting closer to solving them things feel nicer!
Then Astral came out with uv which aims to be a frontend into a collection of their own tools similar to cargo.
Armin and Astral agreed for Astral to take over Rye some time during uv development with (I assume) the goal for uv to fully replace Rye.
Use uv. As of 0.3.0 it covers most of rye now anyway. Especially if you’re writing projects and not consumable libs/apps (I haven’t used uv for anything other than package management so far).
I don't understand how we could lose so much flexibility and yet gain so little in return.
P.S I've only ever encountered minor dependency issues in my admittedly small projects using just pip and venv.
What path would you put in setup.py anyway? A different one for different distro preferences, a different one again for Windows, for macOS?
(Note that I can certainly complain about how `bundler` works in ruby, but these discussions in python-land seem to go way beyond my quibbles with the ruby ecosystem)
Here is a story about the nightmare it takes to compile the Fortran code that supports much of SciPy (backbone numerical computing library with algorithms for wide swaths of disciplines) https://news.ycombinator.com/item?id=38196412
[1]: https://pixi.sh/
It feels like we were driving cars since 50 years and still haven’t figured out a way to distribute gas.
Is there any research going on? The situation is totally crazy, especially for python.
I would like to see this done by top scientists. I would love to never have to spend any time again on the newest packaging tool.
What is the core of this problem and why is it not solved?
As the author of Rye, a package that was heavily influential on the development of uv, Armin is very well positioned to share his opinions about this.