Thoughts on the Python packaging ecosystem
pradyunsg.me
pradyunsg.me
Pradyun's point about unnecessary competition rings especially true to me, and points to (IMO) a hard reality about where the ecosystem needs to go: at some point, there need to be some prescriptions about the one good tool to use for 99.9% of use cases, and with that will probably come some hurt feelings and disregarded technical opinions (including possibly mine!). But that's what needs to happen in order to produce a uniform tooling environment and UX.
Not to mention that all efforts would be focused on improving a single package manager. Even if the one that was standardized wasn't the "best" at first (by whichever metric you prefer), giving everyone an incentive to improve that one will likely make it the best within a few years.
The npm/yarn fork that formed in 2016 is the closest analogy I can think of. My impression is that npm improved on its biggest weaknesses (e.g. a lock file) and has remained the default choice in that ecosystem. I suspect the same would happen with pip if it too made significant improvements.
python3.9 -m pip
python3.10 -m pip
The standard library is the reason why python is popular to begin with.
IMHO it’s absolutely the toolchain’s job to manage that for me. I’d never adopt a toolchain that wouldn’t.
You can just use 3.11 with everything.
Generally speaking, if you're on a modern OS - like recent Ubuntu LTS or MacOS you're getting 3.8 -> 3.10 and compatibility is very good in these releases.
The same goes if my project depends on a library or framework that doesn’t support 3.11.
There’s no doubt it’s the default though.
If poetry (or similar) were renamed to Pip and it were to be included in the stdlib id obviously switch.
It doesn’t solve the single static binary problem or docs problems, but it’s much much closer.
You didn't even understand the point they made: the need for Pipenv and poetry would pretty much go away if pip added support for a proper lockfile and venvs. And that's the only correct choices as pip is already pythons package manager.
If there two steps, one part of the community will insist they belong in: a makefile, bash script, python script, lambda network service, bazel, pants, scons, terraform, ansible, npm …
I spent a long time migrating a project to use poetry. One of the reasons I opted for poetry over others was that the lockfile retained all of the environment markers in the packaging metadata, so that the lockfile could support multiple interpreters and interpreter versions, multiple platforms, etc.
For example, your package might depend on `foo`, which in turn could sniff the host OS and select the appropriate subdependency. You'd then end up pinning that subdependency, which would be incorrect on a different host OS.
(Similarly for Python versions: a subdependency might be required on < 3.7, so re-installing from a lockfile generated from an older Python could produce a spurious runtime dependency.)
Metadata for the ultimate software package should probably include a sufficient number of attributes in its declarative manifest:
Package namespace and name,
Per-file paths and Checksums, and at least one cryptographic signature from the original publisher. Whether the server has signed what was uploaded is irrelevant if it and the files within don't match a publisher signature at upload time?
And then there's the permissions metadata, the ACLs and context labels to support any or all of: SELinux, AppArmor, Flatpak, OpenSnitch, etc.. Neither Python packages nor conda packages nor RPM support specifying permissions and capabilities necessary for operation of downstream packages.
You can change the resolver, but the package metadata would need to include sufficient data elements for Python packaging to be the ideal uni-language package manager imho
And we've always had such tool, the humble "pip". Now it even works great with pyproject.toml. Plus built-in "virtulalenv" if you need isolation.
That shit has to stop. The default needs to be project local installs. Node might have issues but one thing they got right is defaulting to project local installs vs python where you need various incantations to get out of the default "globals" and you need other incantations to switch projects.
However, you'll find that as with all packaging discussions there are people opposing it, because their workflow doesn't match yours and they don't want to change how they work. We, as a community, need a way to resolve such stalemates or I fear we won't make much headway
Unfortunately (or not) that basically means no more python. Because it is definitely a “Jack of all trades, master of none”.
Focusing on programming languages that just want to be programming languages and not also a system service makes life so much better. You end up using languages that produce self-contained, easily shippable binaries, or languages with easily embeddable runtimes, instead of trying to write code that has to somehow survive in a “diverse ecosystem”, which generally makes it overly bloated and brittle as it grows so many appendages to solve so many orthogonal incompatibilities it comes to resemble enterprise open source…
This PEP is meant to ensure you don't clobber your Linux distro's base environment and break key OS applications.
Also, this problem doesn't exist on e.g. Windows, where there is no OS installed python to clobber. Other folks have taught themselves to always use virtual environments for this very reason and therefore don't share your problem.
Hence, there are tutorials out there that don't talk about your problem and tools exist where the default behavior might be dangerous on your Linux distro.
[0] https://realpython.com/python-virtual-environments-a-primer/
In node
cd projectFoo
# do stuff, node uses packages local to projectFoo
cd ../projectBar
# do stuff, node uses packages local to projectBar
There is no "activation", the default is it just works.Installing packages local to the project is also the default.
$ cd projectFoo
$ ./ve/bin/python whatever...
(more realistically, it's `make whatever` which then builds the virtualenv into `./ve` if needed, pip installs required packages into it, and runs the command).Yes, I agree that it would be nice if the default behaviour of `pip install -r requirements.txt` was to install it in an isolated virtualenv specific to that project, but it's not also not like it's completely impossible magic.
But as someone that went through the transition, I’m certainly not going to deny how much of a shit show it was. And I get a feeling that the two situations were/are caused by the same political / organisational / philosophical factors.
I’d be miffed if the Python / PyPA mob got so distracted with their internal politics, certain terse personalities, or the hands-off competition-is-good ideology, that the packaging story makes Python an increasingly unappealing choice. Even worse, we could end up with a packaging ecosystem run by effing Microsoft like the JS people do, because “competition”.
And yes I’m totally conflating aspects of packaging here. At least we all settled on PyPi, except for all those that haven’t… :)
The industry I am in is extremely conservative about its upgrades and changes. Between that, our custom software interfacing with the 800 pound gorilla for our market, and a few other things, we're behind.
Now, this software depends on a very specific version of ArcGIS. Desktop, not Pro. Not just version 10. Not just version 10.2. But version 10.2.1.
If you go look at ESRI's page for Python and 10.2.1, it says, in large and bold letters (which is something ESRI doesn't do in its documentation very often), NOT TO UPGRADE OR CHANGE PYTHON VERSIONS, EVEN A TINY BIT.
And that version is 2.7. I couldn't even get pip working when I wanted to install a very old version of lxml I wanted.
I guess what I am getting at is, as a language becomes successful and has market penetration, it also seeps down, way far down, in the chain of dependencies. It's the price of success, essentially. I think language maintainers should really pay more attention to that. Just as you know how a program ends up sticking around longer than you originally thought, so too do versions of your language. It is in some ways akin to trying to explain to da Vinci that very far in the future, some of his artworks will adorn clothing and coffee mugs: the good stuff just gets to places you would never dream.
I think a Cython extension or a second process that uses ArcGIS, with a simple interface between them, might be a better choice for the future? Never needed that though.
Making things better for dozens of maintainers at the expense of millions of users is the wrong tradeoff.
Apple does something similar (e.g. with Swift and with iOS), offloading a constant support burden onto developers.
1. Installation. Just install everything in a bloody application-specific virtualenv by default and use the existing lockfile mechanism by default.
pyinstall='python -m venv .pip_modules && ./.pip_modules/bin/activate && pip install --editable ./ --constraints=constraints.txt'
2. Lockfiles. Encourage and always always use the constraints.txt-file. Maybe take the oppurtunity to rename it to lock.pip instead. pyfreeze='./pip_modules/bin/pip freeze > lock.pip'
3. Running. Remembering to activate a virtualenv is too easy to forget and leads to mode-confusion. We need an equivalent of npm run. pyrun='./pip_modules/bin/$1'
Done. That took 5 minutes tops. Now if you were to do this just a little more seriously its of course a bit more effort but still within the realms of weeks/months, not years. It seems they are more busy bikeshedding forever about the build-backends or whatever else that only very few packagers cares about. The packaging of course has its issues but they are all minor compared to preserving end-users sanity. The UX is the #1 issue that impacts millions of users daily, it can and just need to be fixed yesterday.My feeling is that that approach would cause dozens of itches to pop up. Things like, “my coworker is on Windows and I’m struggling to come up with step-by-step instructions for her to set up the project.” Those are valid UX issues but likely just the beginning of a rabbit hole.
I think that tackling those issues would inevitably lead to yet another Poetry/Pipenv clone – because with UX, the devil is in the details.
I don't know enough about node, but aren't there at least two or three package managers (npm, yarn, maybe pnpm)? Then there are half a dozen different things for the "build" / transpile / compile stage of frontend work, and
> Pick from N ~equivalent choices is a really bad user experience
this has caused me to bounce off of getting into frontend work several times. It's so aesthetically displeasing that my brain doesn't want to learn it.
I'm not sure about yarn, but pnpm has exactly the same API as npm has. It simply has a different disk organization and caching strategy. If npm maintainers so chose, pnpm's behavior would be an option within npm. Since they haven't chosen that, it seems completely reasonable to "compete" in the way that pnpm does.
I realise you’re talking about pnpm, and not yarn, but when I think of issues like this, yarn is what comes to mind. A bunch of projects prefer yarn over npm, or at the very least place them side by side, in the project documentation. As someone reasonably familiar with JavaScript, but that only works with it sporadically, I always find myself going back and trying to reverse-engineer the reasoning, and I’m never entirely confident.
So when I circle back to my 6-monthly journey into writing some JavaScript, and see another thing called “pnpm” making its way into package installation docs, I cannot be blamed for sighing and being a little irked. Especially since my daily driver is Python, so I’ve got loads of packaging trauma.
You can usually substitute the package manager in the docs with your favorite, but in doubt just use the one they use.
It is de facto standard tool for installing packages. It can be extended to do more.
Also, the standard should favor pure-Python packages, in order to untangle Python from it's legacy as a glue language for modules written in C(++). Otherwise, no progress will be made at the language level and Python will die out once C becomes legacy language, replaced by safer systems programming languages.
We need to learn from Java's and Javascript's ecosystems. Those languages are now used mostly in their pure form. They are portable, optimized and more stable than most languages. With enough care, Python can become the same.
Never gonna happen. Python's explosion has been through data science and ml, which is based entirely around calling Fortran/C/C++/etc.
> Otherwise, no progress will be made at the language level and Python will die out once C becomes legacy language, replaced by safer systems programming languages.
I assume you mean Rust here. Do you think python can't already call rust? It's been able to do that for the better part of a decade.
I mean, this is once of python's main selling points. It doesn't matter what your numeric/scientific library is written in, python can call it, and it's vaguely convenient to write (depending on your taste for comprehensions).
Why would the python ecosystem shoot themselves in the foot like that?
What's Django, then?
"I assume you mean Rust here. Do you think python can't already call rust? It's been able to do that for the better part of a decade."
Python can only call C code. The fact that Rust itself (or Fortran, btw) has good C ABI is on Rust, not Python. Also, by calling Rust's C ABI, you are making a double indirection, thus severely reducing performance.
That's why I said we need to optimize Python, and even make progress towards making Python - in Python itself. Just like Go can compile itself, Python should be able to run itself. But, that is all for the long run.
Pretty rare compared to data work, at least among python projects I've worked on.
> Python can only call C code. The fact that Rust itself (or Fortran, btw) has good C ABI is on Rust, not Python. Also, by calling Rust's C ABI, you are making a double indirection, thus severely reducing performance.
That's a bit of an implementation detail, but fair. There is some overhead. But every language since C has shied away from committing to a stable abi (or at least the C ecosystem is willing to commit to one, the standard mentions nothing about it).
That's only overhead at the boundaries though, and the common pattern of using python as an orchestrator with the vast majority of logic happening inside native code really does minimize the overhead you actually experience. At least you can architect an application that way (same pattern works with things like pyspark as well).
> That's why I said we need to optimize Python
By all means, optimize python. I think you'll run into the GIL pretty fast, but I wasn't objecting to that. I was objecting to making things like pyspark or tensorflow second class citizens. Doing that is going to destroy python.
Why? Except for the C(++) part, which was never particularly accurate in the first place except that that’s the most popular set of lower-level languages to start with, what's thr problem with Python being a glue language?
> Otherwise, no progress will be made at the language level
Clearly false, as progress continues to be made at the language level.
> Python will die out once C becomes legacy language, replaced by safer systems programming languages.
Why? Python works as glue for Rust as well as anything else. If anything, the biggee threat to Python as a glue language is more dev-friendly system languages, not safer ones, but even there I don't see how packaging deemphasizing support for non-Python modules does anything but accelerate Python problems.
> We need to learn from Java's and Javascript's ecosystems. Those languages are now used mostly in their pure form.
They always were, though; both were designed for use cases without any reliable lower-level besides their own VM, and became popular in that environment. That's not what Python’s ecosystem grew on, and arguably what Java and JS teach here is lean into what made you a success.
There's a lot of talk about Python packaging but as a semi-casual user I've actually had very few problems with it.
What I want is for the Pipfile to get merged into the pyproject.toml at this point, so that we can speed up the move to pyproject.toml for everything.
Moved all my projects to Poetry for reasons. But in hindsight, I must say that Pipenv used to work just fine for me, as does Poetry today. No major issues so far with either tool.
What I do miss is being able to `pipenv run myscript`. That UX has degraded slightly for me, having to say `poetry run poe myscript` now. Barely an inconvenience though.
Used to spend much time to optimize Gentoo Linux but eventually switched to Ubuntu to be productive. I feel Conda is similarly useful, especially if not considering Docker i.e. local dev, and am disappointed hackers seem to be dismissing it.
The language needs to stand on it's own to be competitive in the long run.
Pure Python modules are much easier to install and keep updated.
Maybe such a community deserves obscurity.
Python has real problems that need addressing, and that was true 10 years ago as well, packaging included. Pretending this isn't the case makes more serious techies just laugh at Python, and it's fully deserved.
We do tech, for God's sake. We don't do religion. Merit is what should call the shots.
Python also works better in Codex and chatGPT, as an output of AI. Almost all AI research comes in Python, most of it never gets reimplemented, and the cutting edge stuff are usually just in Python. So if you want to try the latest toys - you guessed it - you need Python. Or be prepared to wait for 6-12 months/forever to get ported to your lang, and be prepared for bugs and less support
AI is going to be a differentiator for programming languages. The languages with more code out there, more questions answered on SO, will work better, causing more adoption, it's the rich get richer problem.
I got no horse in the race and I don't care if Python is "going away". It likely will not since it has a huge inertia and devoted fans.
Furthermore, you citing the current state of affairs is not convincing. You're basically saying "the sun is shining now at noon, surely it will keep shining during midnight". Or "right now it's raining so surely it'll keep raining 24/7 for the next year".
Also who predicts stuff for 2030 with 90% confidence? I'd like to have that person's self-esteem because nobody can predict as far into the future.
I've been part of a number of communities and Python's seems dysfunctional and anarchic.
Maybe that's a good thing and breeds creative forces -- the proponents certainly make that case, I heard, and I'm not opposed to the idea, just a tad skeptical. Time will tell, right?
In the meantime, Python is still missing some stability guarantees and a good package manager. As a programmer that's a turnoff, so I work with other languages. Make of that what you will but I'll restate that I got no horse in the race. I'm using my experience to judge if something seems a good fit to work with, and maybe -- does it have a future.
The greatest strength of Python, I think, is that it allows cross-disciplinary collaboration, because it's easy for non-programmers to learn.
Which makes the lack of a functional easy-to-use package manager all the more frustrating.
I've seen people openly admit they'd learn their second language much faster if they didn't try to constantly compare it with Python.
As usual, hindsight is 20/20, and nobody is telling you these things beforehand. And those who are lucky enough to have wisdom shared with them are usually quick to dismiss it and not listen to it.
Why doesn't somebody just fix pip? Or is pipenv just understood to be the fix?
It doesn’t have a lot of features, but its position isn’t one to have those features.
It’s simple ina good way. I don’t need to futz with a bunch of stuff I don’t need to get it going
Being able to install multiple versions of the same module seems like a fairly basic feature for a packaging manager. And it would seem to me that having a shim executable to examine a profile of some kind for an app at startup, and load/install the correct version of python doesn't seem impossible.
If nothing else it seems like a good way to let packages be lax in keeping their dependencies up to date.
Python makes backwards incompatible changes to the language all the time, even in dot releases. This means that you can't just install everything into some shared directories since some code would need version X of Python and some would need version Y. So you need virtualenv or something like that.
But virtualenv isn't enough either because most actually useful Python code calls into C libraries (Pandas, Tensorflow, Matplotlib, etc.) And those need to be managed by something like RPM, deb, or even docker.
This is leaving aside implementation issues. For example, many python dependency management tools are incredibly slow and there are some differences of opinion on how to express dependencies.
Even when it is possible to install bundled libraries, someone has to do the work to set up Python packages that bundle the native libraries. This work is done for some of the most important pacakges like TensorFlow, but probably not for some more esoteric library you want to use. I also suspect that people seriously using TensorFlow probably want careful control over which version they are using, rather than letting a random Python package manage that.
Basically, interacting with native libraries is a hard problem to fully solve for any language. A lot of other languages like Java have struggled with this as well. But at least in Java you seldom use native dependencies. In Python you use them all the time (arguably, serving as glue code for crusty old FORTRAN and C libraries is where Python shines.)
Python's stubborn determination not to standardize on anything for packaging (we have easy_install, conda, pip, pipenv, poetry, and who knows how many others) as well as its insistence on breaking its own compatibility certainly don't help either. And Python is more hostile than most languages to integrating with the system package manager (RPM, deb, etc.) because of how it sprays its files around the filesystem.
I've used pipenv, pyenv, poetry and settled for my own use-cases on just pip with virtualenvwrapper. So any program goes into its dedicated folder with requirements and virtualenv. However there's something that I've never managed to get working. You have say a 3.8 virtualenv, and package file mentions python should be >=3.8. But when calling python -m build, the build process creates a new environment using the system python. So if you are on an old system that is on Python 3.7, the resulting package cannot be installed because of the inconsistency between declared python version and the one that has been used. It may be a very stupid problem, but I haven't been able to find any documentation for it. Any search engine is just too happy to throw any 'python packaging result' before anything so specific. Maybe I should use one other tool for that? But there's no clear (default?) path documented.
Otherwise you'll have to be rather specific about what Python interpreter to use. For example using the python launcher for Linux (https://github.com/brettcannon/python-launcher).
Looking through the docs of pdm, Hatch or Poetry, I can't really find a definitive answer if they will use the Python version you specify. They all at least need to be able to locate the correct Python version in order to do so. It would be great if managing Python versions was something that came along with Python, so that these tools could rely on it more.
Pick from N different tools that do N different things is a good model.
Pick from N ~equivalent choices is a really bad user experience.
Picking a default doesn’t make other approaches illegal. Picking a non-default tool needs a written explanation with sound reasoning
When there is a default, new members to the community will gravitate towards it and, if there isn't a good reason, will avoid learning other options.In time, this ends up being anywhere from a minor annoyance to very frustrating, as not only are you losing time to learning, but potentially losing time to not having access to extras as the community around the default expands and your now special snowflake project goes without.
For very specific tools and libraries, it might be okay if they are easily interchangeable (but is apparently bad UX?) But for things with larger responsibilities - a full blown ORM, package manager, etc- it can end up being a millstone.
Packaging pure Python often boils down to where to place the packages, the hierarchy between them and where to find the starting point for an application (if it is one and not a library).
Insert native extensions into the mix and you get all kinds of issues.
Like, which library to link to? Which version? Which foreign function library to use? Will my Python version be compatible with that library? Et cetera.
Written by one of the maintainers of scipy
Like why does Ruby have gem and bundler, each doing one thing, whereas python has fifty bazillion tools that all do nearly the same thing if you squint but all have their own weird problems?
I've personally just ended up using poetry and that has mostly stopped me from having to care overmuch about the tooling.
I feel like I'm spoiled coming from Java land, where aside from too much XML, maven just works and if you need to do weird shit (like if you're android, for instance) then you use gradle, which still doesn't fuck up the maven repository format and we can all just live happily. (And nobody uses Ant anymore).
In others, it's less weird: Ruby's upswing happened with Rails and a handful of other "killer" frameworks, which helped to solidify developer workflows (and expectations) around packaging. Python, by contrast, has had multiple generations of "killer" usecases, each with their own baggage (and many predating any real packaging standards).
The history of Python packaging is also much older, and much more devolved than anybody in 2023 would consider reasonable for a packaging ecosystem: the earliest generation of PyPI, for example, was just an index that pointed to other webhosts for downloads, rather than a full package host. This helped ossify manual workflows that developers at the time were content with, and some of that cost is still being paid forwards.
Nokogiri was historically a bit difficult to install but that was because of dependencies and not packaging. To overcome that limitation, nokogiri now has prebuilt binaries for most platforms.
Packaging in ruby using bundler is amazing. It also evolved naturally. In the beginning there was only `gem install`. Then came `bundler` with package list and lock files. It had widespread adoption and now it is shipped with ruby.