Python: Please stop screwing over Linux distros
drewdevault.com
drewdevault.com
This also applies to many others programming languages which have their own packaging systems.
While python packaging is indeed messy the needs of traditional linux distributions are by far the least important. Python packaging needs to better serve the needs of python developers, not those of sysadmins.
Having a "linux distro" interpose itself between you and your libraries is a fundamentally broken system not worth fixing. All decisions related to dependency choices fundamentally belongs with upstream.
There are distributions that keep up-to-date, though, e.g., archlinux.
Serving the needs of python developers vs. sysadmins is a false dichotomy. Python developers develop on a system that they need to admin. One great thing about linux is that everything on a system can be kept up-to-date using just one software tool (a package manager). You are not going to convince me in a million years that this is a bad idea.
Now, for things that are needed in development you sometimes need different versions, in particular if you happen to have an OS that ships very old versions. For that purpose there should be some sort of tool and not a gazillion tools that are all incompatible and behave in slightly (or not so slightly) different ways. 'python -m venv' is different from 'virtualenv'? WTF?
Also, if your distribution is up-to-date (e.g., archlinux) and you are pinning to older version there is also a problem. As the article puts it 'pin their dependencies to 10 versions and 6 vulnerabilities ago'. And if you actually want to maintain your software in the future you, at some point, have to go to the new version anyway.
I couldn't disagree more. Even if the same person is admining a system they develop on, they almost certainly aren't going to admin the systems their users deploy on. The admin role and developer role should be completely separate with different goals and requirements.
The system Python should only be used for system scripts. The end. Nowadays ideally it shouldn't even be visible to non-admin user accounts, and users certainly shouldn't be installing packages into the python being used by system scripts, not even using virtual envs.
The python your system uses and how it's configured should be decided by the distro. If they want to break it up into weird packages and funky paths, whatever. That's their problem because they and maybe (_maybe_) system admins should be the only people using it.
As a developer you should be making your own decision about which version of Python you are using, what modules get installed into it, what venvs you have and how you manage them. This should all be considered with one eye firmly fixed on the target deployment environment and how your Python will be configured there. The right answer, of course, being that you should be packaging the required python along with the application.
Yep... Perl is a godsend... take a code from 20 years ago, run it on a modern system, and everything works.
Python? Three different software versions need three different versions of the same library, and new library versions are not backwards compatible with their older versions... Even Python 2.x -> 3.x was a pain in the ass too, and even within minor versions, you sometimes get breaking changes.
That's more or less like "take a VB6 binary from 20 years ago, run it on some modern Windows, and everything works" - that's just because the ecosystem is effectively dead, so supporting it on new releases just means carrying over some stuff that worked 20 years ago.
In 10 years, with python 4, or even maybe 5, everything will still be broken, and you'll still be claiming perl is dead, and i'll still be using the same stuff I use now, that worked years ago, works now, and will work then.
Tell me how I'm supposed to define my lib python deps ?
How am I suppose to package the lib ? Distribute it ? Deal with os versions differences? Allow isolation for my user projects ?
Edit : arf, sorry, I answered the wrong comment, I actually agree with my parent.
- Linux dictros have curated packaging, often outdated but stable.
- PyPi, NPM, and others are Uncurated, anyone can publish, could be unstable, could be insecure.
It's down to the developer to decide which route they want to take, and at the moment most want to move quickly with the latest tools. To do that you have to go the uncurated route. It only becomes an issue for a developer if their software is published by a Linux package management system, but 99.9% of developers will never have that.
Exactly this! It is expected from linux distros that they have curated packaging. I think that is good and I really expect it to stay that way.
Whether you want curated or uncurated packages depends on the use case.
As the user of a program, I definitely want curated packages.
As the developer of a program, I want to specify myself which version of a dependency I want to develop against and I don't want to be hindered by the linux distro in doing that. I do think that developers in that context are not always supported that well in linux distros. And on the other hand, I do think that tools supported by the programming language can assist in that scenario. (for example installing multiple versions of the compiler and runtime in the users home directory and being able to easily switch between those versions on a project basis...)
Please substitute s/sysadmins/users/, and realise your developers are users too.
I've been doing Python for almost 15 years; and I'm getting really fed up with some things. Packaging is a mess. Distribution is a mess (for servers/IoT - Docker saves the day; for desktop - I feel like giving up). Managing the installations is a mess; upgrades can be impossible - I'm hard-stuck on Python3.6 on one project!
I find myself rewriting many smaller tools in Go or Rust, just because I can upgrade the toolchain at any time, and/or ship a static binary. But Rust has a very high barrier to entry, and Go tends to be simplistic.
I'd fully jump ship today, but Python has just too much momentum behind it.
Cadences:
Distribution - e.g. Ubuntu 21.10, 22.04; RHEL 8.4, 8.5, ..., 9
Library - e.g. simplejson 3.17.4, 3.17.5, ..., 3.18.0
Language - e.g. Python 3.8.0, 3.9.0, 3.10.0 (This makes Python particularly annoying because for a C project it shouldn't matter if you used gcc 8 or gcc 9 when producing binaries, but with Python it very much matters which version of Python you run with)
Application - e.g. Gimp 2.8.22, ..., 2.9.0
Roles:
As a sysadmin you want to use the distro package manager to install tools for administering your system so that you can connect to wifi, monitor resources, etc.
As an application developer you want to talk to the language specific package manager so that you can use the latest versions of libraries so that bugs are fixed, and you're not shackled to people using RHEL 6 or 7.
As a Debian/Fedora packager you want to talk to the distro because that is required if you want to submit an application upstream to support users who want everything in the distro package manager.
As an application packager you may want to talk to the distro because of the above, but you could also target flatpak (or snap) so that you can use all the latest libraries without worrying about packaging for slow moving distributions.
As an end user you want to use the distro package manager because that's your embedded mental model and workflow. But you should definitely consider using flatpak (or snap if that floats your boat) so that you can use the release stream wherein the libraries are unpinned from the underlying distribution's cadence.
As an end user you should never need to deal with the language specific package manager to install applications or libraries. If you need to install using the language package manager (npm, gems, rocks, cargo, pip) then the application has not really been 'distributed' or 'packaged' imo. If you're doing this, you're off-piste trying out new stuff.
There are more details and a great conversation that can flow from this, but language specific package managers are not the antithesis of package management and supplement it; just not for the 'end user' role. Well, yes for the 'end user' role, just not directly.
Why not just upgrade applications ASAP? Sometimes things are removed over three or so versions. So you can't move from e.g. 3.9 to 3.10 without making sure you're importing types from the correct location:
https://docs.python.org/release/3.9.0/whatsnew/3.9.html
> Aliases to Abstract Base Classes in the collections module, like collections.Mapping alias to collections.abc.Mapping, are kept for one last release for backward compatibility. They will be removed from Python 3.10.
These things are planned. It's just much much harder to handle this because you need the interpreter at runtime while in compiled languages you just use `-std=C99` and you can compile the code and then link.
All the other observations on this topic (and I've read very many over the years) are akin to the six blind men and the elephant.
I used to think OS package managers should be the end-all be-all, but the use cases for OS package managers are very different from language runtimes. While different Linux distros were fighting between themselves, they completely ignored use cases for projects and language runtimes. Sadly, the result is a mess for everyone.
I think people need to figure out what approach works best for them. I'm of the opinion that any core technology to your business needs to be decoupled from the OS. It makes OS updates too messy and you're tethered to whatever the OS supports. My field uses a lot of Python and every company figured out quickly they need to run their own binaries in addition to packages.
At one company, they packaged up custom RPMs. It wouldn't be a problem to package up and distribute Python libraries. Others had their own package system (no OS or runtime fit their needs). It seems like most people use something like virtualenv.
Regrettably, this means there's no easy answer for people new to Python and the right choice will probably change as you grow. But I really think the answer is Python should come up with something that works for Python and let OSes do their thing.
Whay do you mean by "needs an update"?
> Do you expect someone working in Python in Windows who wants to distribute a datetime package should package for RPM and APT?
Absolutely not. I expect a serious user of that library on an apt based system to package it and submit the package to their distro.
Another scenario is that the OS shipped with Python 2.5 (supported until May 2011), but we had third-party tools that required Python 2.7 (shipped July 2010). Switching OSes (where things like monitor or hardware drivers weren't yet supported) was a ridiculous pain to test and certify. Decoupling OS and Python+package versioning was a huge relief for everyone, but won't make sense for everyone.
> Absolutely not. I expect a serious user of that library on an apt based system to package it and submit the package to their distro.
I try my best to personally do this and push for a work culture that does this, but even if this was done I can't fathom waiting on an OS update for existing code to percolate down. The risk tolerance, scope of concern, and agility between an OS and whatever project pays the bills are very different.
No. As a user I want dependency management (and all of software distribution, to be honest) to be handled by the party that's best able to keep things working while at the same time keeping them secure. Linux distributions have a much, much better track record at that than most upstreams.
It’s essentially like version locking packages except some random Debian maintainer decides when it’s time to update.
They are more stable because I can keep using the same version for two years, and I'm not being pushed to the latest version that has (intentional or unintentional) breaking changes every two months. Yes, there might be a bug or two in there that have since been fixed, but I very much prefer the failure I know over unexpected failures.
They are secure because Debian (and distros like it) backport security fixes to their packages. You can argue about whether they do a good enough job keeping up with vulnerabilities, but at least I know that once I install the update from Debian, my machine is secure, and I don't have to wait for the upstream authors of all software on my machine to release updates that upgrade their dependency.
> It’s essentially like version locking packages except some random Debian maintainer decides when it’s time to update.
Yes, but version locking isn't my problem. The crucial difference is that distros pick a version and support those for years, while upstreams usually force you to use the latest version all the time to get security support. With distros _I_ get to decide when I upgrade, and the reduced frequency is a nice bonus. Having a single entity for all software on the system is also valuable, as there's just one tool to learn and one place to check for updates.
The combination of this 2 aspects is what provides better stability and better security.
> It’s essentially like version locking packages except some random Debian maintainer decides when it’s time to update.
Not at all.
Distro-supplied interpreters and their associated libraries are there for the applications supplied and supported by the distribution. Unless you are developing something to be part of the distro, they are not for you.
Do not let the distro get between you and your libraries. Supply your own.
. . . forgetting the original purpose and mindset behind Linux in the first place. From a 1990's point of view; it was always intended to be a hobbyist OS; and in most cases, one had to compile all the binaries and kernel one's self.
There was no such thing, really, as a "user" or "admin" - everyone was considered, and expected to be a developer.
Distros solve a real problem, but the trade-off is that some parts of the system must exist to serve itself.
Absolutely not. Most distributions ship versions that are released before the distro release freeze.
The whole point of using a distro is to have a reliable and trusted development and production platform.
Use Fedora, or Arch, or another developer-oriented distro.
I want to easily and safely use some app my distribution ships. I want to receive security updates automatically for all such apps. I don't care what language it's written in or what its dependencies are.
These app packages provided by the distribution have dependencies that are also packaged by the distribution so that dependency resolution works.
Since the point of a distribution is that it can run apps, the value is that a distribution works at all.
But as a rant this, like most of the ‘but just fix it’ rants, fails to acknowledge the hugely diverging needs of different users. I could not live without conda, since it’s the only sane way to get a working recent geospatial stack. Others need to run embedded environments, or portable ones, some need long term stability while others need bleeding edge packages that haven’t been released yet. Solving for a single case is straightforward; solving for all of them , not so much.
If I can make an observation - it has nothing to do with the number of maintainers. The problem is deep, cultural and occupies a difficult space where it might be a bug or a feature.
The root cause here is that the Python project, and surrounding community, have little real respect for backwards compatibility. The complexity of Python setups is driven by the need to run multiple - potentially even mutually incompatible in the case of 2.x v. 3.x - versions of Python.
All languages have packaging problems, but Python is unique in my experience in the sheer number of Python installs that I need to manage simultaneously. I still have C code that works from around the time that I learned C. I'd need another Python environment installed to say the same thing about Python.
The Python project went absolutely above and beyond to support users who wanted to drag out making not particularly complex changes to their codebase for over a decade.
During this same time they improved Python 3 in response to feedback and among other effects, made it less different and easier to port, from Python 2.
Many common packages supported 2.7 up to its end-of-life as well.
One of the reasons that Python is often a source of compatibility errors is that both distros and large standalone applications embraced Python in the early 2000s, became dependent on a particular version, and then refused to work with newer versions.
Python is not responsible for all the engineering decisions everyone writing in the language has ever made.
I think what it makes it more difficult in the case of Python is that it has decades of legacy to deal with. No consistent semantic versioning, packages that expect that they can modify their package path in-place (this is a nightmare for immutable systems, like NixOS or OSTree-based system like Silverblue), a wide variety of build systems that sometimes hook into make, etc. Solving this is a hard problem.
This is why an authority like PSF has to step in and say: this is how it is going to be done from here onwards.
[1] E.g. Rust's Cargo.
I’ll add that most of Linux distro packaging contributors are generally very nice people, understand the problem at hand, and are very open to collaboration. But sometimes you see this kind of “it’s all your fault” complaints and it’s doing exactly the opposite of helping the cause.
Available Packages
Name : rust
Version : 1.56.1
Release : 1.fc34Yes, Python packages often make poor assumptions about what setup.py can do (i.e., _anything_), and so you end up choosing between "tested, supported by the author, and old" or "untested, unsupported, but up to date".
Some crates have the same issue where build scripts rely on outside tooling being installed, but it's definitely not common to (unless you're relying on compiling C/C++ code for FFI for example, in which case it's somewhat frequent).
Consider that, even if you want to use packages only from your application's virtualenv, the default (footgun warning!) is that Python will still use the "system" packages -- this means you may have installed Ansible or some other tool that relies on Python and many packages from the distro package manager. But your app could pick up one of those dependencies! At best, this will work fine. But in the worse cases, perhaps it subtly behaves differently or simply does not work at all.
My understanding is that Rust will, by default, statically link all of these dependencies. This, in Python, would be like a "pex" or "par" (or one of the many other options :^)), which does make the distribution aspect much simpler. (At the cost of build-time complexity, slowdowns, and occasional incompatibility.)
With Ruby, you are expected to update frequently. New ruby versions are eagerly awaited and all the major packages are updated pretty quickly.
Rust's Cargo is 20 years newer than Python and benefits from those decades of experience. It's a very different proposition to start a new system from fresh than to try to migrate a huge and diverse community towards it.
Distros ripped it out.
They added the ensurepip module so users could add it after the fact.
Distros ripped it out too.
Arguably the advantage newer languages like Rust and Go have in this regard is they don't even consider the distro use case - you're going to get static linking and you'd better like it. Whereas Python is from an older era and tries to fit in with the local customs and so gets hammered for its inconsistency
I can literally take a 20+yo book and all the examples still work. CPAN still works. I literally have 20+yo scripts still copied from server to server, from laptop to laptop, without any changes.
As one of the sibling commenters mentions, there are good examples that it is possible to standardize packaging better. Maven replaced IDE-driven builds and Ant for in the Java ecosystem and added proper package management. Additionally, it required that projects start conforming to standardized layouts, by taking convention over configuration and being largely declarative. I think the Maven success story lost some of its shine with Gradle, but that's another story.
Are you the end user of the Python code?
Yes -> Is available in your distro?
Yes->Use package manager
No->Use pip install in user mode
No -> Create virtualenv with the Python version you want (including pypy!) and do your pip thing there
Some extreme use cases may benefit from anaconda, but personally I've never needed to use it. My only pain point is dealing with legacy code that relies on PYTHONPATH. Nothing good ever starts setting PYTHONPATH.You might benefit from using Pipx in this case: https://pypa.github.io/pipx/
Pipx is good for the case of "I want to run a standalone Python application that is available through Pip, but not my system's package repo." This is a more common case than you might think.
It's a sensible alternative to `pip install --user`, and having self-contained deps for tool is a bit like `npm install --global` or even `volta install`.
It doesn't address the greater issue tho: that it's getting harder and harder for distributions to package things right, and provider packages for their users (evidenced by the fact that you need a second package manager just for python stuff).
Distros have a hard job, but at the same time programming language tooling devs have more "customers" than just distro maintainers.
This is a great recipe for disaster. Whatever you install in user mode will shadow anything installed system-wide, so when you try to run some system-wide project, it may now fail. I'm also not a fan of how it drops scripts into `./.local/bin`, since that's where I keep my own script, and is version controlled.
The installation will also be frozen and never get updated -- unless you remember to do it manually.
Finally, and worst of all, this leaves you in a dead end if your packages have conflicting dependencies, which is too often the case in Python-land.
I used to just use pip to install to the system. Months/years later I would try to untangle the mess of packages I was just playing with, what the OS wanted/needed, I got those conflicting dependencies you mention, etc. I usually ended up reinstalling the OS. At the time I may not have been as knowledgeable about where the OS package manager keeps packages vs pip--but the whole thing wasn't very user-friendly either.
For years I've been installing into user knowing I can just blow it away. I've dabbled with virtualenv, but it's such a pain to set up and activate. If I have a few projects with similar libraries it's more of a pain to set them all up and switch around. If I end up using a script for something important, I just spend the extra time at that point to "package" it.
Anaconda is a distro, and conda is a package manager, that works across OS platforms and hardware architectures, and installs cleanly into userland without requiring admin privileges. The only way we achieve this difficult goal is by creating a distro and build system that creates "portable" packages that can be relocated/relinked at install-time.
Ultimately, Python's challenges in this department come from the fact that it has such great integration with low-level C/C++ libraries. This gives it super powers as duct tape/glue language, but it also drags it down into the packaging tech debt of C/C++. Hmm... maybe I should write that blog post: "Python Packaging Isn't The Problem; C/C++ Is." :-)
* Some distro software uses python. Let the package manager take care of dependencies for that.
* For everything else, use an dedicated virtualenv for each codebase you are working with.
> I used to just use pip to install to the system
Never do this, for the reasons you cite. > I've dabbled with virtualenv, but it's such a pain to set up and activate
Setup for virtualenv: "python3 -B -m venv venv". Have a shell alias 'alias v=". venv/bin/activate"' that allows you to activate it if you need to install libraries or access a shell. "pip install blah" for library install. That should be all you need. > If I have a few projects with similar libraries
> it's more of a pain to set them all up and switch around
Have a think about why you feel this way, and whether you could mitigate the problems.Here is what I do. Once my libraries are installed for the current project, I rarely activate venv in the current shell. Rather, for each python project, I have a bash script "app" in the root of the project, and a dedicated "venv" directory.
The app script does the following: (1) sources the local venv; (2) does pip freeze > requirements.txt to capture any dependency changes; (3) launches the project. Often I will have multiple launchers in that script, with all of them commented out except for the active one. Be in a habit of always launching from that app script.
To reiterate the approach above, whenever you sit down to write some python code, ensure that you have a dedicated venv for it, and that you are only ever launching code from that local venv.
I have spoken to developers who get upset at the extra hard disk overhead. You don't need to optimise for hard disk usage. Hard disk space is almost free.
I don't bother creating setup.py files, except for the odd occasion that I want to publish code to pip. Good luck.
That's sounds like the general approach I take for "projects" even toy projects. My day jobs have never fit the virtualenv use-case. So at home I often have to look up how to use it. It's so rare that when I make an alias I even forget those.
Most new things are one-off scripts; move or rename some files, extract data from something, or pull from a resource. Something that requires libraries or is too big for a shell script. For example, the last one I see in my bin is a web scraper for appointments. It pulls a website, fills out a form, and gets the result a few times--about 70 lines. What's annoying is sourcing some environment just to run this one tool.
Most people have a directory of scripts (a mix of shell, Perl, or Python) they use if they spend a lot of time at a commandline. It's quite a pain to source the environment just to run a quick script. That's generally the libraries I install into user. I don't care much about the version and troubleshoot things as they come up.
I set PYTHONPATH, but the code in that directory is solely small debugging utils of mine that I want available in every Python interpreter, and I make sure not to put anything more complex in there.
It doesn't have to be. You can have a launcher for a project that sets PYTHONPATH just when you launch that project.
What is bad practice is to be setting PYTHONPATH in your .bashrc, and for the reason you give - that makes it global across python launches.
For example your application might interact with command-line tools written in Python, and unless you delete PYTHONPATH from the environment variables prior to launching any subprocesses, they'll inherit it. This could lead to subtle and confusing breakage.
My pain point in particular with PYTHONPATH (or playing with sys.path) is that people tend to use it with the only purpose of making import lines shorter, which brings naming collisions of all sorts when you aren't creative enough.
PYTHONPATH is simple and obvious how to use, and is similar to using LD_LIBRARY_PATH and friends.
And agreed, there are two separate use cases: development and using the software.
But, are distros creating too much work for themselves by trying to package every itty-bitty python library (and for that matter, every npm library)? Are distros doing anything more than scanning CVE databases with the library versions, or are they _actually_ auditing the versions they choose? (Not that there's much choice, since python also has a shitty story when it comes to backwards compatibility; if you're going with 3.10, there's possibly only one version of a given library that will work.)
Java has a commonly used "fat jar" approach which rolls up all dependencies into a single file. It's excellent. In the python world, this doesn't exist, because virtualenvs aren't portable. If that can be fixed (perhaps a specific section in requirements.txt that captures anything that needs to compile C for the platform) then a distributable virtualenv would become possible. Distros would then scan the application for vulnerabilities (via requirement.txt's manifest), build the distributable-virtualenv, and ship _that_. Python library maintainers don't have to do anything different (except, of course, use the standard way to declare dependencies).
Debian Developer here. Part of packaging work, for Python libraries or anything else, is to verify the reliability of the upstream developers, audit the code, set hardening flags, add sandboxing and so on.
I spotted and reported vulnerability myself and it's not uncommon.
But come on, there are 300k entries on pypi, 200k more for perl and 160k for ruby. I'm not even counting the whooping 1.3M on npm because I assume this is considered taboo at this point.
You cannot package 0.0001% of that, not to mentions updates.
And unless distros make it as easy to package and distribute deb/rpm/etc than it is to use the native package manager, this means distro packaging will never be attractive to most users because:
- they don't have access to most packages
- the provided packages are obsolete
- packages authors have no way to easily provide an update to the users
- it's very hard to isolate projects with various versions of the same libs, or incompatible libs
And that's not even mentioning that:
- package authors may not have the right to use apt/dnf on their machine.
- libs may be just a bunch of files in a git repo, which pip/gem/npm support installing from
- this is not compatible with anaconda, which has a huge corporate presence
- this is not compatible with heroku/databrics/pythonanywhere, etc
- this is hard to make it work with jupyter kernels
Now let's say all those issues are solved. A miracle happens. Sorry, 47 miracles happen.
That would force the users to create a special case for each distros, then for mac, then for windows. I have a limited amount of time and energy, I'm not going to learn and code for 10 different packaging systems.
It's not that we want to screw over linux distros. It's that it's not practical, nor economically viable to use the provided infra. The world is not suddenly going to slow down, vulnerabilities will not stop creeping up, managers will not stop to ask you to use $last_weird_things. This ship has sailed. We won't stay stuff only published 5 years ago with delays of months for every updates.
Nuitka can also apparently compile an entire dependency tree into one binary.
I don't know what I'm doing differently.
1. Server applications that run in a dedicated environment.
2. Tools you write and run just on your machine (or some virtualenv, whatever).
3. Redistributable cli or desktop applications which end users will install and use.
For the first two types, you should never have any issues with Python and its dependency situation. You pin everything, and that's it.
For the third kind tho, it's complete pain. Different distros ship different Python versions, so you need to support all of them. You also have to consider that dependencies can't be an EXACT version, you have to support a range of them, and a variety of combinations.
And then, one dependency has a version that works in Python3.6, and another for 3.9. But they had an API change, so which one do you use? It'll break for half your users either way. Of maybe just put some `if version <= 3.6` all over the place, like we did during the py2->py3 transition?
If a distro ships python 3.6 and the app wants to use 3.7, then the end result must include python 3.7 as well, either by distro being capable of having both versions at the same time or the app needs to ignore the distro-python and ship its own version in the package.
Good luck trying to get such a bundle packaged into any mainstream distribution.
https://github.com/openstack/pbr https://github.com/codrsquad/setupmeta
What they should just do is offer a bunch of packages like python3.7, python3.8 that install the official python package wholesale into /usr/python or someplace and then symlink one of them to `python`.
If I would get to redesign package management (both for Linux distros and for languages), I would have one package manager that installs everything into a global package cache, and then pick the correct version of libraries at run time (for Python: at import time). Get rid of the requirement that there is only one stable (minor) version of a package in the distribution at one time. This has become unworkable. Instead, make it easy to get bleeding edge versions into the repositories. They can be installed side by side and only picked up by the things that actually use them.
> If I would get to redesign package management (both for Linux distros and for languages), I would have one package manager that installs everything into a global package cache, and then pick the correct version of libraries at run time (for Python: at import time). Get rid of the requirement that there is only one stable (minor) version of a package in the distribution at one time. This has become unworkable. Instead, make it easy to get bleeding edge versions into the repositories. They can be installed side by side and only picked up by the things that actually use them.
You may want to check out Guix and Nix - their approach is pretty close to what you're describing.
A common solution to this is if you still want to run traditional distros is to just run "bare infra" (whatever that means) on the host OS and everything else in containers or Nix.
I think this requirement made sense when disk space was scarce.
I think this requirement makes sense if you trust that your distro is always better at choosing the 'best' version of a dependency that some software should use than the software author.
Nowadays, I think neither is generally true. Disk is plentiful, distro packages are almost always far more out of date than the software's original source, and allowing authors to ship software with specific pinned dependency versions reduces bugs caused by dependency changes and makes providing support for software (and especially reproducing end-user's issues) significantly easier.
Isolating dependencies by application, with linking to avoid outright duplication of identical versions (a la pnpm's approach for JS: https://pnpm.io/) is the way to go I think. Honestly, it feels like the way it's already gone, and it's just that the distros are fighting it tooth & nail whilst that approach takes over regardless.
Ah JS, how many days has it been since the last weekly "compromised npm package infecting everything" problem? If you are upholding that as gold standard you have to be the worlds laziest black hat.
> Disk is plentiful,
I recently had to install a chrome snap because it is the new IE6 and everyone is all over chrome exclusive APIs as if they were the new ActiveX. Over a gigabyte of dependencies for one application and given the trend of browser based desktop applications? I would like to have space left for my data after installing the programs I need for work.
Distros assume responsibility for fixing major bugs and security vulnerabilities in the packages they ship. Old versions often contain bugs and vulnerabilities that new versions don't. Distros have two choices here: either ship the new version and remove the old version, or backport the fix to the old version.
Continuing to ship the old version without the fix is not an option -- even if you also ship the new version -- because some programs will inevitably use the old version and then the distro will be on the hook for any resulting hacks. Backporting every fix to every version that ever shipped is also not a realistic project.
Here in the startup world we often forget that there's a whole other market where many people would gladly accept 3-year-old versions in exchange for a guarantee of security fixes for 5-10 years. Someone needs to cater to this market, and the (non-rolling) distros perform that thankless task because individual developers won't.
I think they should just ship Python programs, not libraries. They could check what libraries given Python program uses are safe in the version that it uses them.
And just don't care if each of Python programs has a separate copy of the libraries or if particular version of particular library is shared between Python programs by Python environment.
Distributions might just give up responsibility for sharing Python packages between Python programs without giving up the responsibility for security of those programs.
What makes you think so? SSDs aren't exactly stellar in the cost-per-TB department, as will be the case with each new higher-performance storage technology. Plenty of people cannot afford the prices of new Western tech either, what about them?
First of all, 1TB for binaries and libraries may as well be infinite. Secondly, you can get a 1TB SSD for under $100, which is pretty damned inexpensive when you consider it took until 2009 to get HDDs that affordable.
There you have it: You measure in TB, not gigabytes, not megabytes.
Python packages are megabytes.
No, the main reason is security. I need a distro to guarantee me that the library that I use are going to stay the same for the next 3 to 5 years, while also receiving small targeted patches.
> I think this requirement makes sense if you trust that your distro is always better at choosing the 'best' version of a dependency that some software should use than the software author.
No, it just has to be better at choosing version than picking them randomly using pip.
Furthermore, when thousands of developers use the same combination of libraries from a distro the stack receives a ton of testing.
The same cannot be said when each developer pick a random set of versions.
I believe Nix and Guix also work this way.
Agreed, especially on Windows.
It just works.
This is pushing it. It's not hard to break conda or put yourself in situation where the updater/dependency checker gets stuck and doesn't know what to do, especially once you start adding conda-forge packages. But it does do a better job than anything else I've tried (although poetry + pyenv on Linux is getting much better)
For starters, there is no /usr/local symlinking process. It's also possible to have multiple versions of e.g. python installed and active. Homebrew is like a poor-man's Nix.
I do not have a good idea, but other ecosystems evolved much more sane in the realm of packaging. While not ideal, Go has done a fairly good job - and the "module" operations are instant - which they should be.
Sketch for python: Create a ~/.cache/python/packages directory. Manage all dependencies there. Make the python interpreter "package aware" so that required dependencies are read off a file from the current project (e.g. "py.mod") and adjust "system path" accordingly and transparently. Or something along those lines.
No extra tool, a single location, an easy to explain workflow (add a py.mod file, add deps there with versions, etc).
I'm just thinking out loud, but it does not need to be hard.
Are you saying "put the different version of every dependency you need in there if you have to"?
Because I don't think package managers are ready for that, they usually like to have one version installed per package.
The HN bubble can be amazing sometimes.
now the build tooling still needs some work around versions but its only a minor problem generally as much as it annoys me personally.
You can both allow different versions of the same packages to coexist while also managing updates and installation/removal of software. It doesn't have to be this way. Software should be able to ship with its dependencies included and work and not rely on the whims of the OS getting it right.
I’m not exactly sure how it works but I think I’ve heard that newer releases of Enterprise Linux (EL8+) support multiple channels of the same package or something similar.
Essentially what Bundler has done for Ruby since pretty much forever.
I would love to see this pattern elsewhere.
???
$ apt-file search /usr/bin/venv | wc -l 0
I do agree with you however that python3.9 should be moved into libpython3.9-stdlib, and hence be available to all installations.
However, I thought the point was helping distro package management, which, to my knowledge, is not really built to support multiple installed versions of a package at the time: `dnf upgrade` for example, will upgrade all single instances of each packages to their newer release.
Yes you can use Anaconda if you want, and people who do that are probably data scientists or something and know what they want to do and why. It's well documented and has it's own robust ecosystem.
I say this as someone who's been on Macs at home since 2007 and works professionally on Linux, but I started with Python on Windows back in 2002.
The you have powershell permission system.
It's a ton of fun.
Lol, no. There is one in the Windows appstore too, which has issues with privileges. And I'm almost sure now MS ships it in some other tool too.
> I would have one package manager that installs everything into a global package cache, […]
There is exactly one package manager. If you're on Debian or Ubuntu, it's dpkg. If you're on RedHat, it's rpm. If you're on Arch, it's pacman. Yes, some of the BSDs have two (base packages + ports tree), but they're the odd ones out.
pip, cargo, go*, etc. are not the same thing. I know they're called that but they don't perform the same function: none of them can create a working system installation. Let's call them module managers to have a distinct label.
> and then pick the correct version of libraries at run time (for Python: at import time).
That's easy for Python, and incredibly hard for a whole lot of other things. A module manager can do that. A package manager needs to work for a variety of code and ecosystems. Could it try to do it where possible? Maybe. But then the behavior is not uniform and made harder for users to understand. Could it still be worth it? Sure. But not obviously so.
I would also say that this is just giving up on trying to keep a reasonable ecosystem. It's not impossible to reduce the dependency hell that some things have devolved into. It just needs interest in doing so, and discipline while effecting it. I'd really prefer not giving up on this.
> Get rid of the requirement that there is only one stable (minor) version of a package in the distribution at one time. This has become unworkable.
This is to some degree why distros are breaking apart Python. Some bits are easy to install in parallel, some aren't. There can only be one "python". Worse, there can only be one "libpython3.9.so.1.0".
> Distros, please stop screwing over Python packaging. It is incredible that Debian/Ubuntu pick Python apart and put modules into different packages. They even create a fake venv command that tells you to please install python-venv.
They're trying to achieve the very goals you're describing. Trying to give you a working python without having to download and "install" some weird thing somewhere else. And at the same time trying to keep the module managers working when they're replacing some module but not all of them.
On a subjective level, it's obvious you have a strong distaste for this ("they even create a fake") — but could you please make objective arguments how and why this breaks things? If you're getting an incomplete Python installation, that seems like a packaging bug the distro needs to fix. Is that it? Or are there other issues?
> If I would get to redesign package management (both for Linux distros and for languages),
Well, and now we're here: https://xkcd.com/927/
And, I'm sorry to say this, but your post does not convey to me the existence of any essential C codebase packaging knowledge on your end. I don't know about other ecosystems, but I have done packaging work on C codebases (with Python bindings no less), and you don't seem to be aware of very basic issues like header/library mismatches and runtime portability.
If you are interested in this topic, please familiarize yourself with the world of existing package managers, the problems they run into, and how they solve them. There's a lot to learn there, and they're quite distinct from each other on some fronts too. Some problems are still unsolved even.
Meanwhile when I use Rust all these things are taken care of in cargo. It is part of the language. There is one right way to do things and the way is supported by comfortable tooling, that works so well that you literally don't even consider thinking about anything else.
The way python does dependencies is totally unpythonic. The fact that it is 2021 and this isn't fixed or at least the number one priority of things that need fixing sheds a dark shadow on the whole language – a language that I like to use.
Poetry is good. But it isn't as good as cargo, because it also has to deal with all the legacy cruft. To run code developed with poetry on a non-poetry system you have to figure out all the ways of dealing with envs, paths and such.
Issues like these get me a little fired up, because the collective brainpower wasted on something that should have been elegantly solved in one place is gigantic.
Now I know how to do this, so this is not a problem. My complaint was mostly, that this was a waste of time.
Needless to say, I will only use the distro package manager these days. I know the versions are (probably) compatible, maintainers will usually backport security patches, etc. You get none of that using whatever flavor of the week python package manager.
The distro package managers are probably the best place for that, but bridging the gap between them and the python ecosystem is an obvious challenge.
I'm not a Python dev; I just needed to run a tool that didn't have any other alternatives.
Meanwhile, I can download scripts written in a range of other languages and just fire 1-2 commands and the thing will work.
With Python? Almost never.
Tip: in such situations I start looking at the CI scripts in their repositories. Not ideal but gets me through!
- sudo apt install pip3
- pip3 install --user --upgrade pipenv
(In workdir):
- pipenv install --three package
- pipenv run package --option
Works like a charm and doesn't mess with my system.
So for anyone reading this in the future: don't try to use Pyenv to install Conda. Pyenv tries to set up shims for every binary in the Conda env, which will likely break your PATH.
Pyenv supports installing Conda because Anaconda used to be "just" a Python distribution.
They can otherwise coexist without trouble on the same system.
It does this so that it can set up its shim to handle any executable that gets installed in any Pyenv-managed environment.
This is how Pyenv creates the "foobar is not available in your current environment, but is available in x.y.z" message. It's also a much more reliable solution than trying to explicitly whitelist every possible script that might get installed.
The problem is that this was only designed to work for Python executables and scripts installed by Pip. Conda environments can contain a lot more than that; it's not hard to end up with an entire C compiler toolchain in there (possibly even both GCC and LLVM) or even Coreutils.
If Pyenv detects `bin/gcc` in a Conda env, it will set up a system-wide shim for GCC, which no longer passes the `gcc` command along to the OS, but intercepts it, only to inform you that no such command exists in the current env!
So it's not that Pyenv hoses Conda envs. It's that Pyenv can hose PATH if you have it manage a Conda installation, and if that Conda installation ends up with non-Python stuff in `/bin`.
Obviously I don't know what exactly was broken when you tried to set up that application. But this particular adverse interaction bit me at work a few years ago, and ever since then I have insisted that Pyenv should never manage a Conda installation.
I think that's a reasonable policy anyway, in light of the facts that:
1) Conda isn't really a "Python distribution" anymore.
2) The Pyenv installer just runs the opaque Conda installer script and there's basically no way to control the version that gets installed.
3) They are different tools that serve different purposes and it doesn't make sense to have one manage the other anyway.
4) You probably shouldn't use the Python that's installed in the base Conda environment anyway. You need that to run Conda itself, and you want to keep the list of requirements small to make sure that updates can progress cleanly. It's basically the same as any Linux package manager like APT. Except of course, those tools don't generally support "environments" other than chroot.
Do you have an example of a situation when the two tools had issues? I've been using both for years, and don't remember having any problems.
Is this any different from any other programming language ecosystem? Is Python really doing worse than Node, Ruby, Perl, Lua, Go, or Haskell in this regard?
Python files go in `/usr/lib/pythonx.y`, Python finds said files, programs run.
Yes, Python build and packaging in general is messy. But I am curious why, specifically from a distro perspective, it's any worse than anything else.
- python is more active than most alternatives, you have new packages created every day.
- python is massively used outside of the web, unlike JS, ruby or PHP, that are 99% web. You get Python in SIG softwares, data analytics, automation, pen testing, sysadmin, biology, etc. It's a huge graph.
- python is used by the distro themselves to code features of the OS. E.G: you remove Python, there is no yum.
- python has a rich compiled extensions ecosystem, produced from c, c++, fortran, and assembly. It's very complicated to ship them.
- it's much more common to have several Python installed than for other dynamic languages. So isolation matters even more.
So the difference is the sheer size of the problem.
That's a big one to omit. But alright, I'll play.
It has much less packages, it's almost never used on windows, it has stopped being popular 10 years ago, it doesn't have anaconda because it's not a "corporate tech", you rarely install several versions of perl on the same system, you don't make enough projects with perl to justify one isolation per project, nobody moved to perl 6 so the all CPAN transition never had to happened, perl is not used to script DB/GIS systems/3D engines/IDE, perl is not used by millions of non coders (geographers, mathematicians, physicists, bankers, etc) that have no idea how their machine work.
But to be honest, the nail in the cofin is that distros decided not to split perl and cpan in separate packages. In fact, in somes distros, perl and cpan are already installed and ready to be used.
So for linux:
- python: you must decide among several python, then intall the right package to use pip and venv, isolate your install with venv. The procedure is different for windows. Also 2.7 is a thing.
- perl: you have one perl that hasn't change for years, no new packages, it's already installed, cpan is installed as well, and it's not gonna break your system if you use it. You don't care about windows. Perl 5 forever.
But I don't think any of it justifies the ire towards Python, its community, or its devs that was expressed in the blog post.
I also do want to point out that there are quite a lot of general-purpose CLI tools written in Perl, Ruby, and Node.
I get all my dependencies from Debian and they all work, when I need something that is not yet packaged, I use pip.
What are people doing to get all this issues? I don't understand...
Edit: Ah, apparently the steps to reproduce are: do "apt install python-pip3"; do "apt install python3.8"; when pip3 complains that it's outdated, update it with the command it itself suggests.
What about working on a project with a team?
What about deploying the code on another machine?
Yes, that is in the context of a giant and very successful team using this approach, believe or not. :)
Again, it's all a question of point of view, what we see as a package manager problem and causes us to keep reinventing packages managers, might actually be a problem with how we maintain our packages, my point of view being the latter. But I'm digressing.
When it comes to installing on "another machine", you don't know what Python they have, you don't know what libc they have, and so on, that is exactly what containers attempt to mitigate, so that seems exactly like the tool to use for this problem.
Trying to get the script to run on other OS's than just Debian (or Linux).
Even when the things are not Python-based, I often use a virtual environment, managed with virtualenvwrapper (installed on the distro level). One example is a Terraform config that relies on some Python tools to manage deployments - the Python tools are local to that virtual environment and not usable anywhere else.
If I need to develop something that'll need to run with the distro directly (something that could be distributed as ditro packages) I'd use a virtual environment and tailor the package versions in requirements to the ones available in the distro. This way, multiple distros can be addressed with multiple environments pointing to the same source directory and multiple requirements files for tests.
Indeed you haven't. Worse, you've actively damaged Python's efforts to improve. I mostly work on the JVM these days, and I think one of the main reasons dependency management there is so gloriously simple and effective is that the Debian packagers weren't around to fuck it up.
Here's some work I and others did earlier this year, which I thought was a great example of folks from the core Python packaging world and folks from distros working together: https://www.python.org/dev/peps/pep-0668/ See the massive table of use cases for all the things we had to think about.
You might note that one of the things it needs is additional participation on the Discourse thread from the authors (like myself) and from other distros. Again. It's mostly work done by volunteers, and there's a lot to work through. There's no magic to it.
I can tell you that the following (from TFA) will absolutely make things worse for distros, though:
> I call on the PSF to sit down for some serious, sober engineering work to fix this problem. Draw up a list of the use-cases you need to support, pick the most promising initiative, and put in the hours to make it work properly, today and tomorrow. Design something you can stick with and make stable for the next 30 years. If you have to break some hearts, fine. [...] These PEPs are designed to tolerate the proliferation of build systems, which is exactly what needs to stop. Python ought to stop trying to avoid hurting anyone’s feelings and pick one.
If you want the PSF to fund some engineering work on its own that it can finish much faster than any volunteer packager can even read the proposal, break some hearts, stop proliferating build systems, and hurt people's feelings, they will absolutely do that and say "Distro packaging is not supported, we only support virtualenvs. Users should make virtualenvs. Distro software should ship virtualenvs. Installing a Python package systemwide is meaningless." That's clearly the best-supported option right now, and it's a surprisingly technically defensible answer, but it's not going to make you happy.
The specific model that distros could adopt, if we went this route, is that each Python package builds into a .whl, they're build-dependencies of applications, and applications install .whls into a virtualenv at build time. You'd still restrict packages to come from the distro with the usual policies (built from source code, compliant with licensing, not contacting the network at build time, etc. etc.).
So, for instance, installing "python-somelib" would get you a /usr/share/python-wheels/somelib-1.0.whl, built from source. The build process of "someapp" would create a /usr/share/someapp/venv, pip install that wheel into that virtualenv, and then symlink /usr/bin/someapp to /usr/share/someapp/venv/bin/someapp.
The PEP goes into more details about why virtualenvs are recommended, and I can give you a whole host of subtle reasons, but that's not really the point - the point is that it's defensible, not that it's perfect, and that it's very easy to implement with what works today. So if you ask the PSF to come up with something that magically solves the problems and makes people sad if necessary, this is literally what they're going to come up with. If you don't like that outcome (and there are good reasons not to like it!), then you shouldn't ask for them to arbitrarily pick an outcome some people won't like.
So I guess the problem actually lies in the python library ecosystem that's becoming a npm like dependency hell.
The relevant question here is probably whether the library ecosystem is like that because the packaging tooling sucks or is it the other way around?
- Almost all distros don't just supply one version of a package. They try to avoid it, but I can't recall one with a hard rule against it. For instance, Debian packages multiple versions of autoconf https://packages.debian.org/search?keywords=autoconf2 , the Linux kernel, etc.
- One reason that distros try not to install multiple versions of a package is that it's hard to specify which one you want. If you have, say, requests 1.0 and 2.0 installed, which one does "import requests" get you? The virtualenv approach, where "import requests" simply does not work in an un-virtualenv'd Python, avoids this problem entirely. Each application can independently build-depend on python-requests-1 or python-requests-2, and each user can create their own virtualenv and pip install /usr/share/wheels/requests-1.whl or requests-2.whl as they prefer.
- Even if you do not have different versions of the same library, it may well be the case that you have two different libraries with the same importable name. As a great example, see CJ Wright's talk from PackagingCon last week "Will the Real Slugify Please Stand Up" https://pretalx.com/packagingcon-2021/talk/P3983F/ (I expect they'll post videos online soon). tl;dr there are three packages you can get via "import slugify", and they expose different APIs.
- The possibility you haven't accounted for is "The library ecosystem is like that because things are fast-moving, because people have actual problems they want to solve, and upgrading dependencies and sorting out conflicts is work." It would be great for that work to be done, but we go back to the problem I mentioned at the top - limited volunteer time. In the absence of time to engineer things perfectly, your options are to ship something that's engineered imperfectly or decide not to ship it. We went through the dark ages of "We'll ship the next Debian release when all the bugs are solved" over a decade ago, and it turns out that this doesn't actually help users in any way.
- Debian/Ubuntu Python packaging is actually pretty comprehensive, and Ubuntu LTS ships with reasonably updated (although admittedly not latest revision) third-party packages that are usable out of the box, and that I could just list in Ansible to have usable environments spun up.
- Python packaging is hardly a mess. I have been using pip for ages without any issues other than forcing wheel downloads for unusual distros (like Alpine, where musl makes it chancy to use some low-level stuff). But if you're on a mainstream distro with modern pip, wheels just work.
- If you're not on Linux (or not on the mainstream), pyenv also just works. I have been using it across several years of macOS releases without any significant issues other than knowing to pass it the required build flags to build out the Cocoa bindings here and then (which is easy to do with the brew pyenv).
And, finally, I'm constantly shocked at the number of people who just don't get virtualenvs, or who don't know how to switch python interpreters by using environment variables.
I've never looked back, and they were instrumental in bridging the gap between 2.7 and 3.5 while I converted some code across (I never really found that the switch to Python 3 was as dramatic as many people made it, perhaps because the code I handled worked after a single pass with 2to3 and minor tweaks).
- brew install: fails probably
- pip or pip3 or whatever: fails probably, if it succeeds, breaks something else
- look for program-specific install programs/instructions (like the aws cli for example): maybe works, probably bombs
Python's reputation (and sales pitch) among developers is a language for people that don't want to program or learn software development. It's basically the new BASIC.
Their packaging and installers only reinforce this reputation.
Not that it is sweet roses in javaland or many other language ecosystems. Dependency graphs are complicated, because they are graphs when people want them to be simple trees.
As for virtualenvs, why does a user of software that happens to be python need to know that?
Again, I don’t see a lot of logic or understanding of the toolchain in that comment.
Language ecosystem packaging maintainers say that distro package managers make it hard for them.
The year is 2021 and there is no work towards synthesis. It will be endless, fruitless yelling from each side.
Maybe 2022 will bring change. I'm not holding my breath.
To anyone who finds themselves on a single "side" in this argument: if you have ever said "why do you need to do that" as an accusation instead of with curiosity, you exemplify the problem.
Java sometimes has these issues, but mostly distros aren't going around splitting up fatjars for desktop applications or wars for Web applications.
C# just gets shoved into opt as it was considered a problem child to be left alone for a long time.
Perl has avoided this problem by stagnating at the cost of its long term userbase.
C/C++ mostly get along fine since the distro package model was built for them, but occasionally someone gets upset about meson or ninja or something.
These are your tools:
* pip
* pip-tools
* pipenv
* poetry
* virtualenv
* venv
* setuptools
* pyenv
* requirements.txt
* etc.
These are your outputs: * wheel (binary),
* tarball (source).
The tools in use are completely irrelevant and only add noise to the discussions around packaging. So what actually is the problem here? Are projects delivering broken outputs (ie. bad packages)? Are they not delivering outputs at all? Those are real problems but pointing at the number of tools that exist is not helping.Native dependencies are not really handled except in an ad-hoc way, either.
If you want to import Python packages into a real packaging system, you are confronted with the tools whether you want to be or not.
That is the real problem IMHO. Most python users like to pin dependencies so that "their program don't break", and that's also the reason why so much effort is put in what they call "correct" dependency resolution resulting in the creation of new python package managers that all do "more correct" dependency resolution and make programs "break less" and at the same time require less efforts from the maintainers to actually maintain their code by doing dependency upgrades. And then one day you want to upgrade a dependency and you realize you're 10 releases behind on 10 dependencies, and what was supposed to be a quick maintenance task is now 100 maintenance task.
If you don't pin, your program will break one day or another, a user might open an issue with the traceback, or, your program will break directly in CI where you will see the traceback. Upgrade it, or contribute to the dependency, but just go ahead and fix it, instead of being defensive and trying to have dependency resolution that "doesn't break". At the same time you'll be adopting the actual practice of "Continuous Integration", of your dependencies, which has a better cost/benefit ratio.
I always avoid pinning dependencies, I try to make pip just install the latest version of everything, I am willing to contribute upgrade fixes to any Python package I use, but some dependencies do pin which breaks my own aggressive Continuous Integration practice, heck, I'd even need an option for pip to ignore version resolution at all so that I can make all my contributions to upgrade everything I use. I even remember when I had CI test matrix with all combinations of versions of everything, I don't do that anymore, I just support the latest of everything, we can always have the latest python with containers anyway so that's not even a blocker anymore. If you're not a "techbro using containers", it'd be fine too because you should then be able to make your distro packages at any point in time and expect all of them to work together, minus the delta of the handful of upgrades that are pending to release here and there.
This is an very quick route to a maintenance nightmare imo.
If you have totally unpinned dependencies, and you come back to a project after a year untouched, or 5 years, and it no longer works - which dependency update broke it?
I don't agree that using an outdated package is necessarily a problem at all. Some versions are done! You don't need the latest version of every possible package. You don't necessarily need to update _ever_ (which is why this differs from CI). These updates are often entirely unnecessary churn.
There absolutely are vulnerabilities in some old versions, and those updates are necessary (but tooling & notifications to easily handle this have dramatically improved in recent years, especially on GitHub). There will also be vulnerabilities in new packages though, which may be unknown, and will often not exist in older much simpler versions.
Using a well-tested version of a dependency that does exactly what you need is not less secure than chasing the latest version at all times without a specific reason.
I've found manually updating packages on the rare occasions where relevant vulnerabilities arise, and using existing working versions without changes the rest of the time has been perfectly effective over many years now, and avoid the shifting sands of external dependencies wherever possible means that a project that worked 5 years ago still works _exactly_ the same today.
I rather see more software go in this direction, valuing reproducibility & known correctness (i.e. with isolated pinned dependencies, in some form) over 'always be latest' dependency updates and the complex & hard to reproduce bugs that those shifting dependency interactions can create.
In this case, it doesn't matter to me which dependency update broke it, what matters to me is to have the tests passing again with all dependencies.
> I don't agree that using an outdated package is necessarily a problem at all. Some versions are done!
If a version is done then why is there a new release? Using old versions is tech debt that will one day blow up and cost much more to correct than if it had been corrected over the time, not to mention the security risk.
> You don't necessarily need to update _ever_
Another dependency might decide to use that dependency in a newer version, in which case aren't we all better off using the latest versions of everything? The cost is some effort, the benefit is more features, security, performance, less bugs, basically a better program.
> There will also be vulnerabilities in new packages though, which may be unknown, and will often not exist in older much simpler versions.
Then why upgrade at all when you have a dependency with a security issue? After all, in your upgrade you might be adding even more unknown security issues that might be even more dangerous.
> Using a well-tested version of a dependency that does exactly what you need is not less secure than chasing the latest version at all times without a specific reason.
If I can test well that a newer version of a dependency works for me, why not upgrade it? There might be performance, security or other bugfixes, and I'm allowing other maintainers of other dependencies to also use that newer version.
> I rather see more software go in this direction, valuing reproducibility & known correctness (i.e. with isolated pinned dependencies, in some form) over 'always be latest' dependency updates and the complex & hard to reproduce bugs that those shifting dependency interactions can create.
Ok, but then again, some version might fix a security bug that has not been backported to your old version, especially if it's 5 years old.
Basically you're advocating against continuous integration (I'm talking about the "practice", not talking about "the tool that runs automated tests that people call CI").
Knowing which change (or changes) broke it can make resolving the issue much faster.
> If a version is done then why is there a new release? Using old versions is tech debt that will one day blow up and cost much more to correct than if it had been corrected over the time
Because <new feature> was added, that is totally irrelevant to your use case.
> not to mention the security risk.
As parent mentioned, there is automated tooling for this. Your tooling yells that you are using a version of a package with a vulnerability, so you update.
> Another dependency might decide to use that dependency in a newer version, in which case aren't we all better off using the latest versions of everything?
If you don't update A, it doesn't matter if a newer version of A wants a newer version of B.
> The cost is some effort, the benefit is more features, security, performance, less bugs, basically a better program.
This assumes bugs and vulnerabilities decrease monotonically over time. This isn't true.
> Then why upgrade at all when you have a dependency with a security issue? After all, in your upgrade you might be adding even more unknown security issues that might be even more dangerous.
Because a known (to the world) security flaw is orders of magnitude more dangerous than an unknown (to the world) one, all else being equal. If there is a CVE for it, there are likely large-scale attempts at exploiting it anywhere it can be found.
> If I can test well that a newer version of a dependency works for me, why not upgrade it? There might be performance, security or other bugfixes, and I'm allowing other maintainers of other dependencies to also use that newer version.
Large amounts of labor. Furthermore, "well-tested" may include "battle tested". Some bugs make it through to deployment, and get caught and fixed. Updating dependencies without a good reason means more potential bugs slipping through, which means more bugs being discovered in deployment and a worse experience for the end user.
* Create a new conda environment with the minimum python version I need to support.
* Use poetry to install dependency
* Package the production build into container image for deployment.
The only problem I have is sometimes Poetry choke on compiling libraries with C-dependency, luckily most libraries provide binary wheel now.
Then rinse and repeat for red hat, arch, nix, mac and windows ?
Yeah right.
Also, unless you repackage thousands of compiled c extensions and you play well with anaconda and can plug into the entire python ecosystem of platforms such as heroku, python anywhere, databrick and so on, you then lose 90%
I've seen a few comments here about how Nix/NixOS fixes the whole python binary/library mess, but I'm having trouble understanding how. Does anyone have any insight to share about that?
Additionally, the whole thing kind of makes me want to move away from python wholesale. I was wondering if there are other languages that are great general languages like python that don't suffer from this whole packaging and versioning mess. Ruby? Go? Something else? I'm looking for something high level, somewhat easy to learn, and with good library support for things like working with databases and tabular data. Though I don't know much about them, I just feel like I don't want something like Java or C++ or anything like that. I want to "get things done" and not have to worry about tons of boilerplate or working at really low nitty gritty levels.
Nix is a bit of a cult. Its theoretical aims are laudable, but in order to get there it forces you to do a lot of work and reason strictly in its own way. Whether all this work is worth the rewards, I think is open for debate.
It's easy to learn coming from Python, even more so if you can start with Scala 3 right away. You cannot really get more high-level than that. There are good libraries to work with databases, Quill comes to mind if you want something user-friendly. And you can always fall back to the many Java libraries in the big data ecosystem (even though you should probably avoid it if you can).
For personal projects on my desktop, I have given up and just use pip install, and then don't try to share those scripts.
Our project gets a few issues opened by distro maintainers who have trouble with some part of their build process. Is it really worth our project maintainers time to help troubleshoot esoteric build processes when we already provide source, wheel, and conda distributions?
“Just learn $language_package_manager_of_the_day, and hope it doesn't break anything when your system changes” is the equivalent of the 90s “./configure ; make ; make install”, and we should have moved past that for end-users by now. Most users are not developers, and installation procedures need to cater to both.
If you’re distributing an application, shouldn’t you just ship an environment with our package bundled?
No-no-no-no! If you (application author & distributor) do so, you are obligated to release new version of your application (which includes 3rd party libraries as "environment") each time your dependencies are updated due to security reasons. Do you know many application maintainers who want to do so?
If your application depends on system-installed (distro-provided) libraries (python modules, native libraries, whatever) it is responsibility of distro maintainers do not have versions with known security problems in system-wide repo. It is much better for you (app author & mainitainer) and me (distro user & admin), IMHO.
That's how I would prefer it, but, if the intended form of distribution is a distro package, then it should pin itself to the versions the distro provides and avoid (or vendor in) packages that aren't available.
It is possible to just place everything in the app directory and distribute it this way, but it's kind of ugly (and doesn't pick up security updates from the distro)
As a matter of principle, I prefer to avoid software from non-curated repositories like pip and the like. Installing the debian-provided scipy and numpy is more than enough for me.
I don't understand what's so special about python that needs its own "package manager" when the distro-provided one is already good. If I need something, I install it using "apt install", regardless of what language it is written in.
For example, if gnuradio requires numpy then a distro that won't distribute numpy can't distribute gnuradio
* setuptools/pip/poetry is package managers
* egg/wheel/? are package formats
* venv/conda are environment to isolate from the OS
another degree of complexity is added by the OS: their package managers and ability to install as root vs under user. But in any layer, there's not much to get confused.Over 12 years with Python the only serious change I saw was setuptools (was it named like that?) => pip in 2011. The only serious issue I saw was when I installed with several managers (OS, setup.py, pip) and/or as root/user. That got solved in 15-30 minutes after checking the imported package __path__, and cleaning things up.
Teaching people at courses, I saw people who struggled the most never checked anything: doesn't work? They'd try installing harder, or just re-ran things. But they never checked where the packages were installed, and where from they were imported.
Otherwise, I came to following simple rules:
1) python & ipython are installed via an OS package
2) packages that have to be executed like jupyter may be installed as root, or user; just never install both.
3) all other packages are installed as user. Computers are personal, and I almost never execute things as root, nor have other users.
Maybe being able to examine sources of errors was why I didn't feel it was so hard? I'm not a sysadmin, nor a hardcore programmer. It's just I got used to check paths and read error messages attentively.
This does not mean it's easy, I wish things were simpler, there were less ways to do things, but the OP exaggerates things as if there's a thousand of tools.
The fact that it is possible to get all of these confused is in itself a problem I'd say :(
Switch languages, this is a lost cause.
At the core of this is implicit cooperation between system programs, third party programs, and users. This concept underlies a common frustration with Linux: It works reliably out-of-the-box, but customization and installing programs introduces problems.
I wrote a Python dependency and version / installation Manager in Rust to help deal with this sort of thing, as well as related issues like dep conflicts between Python projects.
Right, but then just use pip install --user instead of a virtualenv. Actually --user is the default now when running pip install from non-root user, so, just pip install as a user will work and not mess with the system packages, just don't do sudo pip install.
This seems almost as bad as system python. I suppose it's fine if you only work on one thing, but as soon as you don't, your dev environment will become chaotic and lots of confusing WFM will happen. E.g. this is why npm has a separate node_modules folder for each project.
This sacrifies time and disk space and hides tech debt under the carpet, I'd rather have a solution like `pip upgrade` that would upgrade all packages and fix the environment, like `pacman -Syu`, but people would have to stop pinning versions and actually maintain their codebases and the dependencies they uses.
Not to mention that installing every dependency in the same environment is a recipe for both disaster with version conflicts and bloat when you don't really know which packages belong to which applications.
However, installing every dependency in the same environment is also what distro package managers do, I don't see that as a recipe for disaster, it all depends how well the said packages are maintained in which case just upgrading them should fix whatever problem.
Go ahead and install pip3 from Python 3.6. use that to install pip3 for 3.8, and try using apt. Backup first.
You'll get pip3 from Python 3.6.
apt install python3.8
You'll get system-level python3.
Use pip3 in 3.6 to install the latest pip3 that requires Python 3.8, and drops that package in the dist-packages folder, which is shared between system Python runtime versions.
Behold as further attempts to do something sane with APT blow up, because now you have 3.8 packages on 3.6's runtime path.
I've lost weeks to this particular brand of distro daftness. I thought that surely no one would do that .. Alas...
As far as I can tell, they're a strict requirement, if you ever work on more than one project and want to isolate the dependencies (because you need different versions of the same package, because you want to ensure you've tracked all deps, whatever).
Is it that you don't isolate dependencies between your python projects at all, or is there some other solution you prefer?
In summary: bin/run-app -> bin/venv-python -> .venv/bin/python -- I'm not sure how to make it simpler than this. The cost is perhaps in disk space, but I don't care about that.
Or don't do the above and deploy via docker and use pip to install to the system python in the image.
With Perl modules (like Python packages), the intent is to never have to rewrite them, but to make them so generic that they are just extended by new modules (that inherit the old modules) to gain more features. Their naming convention reinforces this intention, with "Net.pm", "Net/IMAP.pm", "Net/IMAP/SSL.pm", etc. Each one of those files can be uploaded to CPAN by a completely different developer, without having to completely re-write all the base stuff from scratch. You can already do that with Python, but because everyone names their packages "fizzbaggy", "unclib2", "OtherThingHere", etc, there's no sane convention that clearly tells you what this thing is, what it does, or what it depends on.
The end result of all that is a nightmare in terms of regular users figuring out how to install and use most Python scripts, packages, and tools. This is part of why Go is so popular now: no need for an engineering degree just to run a Python program!
I do not believe there is any way for Python to re-invent itself and suddenly become less sprawling or confusing. The shift would be too huge and take too long.
One thing that people do like about JavaScript is that you can import multiple versions of the same package. This solves that depA requires 1.0 and depB requires 2.0.
However this makes auditing nearly impossible. Small projects can end up with well over a thousand dependencies that is just unmanageable.
IMO, the Java or APT packaging ecosystems solve the issue in a far superior way.
The shared dependency model is just too complex to work with, and we have enough disk space that its not really necessary anymore. Sandboxes seem like the way to go.
Actually, this is not sane.
If you have any Python dependencies, you should always develop and deploy your code using virtualenv, never by installing packages into the system Python!
1) If your distro requires Python, don't put it on the path. Refer to it another way if you need it, e.g. have a distroname-system-python package you upgrade at your own cadence.
2) That's it. Then developers can install Python how they like, and it's (probably) all fine.
Of course you still have problems on how to declare those dependencies and resolve upgrades but a code you compiled 10 years ago will probably still work today. Compare that to Python or even worse nodejs and it's bonkers the amount of context you have to be aware of just to make your code that worked fine 2 years ago build again on the current tool chain.
edit: typos
The thing is, people in development wont use the distro python (but different installs, pyenv, etc), and people in devops wont use the distro python either...
So, what values does it offer to whom? Except maybe scripting for sysadmins?
That's pretty sobering look at the state of affairs.
Literally, as one of the reinventions is actually called a "wheel".
And yes, both 'are' a Linux distribution too.
The scale and quantity of software being developed is just too much for traditional Linux distributions approaches. I personally rely on Debian to know that most of the packages I use are reasonably secure. But often I need to sacrifice flexibility and bleeding-edgedness.
And it couldn't be otherwise. There are 339,267 projects on PyPi.org. I bet a good percentage of these are riddled with security flaws. But people use them anyway. How many distro maintainers would you need to handle this workload? Is it worth it? Does anyone care?
It looks like people care much more about experimentation and speed of development than stability, security or coherence and cleanliness of the solution. The author seems to feel this is wrong from an engineering (or even moral?) perspective.
I am also uncomfortable when I have to deal with Python and I have my own favourite few tools... but perhaps if Python users really wanted a single solution then it would already exist.
If you are instead convinced that there is this need and no adequate solution, then congratulations, you just found a gap in the market. Go on and do better than everyone else before you. Relevant XKCD comic is already in the article...
Debian apt installs with using python3 and specifying pip as a module call is the only thing keeping me sane
Using JS or rust seems easier with the unified build system.
It is not the end of the world to be using one or even two minor versions older than the current Python version. New features must be implemented in your code and relying on bleeding edge features of your language (in production code) is outright bad design. If the core Python developers like to experiment and break their language, you are best to avoid jumping head first into that mess.
TL;DR: Keep your requirements conservative. You are not bewilden to the Python core developers, your target are your users.
The biggest insight in software dependency management is that applications and libraries are different. Both have dependencies but they sit at different places in their dependency graphs.
A library can be used together with other libraries and so cannot pin its own dependency versions (if all libraries did so there would conflicts everywhere) but instead can only specify constraints on its dependency versions (eg >=1.0.0).
An application sits at the front of the dependency graph. Nothing depends on it. It can therefore lord it over its dependencies, pinning everything in the graph to specific versions. This only works, though, if it doesn't have to share an environment with other applications and unrelated libraries.
Systems like poetry (akin to npm or cargo) allow a "lock file" to be generated with pinned versions for all dependencies, satisfying all the version constraints. Applications must commit this to revision control, libraries can if they want (should IMO). This is great as it allows CI and other devs to use consistent versions.
The missing piece is that for an application, the lock file should also be used for deployment of the app. If you can ensure that an application is installed in an isolated environment with the dependency versions from the lock file (i.e. that the app was tested against) then a lot of the pain disappears.
So, the suggestion:
* Add support to standard python distribution (wheels, pypi, etc) for application packages to specify (in addition to ordinary version constraints) a pinned set of "preferred" dependency versions.
* Have tools like poetry/pipenv set these to the lock file versions.
* Allow a notion of "application packages", which are required to have this information in.
This should work very nicely with tools like pipx that deal specifically with python applications. The relevance to linux distros is that linux distros also should only be packaging applications (and their dependencies). Developers using libraries should be managing them with tools like poetry/pipenv, not the system package manager.
If the system package manager could install python applications in their own isolated environments along with their pinned dependency versions, most of the pain goes away for distro maintainers. If an application isn't working with the dependency versions it has specifically asked for, it's a clearcut problem with the application as published, and needs to be fixed upstream.
I realise this would be a significant change for package managers, but I think the same model makes sense for other languages with similar tooling and at least some of the work should only need to be done once.
Go on then, tell me why I'm wrong :-D
> A library can be used together with other libraries and so cannot pin its own dependency versions (if all libraries did so there would conflicts everywhere) but instead can only specify constraints on its dependency versions (eg >=1.0.0).
This is a Python defect. IIRC, for example, Node libraries can and do pin their dependencies without conflict, because Node has no problem using multiple versions of the same library in a process.
Python can't do this. It could (there have been proofs of concept that make this work), but it doesn't.
In many cases, applications depend on other applications.
This can easily been seen in most Linux package management systems and also for many applications that act as more user-friendly front-ends or automation for command-line utilities.
Examples:
* Debugger front-ends
* Package management GUIs
* Multi-media processing pipelines
* Many others
If they're depending on them as python modules to import, they're libraries not applications and should be installed as such.
Your honor, I have no more questions for the alleged victim. <drops mic>