As a "professional" python user (who got there in a roundabout way), I'd say the biggest problem with package management is all the conflicting advice and different ways to accomplish the same thing (there are other problems, but they are surmountable). Once you get a workflow down, it ends up being pretty easy to set up and manage an environment for a given project. Not that it's perfect, but I think a lot of the stereotypes about how bad it is come from how bad it appears as you try to converge on a setup that works for you.
As a tongue-in-cheek reaction, Python espoused "TOOWTDI", or There's Only One Way To Do It :)
reference: https://wiki.python.org/moin/TOOWTDI
The problem is, when it comes to the Python package management ecosystem, there are SO MANY ways to do it. And they aren't equal and require a deep understanding of each tool to select the correct choice.
It is a problem for Python, although I often regrettably see it dismissed as not a true issue because $latest_package_management_system fixes it.
Python has allowed me to be profoundly productive in many ways, but this is a huge sore spot for the Python ecosystem.
Python scripting today requires shipping your development environment. Python is wonderful until you want to run that code on another machine. At that point, the target system has to venv their way into reproducing your environment...often including the specific Python interpreter you picked, and to download (and possibly compile) all modules and their dependencies.
It can work beautifully and many of the large companies I've worked for have put oodles of engineering effort into making it "easy", as long as you follow their happy path and don't deviate.
But, to your point, on my own machine with venv + pip it "just works". The pain comes when trying to venv + pip on another machine.
How confident are you that you could venv + pip your moderately complex Python application on 200k machines without issue? What if it's a mix of Windows, Linux, and macOS? (this is a real scenario I've experienced). From my own experience I can share that it is painful. Your mileage may vary.
And after a while your system is riddled with Python venvs. Leading to some mildly annoying problems with popular IDE's: https://www.jetbrains.com/help/pycharm/package-installation-...
Does docker fix all the above issues?
Step 1: Docker must be installed on all targets. This is not a given and is a new piece of overhead.
Step 2: The Dockerfile hopefully doesn't source from just "ubuntu:latest" and bring in the whole kitchen sink for this SINGLE APPLICATION.
Effectively if you are using Docker and deploying your script to Windows or macOS, you are saying "Hey, this script requires you to install linux in a VM to run this. Docker makes it easy. Go download 500mb of Docker and a couple hundred more mb of images to run this 800k script.
sigh
For example, it would install ansible and the ansible dependencies to a separate environment somewhere, but still put commands for `ansible`, `ansible-playbook`, etc. into /usr/local/bin for example. When you run the ansible command, it loads ansible from the separate environment.
I haven't actually tried this myself since our environments are pretty uniform, but it could be worth looking into. It would be nice if pip itself provided this as an option (or if it became the default), but that could also become extremely complicated in terms of upgrades. I already hate having to download and build scipy, but it would be incredibly irritating to have to do it once for every tool I have that uses it. I'm sure some kind of cache could be made to work, but it's not trivial.
There's a workaround but it's not ergonomic and I always forget what it is.
It's not just a mild annoyance, it's a sad statement about python as a community of developers that this is not only accepted but recommended.
I write python because some of what I do occurs in the domains where python makes the most sense, but the idea that the python community accepts venv rather than considering the fact that venv even exists at all to be a source of profound embarassment is mystifying to me.
That bothers me too, a bit. But sidethread to you people are comparing python unfavorably to ruby, which as far as I know uses exactly the same strategy (except that it's called "rvm" instead of "virtualenv"). Node.js appears to do the same thing.
To my eyes, the solution to this would be static linking, but I don't know how much sense that makes in the python / ruby / js context. Other than that, what's the alternative? There's certainly a lot of convergence on this solution.
Python package management is VERY easy.
Single file. Copy and run.
Virtualenvs require every single target to reproduce your dev environment:
(1) have internet access and able to reach pypi (or artifactory, or whatever you use).
(2) the ability to install the required version of Python if it isn't already installed. That's another 30mb download.
Ever want to run a complex Python script on a bastion host or a host behind a bastion? Well now #1 and #2 above won't work (or haven't in my direct experience working for some of the big cloud companies). So you have to use something like PyInstaller and hope it works. It might. It might not. A statically compiled C++ binary or Go binary probably will.
This thread is full of pain points in python package management. The first step is admitting there is a problem. I don't think you've experienced the pain caused by Python package managers that others outline in this thread.
Obviously Python is (usually) an interpreted language, so you're going to need Python on user system. If that's a problem Python might not be for you, or you're going to need to do some extra work.
It can be done. For example the popular visual novel software Ren'Py is written in Python.
Statically links all the dependencies and the Python interpreter. Other machine can just run the executable, no need to set up a venv+pip.
Of course the binary is huge, but IME it works. Unfortunately it can't cross-build, so you need a VM or multiple machines set up with the dev environment to run pyinstaller for each OS.
However, it was pushy. It would put its own header/linker paths in front of the system paths (this is how it made "environments," its wannabe containers), which tended to create inadvertent cross dependencies if you didn't understand or didn't remember that the semantics of a conda environment extended beyond python. These dependencies could get baked into binaries and break far down the road, or they could get sucked in as a transitive dependency and trip over the shoelaces of a different build of the same software installed outside conda. Unfortunate. However, around 2019, the problems started growing beyond mere foot-guns. Conda uses a full SAT solver to provide a highly featured versioning system, and this worked great until the big conda channels grew to the point that it started getting really slow. Installing packages went from taking seconds to minutes to hours to forever. They tried caching, they tried fragmenting channels, but it was all very not-seamless.
Eventually, people started migrating back to pip. It turns out that over the last decade distro repositories had gotten their shit together and now Docker existed to sweep up the last few use cases, so nobody needed conda's "poor-man's docker plus curated 3rd party repos" anymore. Now pip is the tool that Just Works, and it Just Works without any of conda's baggage. Virtualenv environments don't hook your system quite as aggressively, pip never stalls when resolving its version plans, and Docker can be used to reproducibly experiment and find the happy path.
Be glad that you missed out on pre-conda pip and the conda arc.
Compare this with Node. It always installs locally by default and always installs in node_modules regardless of which package manager you use. This is integrated into Node so you never need to modify the path like in Python. You don't have to guess whether your packages are installed in .venv, env, environ, or whatever someone else decides.
`pip` actually has a lot of issues with regards to deciding which version to use. That's why people moved to pipenv... then pipenv stagnated and people moved to poetry.
Using `requirements.txt` is dead simple but it ignores issues such as version locking. If you just add the packages you need to `requirements.txt`, then every time you install you could get a different set of packages. If you do `pip freeze > requirements.txt` then you don't know what comes from what.
I hate the most popular ORM (sqlalchemy) and alembic has a lot of footguns: for instance, if you autogenerate a migration where you change a table name, it will try to drop the old table and create a new one.
In web dev a lot of the OSS community has moved on to more appropriate or exciting languages, so the tools I'd normally reach for frequently are OSS projects that haven't been maintained in seven years.
As for pip: it needs a major version bump that creates a lockfile, or it needs to be intentionally sunset. It's easy for a beginner to start using (part of why it's heavily used), but extremely dangerous because of the lack of guarantees. I wish it didn't have as big of a footprint as it does.
A classic example is version pinning. Rust has Cargo.lock, Ruby has Gemfile.lock. Python? It depends on which one of the multitude of options you pick. But at least a couple of the most popular ones basically don't do this (pip) or do this in a hacky, ugly way that makes you want to tear your hair out (Conda). As a result, at least in the projects I work on, people tend to skip this.
I hope it should be obvious what bad things can happen if you don't pin your dependencies, but for the uninitiated: I'm talking about things like projects breaking inexplicably after 6 months because you rebuilt some Docker container and something or other got upgraded and is now incompatible.
Beyond the technical issues, there's a human engineering problem of getting everyone on the same page about the right processes, and Python makes that vastly harder by not having a One True Solution that everyone just uses.
python3 -m venv .venv/
if [[ -f versions.txt ]] && [[ versions.txt -nt requirements.txt ]]; then
pip install --requirement versions.txt
else
pip install --requirement requirements.txt
pip freeze > versions.txt
fi
(I place the files in etc/pip/ in my projects (and check them into git), but I've omitted the paths for clarity. One could embellish this by including python version number in the filename, as package requirements can change between python versions.)I also have a bin/venv-python wrapper which sets PYTHONPATH, PYTHONDONTWRITEBYTCODE before chain calling .venv/bin/python3 with the arguments, and this is how pip above is called. (again, omitted above for clarity.)
This won't cover everyone's usage scenario, but it works for me. YMMV.
https://iam.georgecox.com/2021/09/25/python-3-venv/ explains the details.
Also, `pip freeze` doesn't include platform-specific dependencies for other platforms. So if you run the following on linux and again on macos, you'll get different results because `ipython` depends on `appnope` only when running on macos.
python -m venv env && env/bin/pip install ipython==8.4.0 --quiet && env/bin/pip freezeThe beef I have with so many explanations of pip is that they tell you to source ".venv/bin/activate" and then "just run pip install whatever" without a) separating the root requirements from the effective/complete requirements, and b) they don't suggest using a wrapper for the ".venv/bin/python3" binary so that execution is the same in all environments.
I don't think this practice is widespread, which is exactly why I made a point about human engineering in my original post. But I do appreciate that solutions like this exist, and I should look into driving more of this sort of thing in my projects.
Does it matter? I get a different way to define dependencies. But the lock file itself is an implementation detail. You use the package manager the project uses and the lock file can be opaque. It's the same for package-lock.json / yarn.lock in js land.
for pip file I usually make a requirements-to-freeze.txt, make a fresh virtualenv, install requirements-to-freeze.txt and then pip freeze > requirements.txt poetry does this automatically
I use poetry every day and so far it's been pretty great. The main complications I have had are with accessing private pypi repos in Azure DevOps pipelines which use short lived tokens but it just looks a little clunky, still works.
> Also its odd that pyproject.toml (not poetry.toml which is the poetry settings for repo) is the dependency definition file and poetry.lock is the lock file
This is because poetry is using Python's PEP 518[1] specification rather than define their own build requirements format. It also isn't limited to just building, you can also include the configuration for other python tools like `pytest`[2].
[1] https://peps.python.org/pep-0518/
[2] https://docs.pytest.org/en/stable/reference/customize.html#p...
I sort of get this but its still a bit odd. especially since I have a small poetry.toml as well.
> I use poetry every day and so far it's been pretty great.
I like poetry, it works a lot better for me then pipenv did(to be fair that was a few years ago). But I do seem to have to delete lock file to update it and sometimes poetry add/install breaks oddly(or at least with awkward error messages). Also sometimes it interacts with venvs in odd ways. python really needs to ship with it
In my experience, this is basically no one; in fact I've seen many (most?) projects not even pin the versions on their top-level dependencies. (Ask yourself how many times you've seen just "numpy" as a dependency without any version bound. Far too often in my experience.)
And this is exactly why stuff breaks: even when the root dependencies are pinned (and are they?), a transitive dependency could get upgraded and break something (or fail to build entirely).
So you end up with all these different Python package managers that are all good at individual different things but no single package manager that is good at all of the things.
I swear that is a major problem with most libraries or tools (and products overall actually) that people make. It’s like people fix their pet peeve but forget that the whole picture matters way more.
BTW, I know setup.py method. How much is that in percents?
I hope you'll forgive me for adding one additional piece of advice: for many Python packages, the only packaging metadata you need is `pyproject.toml`. You don't even need `setup.py` anymore, so long as you're using a build backend that supports editable installs with `pyproject.toml`.
Here's an example of a Python package that does everything in `pyproject.toml`[1]. You should be able to copy that into any of your projects, edit it to match your metadata, and everything will work exactly as if you have a `setup.cfg` or `setup.py`.
For managing python itself and binary libraries I started using Nix package manager.
It allows to describe all dependencies via code, but with time that code became a boilerplate, so I created this: https://github.com/takeda/nix-cde
It works very well for me so far.
You do need to have Nix[1] installed, but hopefully that should be the only thing needed and everything you can just list in project.nix file (in the example in README.md you should have access to `dive` command even if you never installed it on your system). Which will make it available only for that project.
Here's also an example of a very simple project: [2]
[1] https://nixos.org/download.html#nix-install-macos
[2] https://github.com/takeda/nix-cde/tree/master/tools/aws_assu...
1. Install pyenv to a central location of my choice, e. g. ~/.pyenv (this helps manage all those different Python versions piling up from all the isolated projects I have.)
2. Install pipenv as a stand-alone tool. Doesn’t matter which Python I’m using, I just want to have it in my PATH.
3. Now I’m ready to create fully isolated projects using pipenv. Not only does pipenv create a venv right inside my project directory for me, it also helps me manage and lock dependencies. It also talks to pyenv so it can fetch the per-project Python version for me; this is especially useful when collaborating with others on the same project; they can reliably check out the project without having to manage specific Python versions themselves.
YMMV but I’ve had a stable experience using this setup so far.
There are solutions to this (pipenv being a popular one, though sadly often not explained to Python newcomers) and there are alternatives for that as well. You probably don't actually want to install a package like a package manager, most of the time you want to install an application and the specific requirements for that application, which is exactly what pipenv is for.
Your global package manager and your global Python package manager don't agree on what version is "the latest stable version" of a package so you just can't expect global package installs to work like that. It's like trying to make pacman, apt, and dnf work on the same file tree: you can probably get it done, but the end result will grow unstable or unusable within days.
It's true that people have made a lot of different ways to install Python packages. Each of these people has an axe to grind and they have supporters who are trying to boost their preferred solution and keep people away from the others. That part really is a mess.
The standard virtualenv+pip has always been rock solid for me. Maybe this is because I do all my development inside virtualenvs and I only ever use pip inside a venv. And I run Linux, though I do ship my Python code to a fleet of servers.
Sometimes I think maybe this Python-ecosystem-is-horrible stuff comes from people who are trying to steer people away from Python towards their preferred language.