Pipenv: Promises a Lot, Delivers Little (2020)
chriswarrick.com
chriswarrick.com
python3 -m venv .venv --prompt="foobar"
. .venv/bin/activate
pip install -r requirements.txt
?Whenever I see all these other tools I just get the feeling like there's some big elephant in the room that everyone is battling against, but I've never come across it as a python dev, and the moment I try to user/understand these tools I feel like they're against the "keep it simple, stupid" vibe that python gives me.
There's auto-venv for the truly lazy (like me) but other than that I don't rely on any of these Env/requirement wrappers or tools in any of my projects (virtualenv-wrapper for work but that was in their setup guide).
Is it a legacy thing?
Then even with the new dependency resolving improvements it's way worse than poetry's in my experience.
Also the requirements.txt file tends to get cluttered with non top level dependencies that make upgrading your dependencies an herculean task. Instead on poetry you define your top level dependencies in your pyproject.toml and all the actual pinned dependencies will be compiled into a poetry.lock file.
And finally the user interface is much more modern. `poetry shell` is way simpler than `source venv/bin/activate`. You don't even need to run `poetry shell` because there's `poetry run`. And to install a new project is just `poetry install` instead of `python3 -m venv venv; source venv/bin/activate; pip install -r requirements.txt`.
Not to mention that it's way easier for me to understand where the depenency conflicts from console output compared to even the newest pip versions.
This has now been solved [1] in pip, so no third-party tool is needed. Put your pinned dependencies in constraints.txt and your top-level dependencies plus the line "-c constraints.txt" in requirements.txt.
[1] https://pip.pypa.io/en/stable/user_guide/#constraints-files
As an example, let's create a venv and install some older versions of Django and its dependencies (current versions are 0.4.2, 3.5.2 and 4.0.6)
$ python3 -m venv env1
$ ./env1/bin/pip install sqlparse==0.4.0 asgiref==3.5.0 django=4.0.0
...and also Flask just to complicate the constraints file for the example: $ ./env1/bin/pip install flask
Lock all dependency versions in constraints.txt: $ ./env1/bin/pip freeze >constraints.txt
Create a requirements file that specifies just "django" and references the constraints file: $ echo '-c constraints.txt' >requirements.txt
$ echo 'django' >>requirements.txt
$ cat requirements.txt
-c constraints.txt
django
$ cat constraints.txt
asgiref==3.5.0
click==8.1.3
Django==4.0
Flask==2.1.3
itsdangerous==2.1.2
Jinja2==3.1.2
MarkupSafe==2.1.1
sqlparse==0.4.0
Werkzeug==2.1.2
Now we can create another venv with the exact same versions of Django and all its dependencies (but not Flask or its dependencies) using just pip and the requirements file: $ python3 -m venv env2
$ ./env2/bin/pip install -r requirements.txt
$ ./env2/bin/pip freeze
asgiref==3.5.0
Django==4.0
sqlparse==0.4.0 poetry init
poetry add x y z
Then later poetry install
And maybe poetry remove yI hear you on the pinning, but "UX" concerns like this drive me batty; I just use shell functions for this:
venv() {
. ${VIRTUALENV_FOLDER}/${1}/bin/activate
}
gitsu() {
branch=$(git status | head -n 1 | cut -d ' ' -f 3)
git push --set-upstream origin ${branch}
}
git_add_conflicts() {
git add $(git status | grep 'both modified:' | cut -d ':' -f2)
}
It's so much easier to add a little function in your shell than to write a new tool.If every tool were like that that would be one kind of world, but we live in one that's more interesting. There's some Mark Twain quite like "the world owes you nothing; it was here first", or maybe the Buddha saying "it's harder to soften the earth than it is to wear sandals". I like to think that sandal-making is a core engineering skill, whether you're a carpenter or a computer engineer.
Separately, I'm curious to hear examples of where dependency-clutter actually caused a problem.
I used to work at a place that was just using requirements.txt files that only included our direct dependencies. There was a project that needed updating after not being touched for a couple of years. The requirements.txt didn't change, but when we built the project again, some of the transitive dependencies used a newer version, and a bug was introduced from one of those updates. A bunch of time was wasted tracking down the issue, pinning the old version of the transitive dependency, and figuring out the damage caused by the bug.
As a result, the requirements.txt was changed to also include transitive dependencies. We had vulnerability scanning on our code, and it found a severe issue with one of the transitive dependencies, but there wasn't a version of that library with the issue fixed yet. Time was spent looking into this to see how we could be impacted. As it turns out, it was a transitive dependency for a library that we no longer used and removed from the project months ago. When you create your requirements.txt by running pip freeze > requirements.txt, you don't have an easy way of knowing which library requires which transitive dependency.
There's ways you can fix this using multiple requirements.txt files, but at that point it's a lot easier to use poetry, especially if you want to keep your development dependencies separate.
Maybe that could work with a `pip --user` install, but at least the venv guarantees that all Python packages and data will be in one self-contained folder.
pip install -r requirements.txt -c constraints.txt
[1] https://pip.pypa.io/en/stable/user_guide/#constraints-filesI think it's a lot of javascript developers who end up having this problem, and all I can say is that the python package ecosystem is not the same. Please don't lock your dependencies to specific minor/patch versions, and please don't use so many dependencies that it becomes tedious to deal with changes. Especially when you're writing a library. If you're locking to anything other than a major release or a minimum minor/patch revision when producing a library things have gone very wrong.
My time and the time of my peers is too valuable to be debugging minor incompatibilities between a library 5 layers down the transitive dependency stack because versions aren’t pinned. It’s too important that if I go back to a 3 year old checkout of the code that it still runs and isn’t a wild goose chase of running down bugs and library incompatibilities.
You can always upgrade overzealously when dependencies are pinned, falling back to the workflow you’ve described, but the inverse isn’t true. If your tools don’t allow you to pin, you’re signing yourself up for churn at unknown and unpredictable times.
That happens very rarely with our dozens of transitive dependencies. What are these packages that cause trouble often?
Still, it doesn’t have to happen often to be a serious hindrance to either developer productivity or your business. Imagine trying to roll out a fix for a production outage only to find a random transitive dependency has started breaking your build. Now your outage is extended to the duration of debugging and fixing an unrelated and avoidable problem.
Because you don't do ml. My requirements.txt is useless in 6 months because the transitive dependencies all have incompatible versions of common libraries by now
Maybe you’re using recent less-stable projects? I found that the main ML packages rarely break things without months of deprecation warnings.
You’re not wrong, but developers still need precise control about which version is pinned so they have a chance to figure out whether a bug is in their code or a regression in a dependency.
> If a dependency is constantly breaking on minor/patch releases you shouldn't be using it.
A more common example would be a dependency of decent quality, but there will still be bugs because there’s no such thing as bug-free software. And when the inevitable issue comes up, it’s super useful to have a switch that lets you pin down the regression and helps you figure out whether it’s on you or on the upstream project to fix the issue.
> Please don't lock your dependencies to specific minor/patch versions
Even if you allow a version range of "*" in your dependency file, it may still be a good practice for your app to have the resolved version number pinned in a dependency lock file, and even commit that lock file to version control.
> Especially when you're writing a library. If you're locking to anything other than a major release or a minimum minor/patch revision when producing a library things have gone very wrong.
You’re referring to the dependency file, not the dependency lock file, right? I couldn’t agree more. At the same time, it may still make sense for some libraries to have a lock file during the development process, and even commit that lock file to version control. You just can’t include that lock file in your release.
I don't want my tools "helpfully" upgrading me to a different version, I don't care if it's "minor". It's different and that's a potential source of heisenbugs. Any change to dependencies should be an explicit action and it should be a commit in SCM.
By all means show an ugly warning message to nag people to upgrade but computers live to serve us, not the other way around. The moment you start devolving these decisions to a computer you've lost control of your core competency, which is ultimately what code is running on the machine.
They also create lock files with hashes by default.
pip still doesn't use lock files. It's subpar compared to other ecosystems.
If your ecosystem is healthy than pinning exact versions with lock files shouldn't be done. Making it so that every dev uses the latest patch or minor when they run your program.
Using lock files for libraries should absolutely never happen, you should at most fix your dependencies to the latest major version.
I think this issue is unrelated to semver.
I've had so many issues over the years with Python packages, even very polished libraries like Flask and Celery.
For example, you could install something like Flask 2.0 but end up with a major difference in versions of its sub-dependencies depending on when you installed it.
That's because Flask has its own dependencies defined like this:
"Werkzeug >= 2.2.0a1",
"Jinja2 >= 3.0",
"itsdangerous >= 2.0",
"click >= 8.0",
The above means installing Flask 2.0 could one day in the future install Jinja 4 or Click 10 unless you lock your entire dependency tree.I've also had all sorts of things break because Flask installed Jinja 3.1 which wasn't a problem 6 months ago when Jinja 3.0 was the latest release. I've also had cases where installing a specific version of Celery in the past worked but failed in the future because it didn't lock one of its sub-dependencies down well enough (vine) which caused a breaking change.
This stuff happens all the time and it's a nuisance. I never experienced issues like this with Ruby and other languages that have the idea of a lock file built into their package manager. IMO it's a desperately needed feature that should be built into pip.
The Rust ecosystem is good evidence that locking doesn’t kill semver. Semver is still widely used and has all of its meaning.
You mean "if all your dependencies are using semver properly". Yes, you're right, and you've found the problem.
I think semver makes sense to humans. I can derive a lot of meaning from it when I see n.n.n. But when it comes to the software supply chain, it's just too rickety. Frankly, when you lose customers after a new deploy broke one of your dependencies, "but the dependency author didn't respect semver" isn't an excuse.
I say this as someone who strongly pushed `dependency>=1.3.2,<1.4` until that happened to me. My argument was "security updates", and now I just don't care. The software supply chain is too chaotic, and you have to be defensive against it.
I agree that lock files are useful and it’s a pity that pip does not offer them.
With sensibly written setup.cfg files, we just “pip install” all our packages with no issues. Pip’s dependency resolver has come a long way since 2019.
Some projects used to have two files for this, effectively managing a lockfile manually with pip freeze, but then it's nice to have a wrapper around this pattern and that's where the first gen like pip-tools came from, and then stuff like poetry/pipenv is aiming to streamline that (and avoid manual use of companions like pyenv for specific Python versons) even more.
Some tools, like poetry do add value to your development process (namely, better management of dependencies)
pipenv really does sound like something for people that are too lazy to activate their envs before running stuff
Any folks having a good experience with PDM https://github.com/pdm-project/pdm ?
My situation is I think unusual, in that I need to use a private pypi repo which requires mTLS for both fetching and publishing. Had it not been for that I suspect I'd still be using Poetry, but given the experiences I've had with PDM I wouldn't switch back even if the situation with my repo changed.
1. Poetry's default assumption on packages respecting semver simply does not hold up in reality. There are very few packages actually sticking to semver. Thus the `^x.y.z` default version range is quite often too loose. I've found that using `~x.y.z` for most packages is far more stable.
2. Imho, `poetry update` is a footgun. Without specifier, it will attempt to update the entire dependency tree. Not only is this slow, together with 1) it's all too likely one ends up with dependencies that actually are incompatible at runtime. I'd much rather have a `poetry update --all` flag instead for the rare instance I do want to update everything. The default behaviour should be to require a list of packages to update.
3. There are some common packages that cause very long resolution times if they are not restricted. Case in point: boto3. Even if one doesn't use boto3 oneself, it's very likely a transient dependency. Many packages simply specify `'*'` as their version dependency (they shouldn't, but it's the unfortunate reality many do). This will cause poetry to consider every possible boto3 version. With hundreds of versions - boto3 has a release every other day - this gets unwieldy fast. So I often end up specifying boto3 myself with some sensible range in my toml file, even when it's not a strict dependency of my own project.
4. The datascience ecosystem needs particular attention. Best to simply pin those, as every pandas update is guaranteed to break something. ABI changes to numpy are a particular nightmare. This is again due to too many packages simply specifying `'*'` for their numpy dependency. Which is further complicated by the fact that most don't distinguish between build-time dependencies and run-time dependencies. The numpy ABI is only forward compatible. Hence one should build with the oldest supported numpy[0].
virtualenvs are a much more supported.
I’ve been using Anaconda and been finding that it install incompatible versions of jupyter and ipython dependencies and that a lot of tools I need only work from pip, so I’ve been wondering if maybe I should just switch to the “pythonic way of doing it” and have absolutely no idea what that is.
For every project you're developing, there will be virtualenv which has all the dependencies that project needs (which may be different than what's installed in the system, and different than what other projects may need).
"python -m venv init project/venv" will create it. "rm -rf project/venv" will delete it.
Usually the virtualenv goes somewhere well known, like "project/venv". Sourcing the activate script ( "source project/venv/bin/activate" ) changes your shell environment to use the virtualenv instead of the system python environment. Once activated "pip install package" installs to the activated virtualenv. "deactivate" will turn off the virtualenv.
This is semantic sugar for what's really happening behind the curtains. There's a copy of python in the virtualenv: "project/venv/bin/python" which runs in the virtual env regardless of whether the virtualenv is activated or not. "activate" just adds "project/venv/bin" to the start of the PATH. "deactivate" removes it.
Regardless you can always see which python you're using by typing "which python". System python (/usr/bin/python) will use system packages. The venv python (project/venv/bin/python) will use the virtualenv python packages.
This allows you to have different virtualenvs to try out things like new versions of python, say. And each virtualenv is isolated from every other virtualenv.
And poetry is just a nice wrapper around this process that also figures out total project dependencies and creates a "lock" file to freeze all deps to a particular version for consistent releases. Pipenv is basically the same thing.
"pip install --user poetry" will install it in your home directory.
It's not recommended to use conda and pip together, basically the recommendation is to use one or the other, though I've heard miniconda is better in this regard. YMMV.
I'd also be curious what others recommend!
https://cjolowicz.github.io/posts/hypermodern-python-01-setu...
I use pipenv on one of my projects, including (in combination with direnv) for setting up my local dev environment on macOS and as part of CI/CD for deploying into a Docker container based on one of the Python images. The project doesn't have a huge number of dependencies, but pipenv has worked well. The only time I fought with it was over its PyUp safety checks, but "pipenv check" is opt-in anyway.
For most of my other projects, I install from either a setup.py or requirements file using pip, often driven by a Makefile. Again, deploying into Linux, but using macOS for local development, again with direnv.
In 20 years of professional Python development on mostly small-to-medium size projects, I've never had any issues I can recall getting packages installed.
Conversely, I semi-regularly fight with cocoapods, npm and gems.
0: https://packaging.python.org/en/latest/key_projects/
1: https://packaging.python.org/en/latest/guides/tool-recommend...
While both pipenv and poetry were nice and had a lot of comfort build in they broke in nasty ways and it was hard to debug so I traded the comfort for less complexity.
Our developer environment is well polished and can be started up with installing a couple of packages then running a single command... very nice, except for the Python part , which is extremely tiny compared to the other stacks.
We used pipenv, but that broke very often, so we moved to Poetry... it looked like it worked for a while, but as we got more people, specially ones using Linux and Mac M1 had to spend lots of time fixing issues with Poetry/Python... something to do with cpython dependencies downloaded for the wrong architecture.
We've had zero, absolutely zero issues with all the other stacks. We're trying to find solutions, but IMHO we'll just have to bite the bullet and re-write the small Python component we have into any of the other many languages we use without any issues.
Or fix what you perceive is wrong with it.
Don’t demand the tool do what you want or insist on a certain release cadence unless the developers are on your payroll.
I admittedly started skimming around halfway through but still want a refund on the time I wasted trying to determine if there was more to TFA than a whinefest.