Why to use ‘python -m pip’
snarky.ca
snarky.ca
# pip should only run if there is a
# virtualenv currently activated
export PIP_REQUIRE_VIRTUALENV=trueThis seems less like a reason to use `python -m pip` and more a reason to endorse languages with saner versioning and less need to sandbox twenty different versions from one another.
Go doesn't have package management, it's just list of git repositories and revisions that should be checked out and compiled with your code.
In fact it is encouraged to have as little dependencies as possible.
- great concurrency (https://trio.readthedocs.io)
- bad package management (it just is. though Poetry's making a dent.)
- tooling? I'm not sure what you mean here.
- docs are ok not great not terrible
- types are rolling out rapidly across the ecosystem
- stdlib is too big but PyPI is right there so who cares
- ok to distribute? pyinstaller and shiv are alright.
He means linters, debuggers, refactoring libs, and other such things...
That being said I would like to inquire which languages _you_ could give as an example that work perfectly well with the system-provided compilers when you are using the latest features.
Honestly virtualenv's solve most of these problems. If using a virtualenv seem to low level or a pain, I can't recommend Poetry [0] enough. It uses virtualenv's in the background, and has from my experience a great dependency resolver.
All of those build tools are pretty complicated. I honestly prefer python and virtual environment. Simple requirements.txt file.
Want to ship all your app and all it's dependencies in a single file, sweet as, use the shade plugin, or maybe the capsule plugin. etc. Etc.
The Python equivalent is, well, take your pick: https://docs.python-guide.org/shipping/freezing/#freezing-yo...
> Simple requirements.txt file
I've never had a JVM dependency break because I upgraded Maven. I have had that exact issue with Python - we're using Apache Airflow, and anyone who installed the latest version of Pip hit an issue - Pip had changed an API that Python dependencies depended on, so Airflow could no longer be installed by that version of Pip, because it was written for the old API. Have fun googling "TypeError: 'module' object is not callable" to enjoy that rabbit hole.
Not to mention that it was impossible to generate a pipfile.lock for Airflow because its dependencies had rules demanding incompatible versions of dependencies (one wanted <= 7.0.0, the other wanted > 7.0.0) - in Maven I can easily exclude one of the two competing dependencies.
Oh, and then there was the time that a Python dependency used by Airflow required me to set an env var to handle a licensing issue. Can't even do that in Maven, which strikes me as a feature.
pip3 install apache-airflow
RuntimeError: By default one of Airflow's dependencies installs a GPL dependency (unidecode). To avoid this dependency set SLUGIFY_USES_TEXT_UNIDECODE=yes in your environment when you install or upgrade Airflow. To force installing the GPL version set AIRFLOW_GPL_UNIDECODE
Also, your simple requirements.txt didn't mention how you manage virtualenvs.Lastly, Gradle is being used quite a bit, but Maven is still pretty dominant. It really depends on if you want a purely declarative build tool like Maven, or one you can drop scripts into like Gradle. I prefer purely declarative for reproducible buils.
Pip explicitly has no API (except for its command line interface), even the Pip namespace implicates this (pip._internal). So if some dependency is using a Pip "API" they're wrong.
Now, I am developing an application that needs to interface with Pip, in this case it may be appropriate to use the internal Pip API if you pin your Pip version in requirements.txt. However a library shouldn't do this ever.
BTW, I don't rate Python dependency management very high: most other languages have better solutions. However it is manageable.
https://github.com/apache/spark/blob/121f9338cefbb1c800fabfe...
At least half of that is figuring out what the classpath should be. And the worst part, it is all exposed to user, too -- when I was setting up Spark, I had to muck with class paths a lot.
So JVM has 3 big problems: (1) No single way to set up dependency path, each project is different; (2) No way to set up classpath in JVM language, has to use some other one (like bash); (3) Somehow, user has to fix it for many semin-advanced programs. No other ecosystem I know of is this insane! Python, node.js, even C++ are all better!
Okay, so, speaking for myself, we generally ship fat jars or capsules these Docker days. I can't speak as to why Apache Spark chose to do it that way, I don't run their project. But if I'm guessing I suspect it's so you can provide your own Hadoop dependencies as needed, as there's a bunch of varying distributions that people use (Cloudera, Hortonworks etc.) and they're trying to play nice with the existing Hadoop ecosystem - probably also to make it possible to extend it with any other jars that you want to use (although spark-shell has a nice feature where you provide additional dependencies on the command line, or within the shell, using Maven artifact coordinates)
For an example of what's involved using an uberjar https://docs.cyclopsgroup.org/jmxterm > java -jar jmxterm-1.0.0-uber.jar
That's it.
BTW, the uberjar equivalent in Python is something like PEX/shiv/zipapps.
looks like a custom one.. and non-relocatable, too -- if I wanted to install "jmxterm" to my home dir, I'd need to manually edit the script.
(and yes, this is totally "I am developer" issue, not just "I am a user" one -- the dev had to write the script, they had to debug it, and dev would be handling bug reports about inability to invoke from command line when installed to other location)
Funny how you had to go searching to find that script, when the installation instructions are simply - 1) download uberjar, 2) run uberjar. Also, you had to go searching to find a script 8 years old, in a project with a commit 20 days ago, to attempt to prove your point.
Surely you must realise you're grasping at this stage.
/s
The articles arguments against this and for always using python -m pip don't convince me, personally. I never work on windows, so that's a mute point, and b/c I use direnv to automatically activate my env when cd'ing into the directory, I'm never not in the project environment. So basically, its impossible to not be using the correct pip version i.e. I have never once had this problem and I am constantly working on multiple python projects + code bases in various envs and versions.
And then python “just works”.
Does your text editor know that? Your colleague? How about your CI-system? What about the deployment story? When are things “unexpectedly” going to break because not everyone knew all these informally applied requirements?
Admit it: python is worst in class here. Almost no other popular language or platform forces it’s users into taking these absurd, error-prone steps, just to have basic shit “just working”.
Take a look at Node: npm may be a mess for other reasons, but boooy do you have clean per project dependencies. Look at Cargo for something even better. Clojure has Lein. The new dotnet CLI is slick too.
At this Python just plain sucks. It’s about time you guys start at least admitting it, or else it will never be fixed.
As far as the CI-system knowing, there is not much difference in installing a particular version of rust/cargo say in a docker container and having a particular version of python installed in a container, so I don't really know what you are getting at.
Could it be better? Sure, I do some programming in rust and IMHO the best build tool I've ever used is cargo. But compared to some of the other stuff you mentioned which I've used on projects as well (npm, lein specifically) I would say python is on the same level, and I prefer python over them but that's probably just my bad experience with package-lock/caching and long jvm load times speaking.
Whenever a discussion of this comes up, it seems there are always some "just use xyz" patterns, and usually 3 or 4 xyz or some very new thing someone just released and it always seems contradicts python official documentation. For every python packing question, there are 33 answers. I guess the irony is that there is more than one way to do it.
Use poetry, use prose, use pip, use venv, use vidtualenv, use conda, use --user, pipenv, pipfile, pipy,pipx,anaconda,miniconda.
It really is terrible. Npm gets a lot of flack for having vulnerable packages but do you see how many python fanatics push pip freeze? It allows them to ignore security vulnerabilities. I'm scared of upgrading any dependencies in my python projects because any update is virtually guaranteed to break. I can upgrade npm all day long and rarely encounter an issue. The ecosystem is so radically different.
Your advice: forget whatever brought you to Python and start from scratch.
I really can't figure out why conda is not ore popular.
The lack of a Conda lockfile makes it impractical to use outside of toy research proof of concepts.
And while lockfiles are convenient, they are far from making anything “impractical”. I take Conrad correctness over minor conveniences any day.
Lockfiles are not just convenient, they are all but required for anything serious. You can’t have your dependencies break or change under your feet from one build to the next. Your builds need to be reproducible.
I’ll take all that over whatever conda has to offer any day.
I see. So because C, C++ and most other build environments don't have lockfiles, there is no way to do reproducible builds. I better tell the Debian people they've been wasting an awful lot of time on their reproducible build project. /s
Seriously, "lockfiles are required for anything serious"? That's ridiculous. But if you insist, a quick google shows e.g. [0] and [1] provide that.
Most build environments do have lockfiles. And, just to clarify, that doesn’t have to be a specific dedicated file. It has to be something that can be versioned alongside the code, so each build gets the exact same set of dependencies and updates are explicit commits.
This basic principle is a requirement for anything serious (I.e something with customers). I’m sorry that this statement hit a nerve for you, but it’s true. In fact in some industries it’s a legal requirement.
And no, a dodgy third party plugin that hasn’t been updated in a year isn’t a good solution.
In fact, they also specify a specific gcc for extensions that need it, because relying on the system gcc is not reproducible. How do you do it in pip/vent/pipenv/poetry?
You depend on `cool-package==1.0`. That depends on `another-package` with a loose specifier, i.e `another-package=LATEST`.
Now when `another-package` is updated you suddenly have a different tree of dependencies, because the package resolution is run again and `another-package=LATEST` is installed. Imagine you're in the middle of rolling back (or rolling out) a deploy to your product. Suddently what you've been testing and working with has changed, and `another-package=LATEST` breaks the deploy due to some changes or bugs.
What's worse is that it's now harder to roll back, as a re-build and a re-deploy will still bring in the broken `another-package=LATEST`!
The solution is to lock the tree of dependencies, including all sub-dependencies. This has the advantage of speeding up installation as resolution doesn't need to happen. So, your package tree is locked to:
cool-package==1.0
another-package==0.1 # Non-broken!
That makes any updates to your packages safe, versioned and able to be reverted.The lack of this feature makes Conda a no-go for anything serious.
> specific gcc for extensions that need it, because relying on the system gcc is not reproducible
apt install gcc:4:4.9.2-2
Also "conda env export", which adds some more documentation; In fact, all defaults list explicit versions of packages. The only way you can list installed packages without it is "conda env export --from-history", which will reproduce the non-version-annotated list.
So, yes, you don't have a "requirements.txt" and a specific lockfile maintained on disk; instead, conda manages it, and if you want it in source control (and you likely do), you need to make sure your local pre-commit hook runs the export. But everything is indeed reproducible, without any additional package needed.
> apt install gcc:4:4.9.2-2
Ok, so your ci (and everyone) must be root / fakeroot; and you somehow have to manage that outside your requirements.txt and lockfile (e.g. in a pre-commit or post-checkout hook). And you can't have more than one at the same time on the same system (so your ci must have different containers/vms to build two executables with conflicting requirements, even if they are part of the same project). And you can't do it on Windows or Mac at all (whether root or brew user or whatever).
From every possible aspect, Conda is much, much more suited for reproducible and well defined builds. It does it differently than managers using lockfiles; But it definitely does it, and more completely.
p.s. conda also has "conda list --revision" which shows you the entire history of version updates. Very useful for troubleshooting, even if not strictly required.
[0] https://docs.conda.io/projects/conda/en/latest/user-guide/ta...
That's a very inefficient way to run your CI, with conda and pip alike.
Instead, you could build your environment once in a Docker image and use that as your build image.
Saves a lot of time on your builds, guarantees reproducibility, and will work even when package servers are unavailable.
We build images based on their commit hash, caching off the last commit hash. That works but has issues with merges and the first commit to a branch.
We also do it based on the branch name, but Docker has issues around specifying multiple cache from arguments in the CLI. That causes unnecessary invalidations on branches.
That all leads to more rebuilds than I would like.
Whereas I’ve spent a lot more time getting pip packages to install, update, etc before switching to conda.
Conda never broke anything for me - it did downgrade after upstream broke things. Pip, on the other hand, broke things more than once.
For installing global command-line scripts, I like pipx instead of pip. https://github.com/pipxproject/pipx
For development, I like Poetry instead of pip. https://poetry.eustace.io
Poetry is a replacement for pipenv + setuptools.
Pipsi is unmaintained and replaced by pipx.
Yeah but you just kick the can to which python happens to be the system python or activated python referenced by “python”. It’s exactly the same problem.
If I as a user have to be careful of which underlying Python installation is being referenced, then I want to be responsible for activating the conda environment I need and using “regular” commands like plain “pip” after that. The “python -m” idiom is not important for this use case, as it doesn’t actually solve the problem (ensuring the right referenced Python).
Exactly right. And on any sane Linux system the pip command should get you a version compatible with the system Python, which means this is a pointless extra step. Unless these instructions are supposed to be for Windows, or MacOS, in which case it should probably say so. (And doesn't everyone on Windows use Conda anyway?)
On top of activating the venv when (and only when) you're inside your project directory tree, you can use dotenv for loading all relevant env vars (AWS creds, API keys, etc.,) and add your project's root dir to PYTHONPATH with path_add.
Come to think of it, only there was a PIP_DEFAULT_VIRTUALENV_PROJECTDIR variable which would causes pip to look upwards until it found a setup.py file, and then use a .venv directory in that directory by default...
Pip installing all from source is tricky to get working correctly as soon as you require a newer compiler or dependencies. Providing wheels works well but is a pain to set up correctly with good platform coverage.
That said, it doesn’t need to be one or the other. Conda integrates well with pip so you can install a base layer with conda then pip install your more esoteric dependencies.
Conda even adds advanced features, like managing multiple compiler-specific build variants in the same environment, that go way beyond anything other Python packaging tools do for this.
On top of this, you can install pip into conda environments and still use pip to manage packages if you want, or mix and match conda & pip with a pip section in conda’s environment.yml. Conda automatically manages pip installation into the activated conda environment.
Personally, I've come to prefer homebrew python + pip for development on macOS (even though said python is not optimized), and clearlinux + pip for Linux and production use.
Clearlinux actually gives you a pretty simple way to install a highly optimized python and most all common libs (numpy, scipy, ML stuff, even with GPU and Intel MKL support) - then you just run a pip install -r requirements.txt to get the small, pure python stuff that you want on top that foundation.
alias pip python -m pip
then stop worrying about itPackaging is one of my least favorite things about the language I love the most.
Anyway, I think the fact that people has more patience with learning a language than tools is that learning a language is fun, while tooling is just boring.
However, Python is also one of the oldest programming languages still in use and one of the first to have the concept of installable dependencies (easy_install).
Now it is kinda difficult to fix without breaking something, however I think we eventually will have a solution, like __pypackage__ https://www.python.org/dev/peps/pep-0582/.
Unfortunately there are people who complain that pipenv makes the ecosystem as bad as npm, which frankly is an improvement.
Required By: apitrace apm asciidoc bluefish bzr cloudprint-cups dblatex epydoc fio gconf gemrb getmail gmock gnome-doc-utils gtk-recordmydesktop inkscape ipcheck ipython2 java-openjfx java11-openjfx java8-openjfx jcl jmc kig kodi-bin kodi-gbm kodi-wayland lhapdf lib32-apitrace libgda libkate libvirt-python2 libvolk lilypond mailman mcomix mediaproxy mercurial mftrace mod_wsgi2 moosefs mopidy mysql-python mysql-workbench nss-pam-ldapd ntop pluma pychecker pydb pylibacl pynac pypanel pyrex pyrit pysol-sound-server rdiff-backup repo rtaudio scribus sgmltools-lite shedskin singularity spambayes tellico texmacs tuxpaint txt2tags vim-ultisnips wesnoth wicd wifite xpra
Optional For: android-tools arduino armagetronad boost boost1.69 brial bugzilla cairo-dock-plug-ins cantor cryptominisat5 csound dia ecasound ecryptfs-utils efl efl-docs faust freedroidrpg freeradius gammu geda-gaf gif2png git gogglesmm graphviz gvim hivex inn john kcachegrind-common kodi-eventclients kross-interpreters libdnet libevent libglade libieee1284 libnewt libproxy libuhd magma marsyas mate-menus mathomatic mediawiki munin-node ncmpc net-snmp non-timeline notmuch openconnect opensips plan9port postgresql pycharm-community-edition python-pywal recoll roundcubemail rrdtool singular skktools spectmorph spring subversion texlive-core texlive-music vim vim-latexsuite xf86-video-qxl
Regardless, let's not turn this into yet another argument about why anyone should/shouldn't upgrade to Python 2. I was just trying to confirm that the parent is not doing anything that requires Python 2.Most (all? actually I don't know any from your list that do) packages that you listed are no longer maintained
Except for PYTHONPATH variable, 2 and 3 worlds are completely separate.