You don't really need a virtualenv
frostming.com
frostming.com
Never had such a requirement.
Environments, like the interpreter itself, seem a singleton concept.
I have used a Makefile that sources different Bash environment variables kept in files in etc/ within the venv to switch between, say, a Flask and a Gunicorn startup.
Way way harder on yourself.
I had to set up Python projects for some machine learning classes in college and it was a complete mess.
C has shared libraries and static compilation (bundling of dependencies) Java has .jar files which probably contain all your dependencies, go and Rust statically compile to something that has minimal dependencies and python has... no clue. Some packages you have to install via your package manager (in linux), because they have to be compiled against your local libraries, some other dependencies are somewhat included... and if you then try to also develop on the same machine you can choose between your system python libraries or installing them via pip (which makes your system weirder to deploy).
Having said that, I'm no expert in python but I have done some work in it and the deployment question is something that always seemed weird and unexplained to me.
- There's one tool to use, and it comes pre-installed
- You don't have to deal with virtualenvs or packages from different projects conflicting.
- Native code dependencies "just work".
That's only because you're basically assured there will be conflicts inside any single project. Every time I install anything via npm I get a boatload of warnings about insecure dependencies, which are effectively impossible to fix without breaking the whole mountain of hacks.
There are good deployment stories out there. JS is not one of them.
That's not an issue with the package management though, that's an issue with the quality of the packages themselves. The JS ecosystem does have issues there. But the tooling itself is good.
Though, as the sibling comments say, it's debatable whether it should.
There are tradeoffs for sure, like the ridiculous sizes of those directories on development machines. But a giant SSD is cheaper than my time resolving dependency issues on user machines.
No language should replicate anything JavaScript does when it comes to package management
On the other hand, deployment of Python applications is still unresolved, and there is no single standard method to ensure both compatibility with and isolation from the underlying system in a universal way. There are different methods, from virtual environments (both standard and nonstandard ones like conda), to version managers like pyenv and asdf, to "jailers" like pipx and pyinstaller, to containers; but none of them has yet clearly won.
The candidates for distribution are things like pyinstaller or pex (still not fully mature) but I don't think they are as easy to use as Java jars or Go binaries.
> All of the above are suitable for deployment on your cloud servers but not as distribution mechanisms to end users.
virtualenv's are also not a distribution mechanism because they aren't portable (unless that has changed). I can't "distribute" a virtualenv to a host. I can only use virtualenv's to change how pip (or poetry) installs artifacts. Pip is still your distribution mechanism, virtualenv is a runtime customization to change how your distributed artifacts are loaded.
Python is glue language for C/C++/Fortran/Rust/etc. projects. Most of the valuable packages in Python in fact C++ projects (like Tensorflow, Pandas, etc.). When you are talking about Python dependency management you are talking about the Cartesian product of the packages you are depending on, their dependency in terms of compilers and header files. It is not hard to run into missing .h files while trying to install these dependencies. This situation was improved significantly a while back with the WHL files.
Anyways this whole article does not make too much sense, I am using Python for 10 years professionally and never run into the concept of nested venvs let alone somebody trying to use that in production.
They would be if python provided the .h files in the packages rather than expecting them to already exist on the system.
(Nb I’m a maintainer of a python project that’s difficult to build, and the number of issues we get that start: can’t install because foo.h can’t be found is very high, even though there’s other explanations closer to the end of the failed build)
I'd say maybe 25-50% of developer linux workstations have the appropriate packages and will compile out of the box, and if they don't, it's almost always a one line command to install them with a package manager. We've got instructions or docker images for testing for most major linux distros.
MacOS is probably in the 5-10% range for successfully compiling out of the box, some of the headers ship with MacOS, but there are a couple steps to go through. M1 is still a WIP in the python packaging community.
Windows... 0% compile out of the box. You have to set it up, download dependencies and so on. There are at any one time a handful of people who can compile for the windows platform, and a smaller handful of people who can debug problems.
Now, to be fair, I think this project is probably in the most complicated 1% of python projects from a packaging point of view. There are probably more complicated ones, and the only two complications I'm aware of that we don't have to deal with are 1) The chicken and egg of if we break an update we lose the ability to update all the other packages including our own (pip) and 2) GPU drivers.
One reason people like PEP 582 is that, by being something built into the interpreter, it will reduce a lot of the confusion over which isolation tool to use.
Part of the problem, which local installs won't completely solve, is that the python interpreter is still changing fast enough that modules are having trouble keeping up. E.g. tensorflow is still working on python3.9 support [1], and now we'll see python3.10 in the next couple weeks. Even with local installs we'll have to be careful to use the right interpreter.
Another piece of the problem, which is also kind of a good thing, is the huge size of the python module ecosystem. I've got python scripts with over a hundred requirements. It saves me from writing a lot of code, but resolving the requirements is a challenge. The pip package manager has recently added a dependency resolver, which has helped a little make sure everything can work together.
The only surefire ways I can think of to make this huge module ecosystem interoperate cleanly are (1) have every module support every other version of every other module and interpreter (insane amount of work, won't happen) or (2) have simpler interop interfaces like the unix shell convention of only passing strings between programs, but that would take away a lot of the ease-of-use of passing rich objects around between modules.
The current situation is actually not that bad. With my students, the topic is quickly explained and mastered. But it's not a common occurrence.
However, it is true that it's still not as easy as it could be. Some reasons.
- Python is older than java.
- Python is shipped as part of MacOS and Linux, but using the system Python can cause issues, so it's a trap.
- Python deals with a lot of compiled extensions. Scipy embeds C, fortran and assembly. It's a hard problem.
- The Python project has *tremendously* less resources than java.
- Politics. Lots of it.
- Debian et Red Had packagers made a mess by splitting the python setup in many small packages.
- Python 2/3 transition took attention span from the problem during a decade.
- You often install many python versions on the same machine, which makes things harder.
- Python is used by a huge number of non programmers. Much more than java. This shows in the reported difficulties.
- There are way too many ways to install python. So many combinations of failure modes.
- Microsoft screw up with their store not packaging "py" or the core dev by not packaging it on Unix.
- Anaconda brings as many problems as it solves, and is popular because of enterprise limitations.
"Machine learning" here is crucial I suppose. Sklearn, keras are big beasts.
Would you install something like django/flask site than I don't think you would had any problems.
Did you use conda (anaconda/miniconda) or pip?
It's super easy, barely an inconvenience
Gradle is a hot mess and JavaScript is the just a complete shit show
I ultimately fixed this with going to the conda files list, and choosing exactly what I needed. Not straightforward at all.
Regardless of whether or not you choose to _deploy_ with Docker, _developing_ in Docker containers using the VS Code Remote extensions really solved all of Python's (and Javascript's) annoying packaging problems for me, with the bonus that you also get to specify any additional (non-Python) dependencies right there in the repo Dockerfile and have the development environment reproducible across machines and between developers. YMMV, of course, but I find this setup an order of magnitude less finicky than the alternatives and can't imagine going back.
I have resorted to using a throwaway dev container which mimics the local storage of packages. Hopefully someone find the below bash function useful.
The only issue I have using this method is that I cant use VSCode to Debug, which I am hoping to find a nice solution for, but nothing yet.
function python() {
docker run \
-it --rm \
--name python_$(pwd | sed 's#/home/##g' | sed 's#/#.#g')_$(date +"%H%M%S") \
-u $(id -u):$(id -g) \
-v "$(pwd)":"$(pwd)" \
-w "$(pwd)" \
-e PROJECT_ROOT="$(pwd)" \
-e PATH="$(pwd)/vendor/bin/:${PATH}" \
-e PYTHONUSERBASE="$(pwd)"/vendor \
-e PYTHONDONTWRITEBYTECODE=1 \
-e PIP_NO_CACHE_DIR=1 \
--net=host \
-h py-docker \
python:3-slim \
/bin/bash -c '/bin/bash --rcfile <(echo "PS1=\"[$(python --version)]:: \"") -i'
}https://ntietz.com/tech-blog/drawbacks-of-developing-in-cont...
The author highlights some of the problems of using docker in development.
I probably should have been more careful to qualify my enthusiasm above. I’m definitely not suggesting that everyone that has a Python packaging problem should go out and solve it with Docker. As with all things, there are trade-offs and you should satisfy yourself that Docker is the right tool for your particular problem before diving in.
I just wanted to highlight the fact that - as a VS Code user who has some experience with Docker and writes in a couple different languages and ecosystems - I’ve had a great experience with this modality and find that overall it simplifies my workflow.
1. What does PEP 582 differ from virtualenv, it is just another "virutualenv" The biggest difference, as you can see, is __pypackages__ doesn't carry an interpreter with it. This way the packages remain available even the base interpreter gets removed. 2. I prefer Poetry/virtualenvwrapper/Pipenv It is fine. If you feel comfortable with the current solution you are using, stick with it. I am not persuading that PDM is a better tool than those. Any of the said tools does a great job solving its own problems. PDM, IMO, does good in following aspects:
- Plugin system, people can enhance the functionalities easily with plugins - Windows support, PDM has completion support for PowerShell out of the box. Other CLI features also work as good as *nix on Windows.
It is decidedly not true that I want to update my venv for every minor version bump!
I deploy to a cloud service w/ a specific version; my package manager is slow to update; I develop collaboratively using a shared container w/ a fixed version.
Updating shared resources with every new release is not always realistic, and that makes it so that I do need (or at least want) to use a virtualenv.
I see your point, but testing against every minor version is a great way to spot potential problems (that you'll need to deal with anyway).
Until you run into a minor version bump that patches a critical security vulnerability, and then you suddenly may very much want to (and in some corporate environments, have a regulatory requirement to)
Also, my use case is 100% non critical from a security standpoint, so I can afford to be careless... your point definitely stands in sensitive environments
So careful, far from stable.
I reported it (https://github.com/frostming/pdm/issues/247) because the project is great and I want it to succeed, but I won't put it in prod any time soon. I prefer my package manager to be stable than fancy.
It's also why I don't use poetry and so on in training. There is always some bug that students will encounter later because there are so many config combinations they can encounter, but recent tools have not been as battle tested as we think: the enterprise world is a wild beast.
True, bugs can't be exposed by a few users. People have vastly different environment setup for their Python development. System python, in-venv, pyenv, asdf, homebrew, etc. So I post it here to seek for help testing.
The sematic versioning may be a bit misleading, it is just because there was a big breaking change to support PEP 621, and I marked it as "Alpha" stage.
Thank you for working on such a useful project, and so seriously.
I immediately encountered a second different error, and I'll also take the time to craft a good bug report for it.
$ mkvirtualenv -a $(pwd) new_venv
$ pip install -r requirements.txt
When you want to activate this env and cd to the directory where you created it, you can simply do:
$ workon new_venv
That's all you need to know. It just works.
Being completely unfamiliar, their site says:
> Poetry either uses your configured virtualenvs or creates its own to always be isolated from your system.
So it is just a pretty wrapper? I’m missing the value add.
those are the ones off the top of my head. use poetry, it's great.
Handles dev dependencies, scripts, virtualenv management, upgrading single dependencies, has a lockfile, is super fast, is built on pyproject.toml, etc etc.
Honestly, just use Poetry. At the very least it’s better than ad-hoc personal scripts built around manually managed virtualenvs.
The "poetry run" command is extremely nice. You can do something like "poetry run pytest" and put that in a Makefile and be done. You don't have to worry about making sure you're in the virtualenv. I basically don't drop into a full shell anymore (although "poetry shell" lets you do that) because I don't like having to juggle whether I'm in a virtualenv or not.
"poetry build" will build your project into tarballs that pip can install. It's not earth-shattering, but it's nice to have it one tool.
The value add is a large number of QoL improvements. I don't think there are any features that other solutions can't compete on; Poetry just makes it easier to do the same things. I would highly suggest it for anything where other people are going to work on it, because of the QoL, but for personal projects I don't think it really matters.
It's a great tool but there are reasons you might want something else.
Alternative like raw pip, pip tools, or dephell all have various pros and cons.
What do you mean it doesn't work? It works fine on Windows here, unless I am missing something.
- poetry shell didn't start the same shell that you were in: https://github.com/sarugaku/shellingham/issues/42. Looks like they fixed it 3 days ago though.
- It doesn't set the name of the project in the prompt, so you actually don't know you are in the shell. Very annoying if you have many windows.
I agree with you there, it's very annoying and I've fallen for it a few times. Having not used Poetry on other platforms yet (still using Pipenv in a few projects), I wasn't even aware that this was a feature.
A basic example of something that would break if the environment is not activated is binaries, like mypy.
You have a script that calls "mypy"? Gotta be in the virtual environment to invoke it, otherwise the executable might not be found, or worse, it might run another executable.
One of the thing that the virtual environment does is modify the PATH to lookup executables from the virtual environment first.
This is to ensure that when a script calls binaries like mypy/python, it's going to find the same binary from the environment it's currently running in.
Typically if the environment is not activated, python is going to lookup to /usr/bin/python and might be something else.
This is most noticeable when there are different versions of python on the systems and when having scripts that invoke themselves.
I really like writing one-off or sporadically used commandline tools. Loading a virtualenv or navigating to a directory before executing kind of sucks. Generally, I install into my homedir. I haven't read through or tried this out, but it seems like it might better address this use case.
You're right, this seems to be very similar. To me, it looks like it removes one abstraction layer of pushing-popping the environment.
Another issue with wrappers is how quickly it executes. It's not a big deal for most executions, but for --help or building up a command pipeline it's really annoying to wait a second or two on each execution. That slow startup time also shows up when you start integrating it into automated processes.
Regarding nested virtualenvs, last time I checked there was no such thing. When nested, they always linked back to the system level python, and the cloned python had loads of dependencies on the global system. I'm not aware of a valid use case for nesting virtualenvs.
There is one handy feature of virtualenv that I'm not sure __packages__ solves though: console_scripts and entry_points. setuptools will automatically create a console script under /usr/local/bin or similar which hooks in your module under site-packages. virtualenv includes a ./bin directory and adds it to your path so that these console scripts still work locally. Any thoughts on this?
Pytest, pylint, black, pip, poetry, mypy all have one.
Jupyter one is buggy.
Is there a best workflow for using this inside a container?
Two thoughts:
1. Part of the good aspect of the node_modules is its nukeability. It's a good thing that we can blindly nuke and `npm i`
2. With node and react (I have 5 node projects and 2 react projects that I work on regularly), I really rarely have such error that I need to nuke the node modules. `npm i` to update dependencies and restarting the server because it got stuck in a weird state is usually the solution. I nuke a node_modules maybe every 3 months, probably less. On react-native, yeah it's still very much a thing though
Python isn't at all the only language to suffer this problem. e.g. "DLL hell" and its variations like Haskell's "cabal hell". With Ruby, if its packages are installed with bundler, then these packages are installed locally in the project directory.
But it's not as quick and clean as "python -m venv venv", so everyone uses that (and pyenv for different versions).
while python install packages globally and only globally. so a virtualenv is needed to isolate different projects from one another
pip install -r requirements.txt -t libs
and added the folder to PYTHONPATH. I never bothered with virtualenv.
this is why you need to activate virtualenv every time you open a new command line, so that it can set ENV variables for the current session and force python to use the locally scoped folders.
While in C# et all. the compiler knows to check files locally. and you don't need to activate (or deactivate) environment at all.
Looks like a ridiculous idea to me as node and npm dependencies/node_modules is completely messy and inefficient.