I have struggled with virtual environments and runtime executables across various OSs.
What are current best practices?
I have struggled with virtual environments and runtime executables across various OSs.
What are current best practices?
2) Run pip install. myproject/bin/activate; pip install requirements.txt; (download all project dependencies).
3) Start the application. myproject/bin/activate; python myapp.py
If you can assume that there is an interpreter available on the system, say /usr/bin/python3, you can use that instead of creating a virtual environment.
If you want to embed all the dependencies, you can save all the files created by pip install in the /lib directory if I remember the name well.
Should I write a full blog post with example? I used to deploy python applications in a bank, can explain all the advanced usage with and without internet access, with and without dependencies.
Aaaand, now that you're at it, a single .dmg file for MacOS with binaries and scripts so us lazy people with Macbooks can start coding right away :) (it might be a bit inaccurate but I guess you know what I mean)
" you can assume that there is an interpreter available on the system, say /usr/bin/python3, you can use that instead of creating a virtual environment."
Please DO NOT do this. It is so much easier to always create a virtual environment, and then you never have to worry about installing two applications on the same system that may have conflicting versions. Additionally, you don't know what else may be installed into the system's default python environment.Edit: Although if startup performance is important, using a virtual environment might be a good idea even for simple scripts. Running a "hello world" program using the system Python takes twice as long on my system as running it with "python -S" (i.e. ignoring site packages) or from a fresh virtual env.
In my case, and contrary to what one of the other commentators said, I'm using docker.
It took a while to get it ironed out, but I use a "base" folder where I have already downloaded all the packages I use by default (Top of my mind are pandas, numpy and youtube-dl ;) )
I have all the related configuration in an env file that tells pip to save the packages in the base folder so they don't disappear when I shut down the container - that I always run with --rm so they get removed when dead - and for creating a new env I just have to copy that folder. As I use a base folder for all of this I don't need to remember the names of the envs, as just need to list folders.
Only downside to this - because I'm lazy and it still hasn't bothered me enough to fix it - is that new files & folders are owned by root.
I use just a base image with python, and for different versions that work with this - supposedly - I just have to download a different image.
Learning how to use a simple barebones venv is extremely easy, saves a ton of time both in the short and long run, and generalises better.
pip install virtualenv
cd your/project/location
which python
virtualenv -p result_of_which_python env
source env/bin/activate
pip install anything_you_like
and do whatever you want from there. Those commands get 100% of the basics out of the way for noobs, and cover like 90% of the stuff you use to do more serious stuff.> but I use a "base" folder where I have already downloaded all the packages I use by default [...etc...]
Your setup sounds outrageously complex, and I don't understand why you would do any of that.
- It would work the same way on Windows, Linux and macOS
- Most importantly (!), it is not a Python package manager. That is, if a package needs BLAS or MKL or HDF5 or whatever else, you won't be crossing fingers and hoping your system-installed version would work for all your venvs; instead, those binary libraries are properly managed per environment.
Pro tip: use mamba instead of conda to get a free 4x speed boost.
That's the generalisation part I mentioned. The distinction you're making is non-trival for a wee noobie. Better they get the ground work in and expand to where they need to go.
Then they're reading on some blog about protobuf, or hdf5, or arrow or whatever, and want to use it - but using either from Python needs installed C libraries. On linux/macos you'll need to dig into your package managers and then hope things are compatible, on Windows it's a complete pain. In conda, they can just 'conda install h5py' or whatever, and get proper hdf5 installed without having to figure out the nitty gritty details.
Well, already I can tell you're disconnected from what "most noobs" are actually like. Presumably you think "most noobs" are data scientists, which just isn't the case in my experience.
It is the case in my personal experience though. I don't think I'm disconnected from what 'noobs' are as I've taught a lot of them through my line of work.
Assume you're a noob and are exploring Python package universe, fire a new private tab and google 'most popular python packages'; in my case the first 3 links are:
- https://www.activestate.com/blog/top-10-must-have-python-pac...
- https://www.ubuntupit.com/best-python-libraries-and- packages-for-beginners/
- https://www.edureka.co/blog/python-libraries/
All of which list numpy, pandas and the rest e.g. tensorflow and friends.
Please let them actually use python before sending them into tooling hell.
If you're just writing "a=b+c"-level Python scripts, or scrambling together tiny flask/django apps, or whatever else, you probably don't care indeed.
If you're doing/learning any kind of data-science-related stuff, that requires tons of C extensions and libraries. And many 'noobs' learn Python in order to use pandas/numpy/torch/whatever else is hot these days.
Of all the times I have tried conda only once was it as easy as following the guide they post. The other times it was completely broken. I also greatly dislikes that it pollutes your default environment. Imagine if every piece of software would do that.
But isn't that similar to what virtualenv does? The differences are that with virtualenv, folders are created elsewhere.
Besides that, friction so far hasn't enough for me to automate this even more (as in, creating a script that would do the folder creation) but I guess I'll do eventually.
Also, this way I have a set of python packages already installed.
And it was fun learning how to command pip / python to do things as I wanted them. That I guess answers this:
> I don't understand why you would do any of that
P.S.:
virtualenv -p $(which python) env
Saves you one line there ;)Isn't this what virtualenv was created to get away from?
If several of your projects rely on one of the "base" packages, upgrading it (which you might want for one project) could break the other projects.
Also, having implicit dependencies on packages might lead to problems with deployment (or for other developers) because they don't know that this package is required (and you might not have realized either, since it's always available on your machine).
(I might have misunderstood your solution, though. I've never used Docker.)
This other comment [1] clarifies my whole set up, and sheds a lot more light on how it works
virtualenv envThe friction here sounds like a tire fire. I think you are just not aware at the moment of how much less convoluted this can be. You've become used to this complexity.
OK. I understand where are you coming from, and I feel your pain, being a developer myself and have been way too many times that I want to admit on the "But, why?!" position.
You sound really distressed by this, and at the same time honestly curious (And baffled can probably shoved in there) so 'll explain things a little more.
I think all this starts with the fact that I explained myself rather poorly. I _do_ have different envs for different codebases. My set up is like this:
Everything python related is in a folder where I store my "dockerized" projects:
~/projects/docker/python/
In that folder I have a pip.env file that I use when setting up the container. It tells pip where to pick things up from: UN_PIP_TARGET=/work/packages
PATH=/usr/local/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:/work/packages/bin
UN_PYTHONPATH=/usr/local/lib/python3.7/site-packages:/work/packages
PYTHONPATH=/work/packages
PYTHONUSERBASE=/work/packages
PIP_USER=yes
I take it it might be a little messy (and may be thee are some redundant confs there) but it works. Once I got it working, I didn't mind about removing what was not necessary.then I have a "base" project folder, that I use as a starting point for _most_ of my projects:
base
├── packages
├── notebooks
This folder has installed a few "default" packages (pandas, numpy, matplotlib, youtube-dl).
So, as you can see the packages go all in a custom folder. This is because when I start the container, the project folder would be shared inside the container, and python / pip will pick the packages from there. If I install new packages, it will be persisted in the "env" folder system of the guest OS, so next time I "start" than environment all them packages are there.Whenever I want to start a new project, the steps I need to do is:
python.sh env_name
python.sh is a bash scripts that checks if env_name is there (as a folder). if not, it copies the base env and all it's contents (That, just so that it's extra clear, is a barebones folder _except_ for the already stated packages installed in it). In case it receives a second parameter, it uses that as starting folder (So, I could duplicate an existing project, or use an even emptier base folder).Once the folder is there, it just starts a container in interactive mode with the (environment) base folder always shared as /work
In case I need some files to work there, I just copy them.
Again, the _only_ issue so far - that annoys me -, is that if I create files when inside the container they are owned by root. Eventually I'll grow tired of this and will fix it, but not there yet.
for me starting a certain env is just one line of code. Obviously you could do the same with your method.
Said that I didn't have the script, for me creating a new env goes like this:
cp -r ./base new_env
docker run -it -rm --name new_env_container -v $(pwd)/new_env:/work -w="/work" python:3.7 /bin/bash
Removing an environment is just: rm -r new_env
What I like of this set up is that all the files related to a certain environment are in a certain place. Also, moving envs from one machine to another is rather trivial (as long as I didn't install anything not related to python in that container).This way, the container is removed every time you exit it, but it has a name so you can log into another console would you need it. the only issue with this is that it's a barebones OS, so if you need some program for it you have to install it (And that is lost of you don't keep the container: in my todo list there's a change for python.sh where it would be possible to persist the containers, and run them from the one already existing if found. The downside to this is that every container you persist is several hundred mb... But space is cheap, right?). BTW, I do have a container I persist for use with youtube-dl as it depends heavily on ffmpeg.
From my point of view, that works the same a venv. Your mileage might vary.
Caveat: python:3.7 is not the default name of the python image; I renamed it because the default was too long. I used to do it this way until I got tired of writing the same long command.
cd your/project/location
python3 -m venv env
source env/bin/activate
pip install anything_you_like
Also, on Windows, use the Python launcher[1] (that comes with the standard install of Python for Windows) to avoid having to mess with PATH settings as you would have had to historically. Instead of calling "python" or "python3" and hoping for the right version, just call "py" and tell it which version you want: py -3.6-32 -m venv env
That gets you a venv with 32-bit Python 3.6, for example. Use "py -0" to list available versions.[1] https://docs.python.org/3/using/windows.html#python-launcher...
but otherwise, i agree entirely
1/ add a Dockerfile into your project with the following lines:
FROM python:3
WORKDIR /usr/src/app
COPY requirements.txt ./
RUN pip install --no-cache-dir -r requirements.txt
COPY . .
CMD [ "python", "./your-daemon-or-script.py" ]
2/ install Docker on your various environments3/ git clone, docker build, docker run and that's it
-
- Debug containerized apps: https://code.visualstudio.com/docs/containers/debug-common
- Debug Python within a container: https://code.visualstudio.com/docs/containers/debug-python
You can apparently develop directly inside a container too:
- Developing inside a Container: https://code.visualstudio.com/docs/remote/containers
# the base miniconda3 image
FROM continuumio/miniconda3:latest
# load in the environment.yml file - this file controls what Python packages we install
ADD environment.yml /
# install the Python packages we specified into the base environment
RUN conda update -n base conda -y && conda env update && conda install -y -q moto && conda install -y -q -c conda-forge awscli httmock
# download the coder binary, untar it, and allow it to be executed
RUN wget https://github.com/cdr/code-server/releases/download/2.1698/... \ && tar -xzvf code-server2.1698-vsc1.41.1-linux-x86_64.tar.gz && chmod +x code-server2.1698-vsc1.41.1-linux-x86_64/code-server
COPY docker-entrypoint.sh /usr/local/bin/
ADD ./code /code
ENTRYPOINT ["docker-entrypoint.sh"]
Building that and running it with: docker run -d -p 127.0.0.1:8443:8080 -p 127.0.0.1:8888:8888 -v $(pwd)/data:/data -v $(pwd)/code:/code --rm -it <image>
from the directory where your code is will put those files into the container, and start a VS Code and a Jupyter Notebook server on your localhost. The password for Jypter is the default "local-development" and the password for the VS Code instance is in the Docker logs. You can set these via the Dockerfile but I just keep the defaults.
I vastly prefer this to anything else because it means I can install any packages I want without worrying about messing up my environment. You can use virtual envs to make this even better, but I am typically too dumb and lazy for that. Better part still is that my development is the same on my Mac, on my Linux machine, and on my Windows machine. Same VS Code version, same packages, etc.
Biggest issue here is with certain VS Code plugins. Some, like the vim plugin, can be finicky and depend heavily on the version of code server that you use. Some plugins break completely. However, I mainly hate plugins so this doesn't present much of an issue for me personally. I have the vim plugin, the python plugin, and a terraform plugin installed. Once they are installed, they work perfectly for me.
The way my set up works is I have a repo with that Dockerfile in it as well as the accompanying files such as environment.yml and docker-entrypoint.sh:
#!/bin/bash set -e
if [ $# -eq 0 ] then jupyter lab --ip=0.0.0.0 --NotebookApp.token='local-development' --allow-root --no-browser &> /dev/null & code-server2.1698-vsc1.41.1-linux-x86_64/code-server --allow-http --no-auth --data-dir /data /code else exec "$@" fi
and a .gitignore file with this in it:
code/*
data/*
Oh also I found the repo where I took these things from: https://github.com/caesarnine/data-science-docker-vscode-tem...
Repl.it is one way to sidestep these issues but it's not perfect. Making people jump in the deep end and use Linux helps to some extent, but also brings its own set of issues and frustrations.
I don't think "add another layer of abstraction to hide the complexity" is often a good solution. Docker brings it's own problems too.
I wholeheartedly agree here.
> Docker brings it's own problems too
This is a rather complex statement to reply to. Even thought you might be right, I don't think this applies totally to what is being discussed here.
The biggest issue I have with people advising for or against a certain tool, is that they do that from the point of view of the tool, instead of looking at fit from the problem you are trying to solve. that in your case, would be:
> I have years of experience with messing around with all of the relevant parts of Windows, Mac, and Linux and it never Just Works
As long as you manage to install Docker in all those 3 systems (I have no experience with Mac because I don't use it) both for Windows and Linux installing Docker is a no brainer.
There's a slight curve when it comes to fetching the right image and running it, but your problem is not that one; your problem is teaching Python. So you can take care of that yourself, and focus on the teaching part.
Supposing that you managed to install Docker, fetch a python image and running it (Something that is a lot more easy to do than it sounds), you have python, whatever version you want, and in an isolated way.
For me... it just works.
But this approach might be a bit limiting depending on your target audience.
Probably some combination of memory usage and complexity, depending on your application. If you're already familiar with using docker as a development environment, definitely go for it.
I don't use pipenv, I'm still using plain old virtualenv for development. Mostly it's just a matter of familiarity. If there's not an itch, why scratch?
Virtual environments are just separate copies of python with their own libraries installed.
Pipenv is basically just a workflow for organizing virtual environments
I can give someone a repo and tell them to type pip install pipenv && pipenv sync and they'll have everything. Assuming they already have the correct version of python installed, which is one nice thing docker handles, but it is easy to install python these days so it hasnt been an issue.
My biggest problem with docker was that I ended up using a ton of storage just learning about it. I have a feeling that was mostly my own fault, maybe using too heavy of a base image. I've been trying it out once or twice a year for about 6 years now, every time my conclusion is "wow this is really cool, I wish i could justify spending more time to get it right"
Though docker and virtual environments share the same problem, in that they are just a way for a developer to distribute code to other developers, and to production environments. Distributing python applications to end users is a totally different issue. I floated the idea of sending out a local data collection* app to Mac and windows users mostly because I think it would be fun to try.
*data collection of troubleshooting information from users within the same company on company hardware, that are actively asking for help. I'm not trying to spy on people.
If we're talking about web development with Python typically that means PostgreSQL and Redis too, and probably running Celery in addition to a web server such as gunicorn.
It's really nice to be able to just run a single Docker Compose command and be up and running in a way that works the same on Windows, MacOS and Linux.
This post outlines the differences between creating a Python development environment with and without Docker https://nickjanetakis.com/blog/setting-up-a-python-developme.... It focuses on the use case of web development.
And yes, setting up Python is a pain, so I think being able to start learning and running code without any setup is a big deal for learners. Get them excited about coding before they have to deal with that!
- Vagrant (so you don't have to worry about runtime executables) with the stack matching that of the production server (you can even ask ops to provide you with the provisioning script and remove the parts that you don't need - otherwise learning to provision your dev. VMs won't hurt you)
- Virtualenv
- pip
Then simply point your IDE (I only use Vim under duress) to the remote Python interpreter (the one you installed in the Vagrant VM).
It does add processing overhead but it worked with my 2010 MacBook Pro until it died and still works (only ten times faster) with the 2016 model. Your only limitation would be the RAM (I would recommend at least 8GB and if you plan to run multiple machines communicating together as much as you can afford - I do believe that Docker has less overhead, but again, for my use, not needed).
The best practice is what works for you, not the latest trend.
2) source venvname/bin/activate
then you do everything in the virtual environment...
[0] https://github.com/pyenv/pyenv [1] https://docs.microsoft.com/en-us/windows/wsl/
[1] https://havercene.io/blog/no-nonsense-python-dependency-mana...