A deep dive into the official Docker image for Python
pythonspeed.com
pythonspeed.com
I can’t believe I’m just now learning this, but that’s good to know next time someone asks why the names seem random.
https://www.reddit.com/r/FanTheories/comments/1qmtoi/sid_fro...
https://matthewdicks.com/film/2018-5-23-sid-from-toy-story-w...
"<version #>-<stupidname>".
but I'm a FreeBSD user so what do I know?However every few years there might be a not so "stable" update around the new release.
Windows 10 and macOS keep it pretty simple too.
mac 10.15 means what? and the next release is gonna be 11.0. someone who doesn't follow or use macs has no idea to figure out if the latest version is 10.15 or 10.14, or 11.0...
Every now and again I have to translate between a code name and a version string, or between a version string and a code name, and I think 'why on earth am I having to do this lookup work as a human?'
Anyway, the obvious thing is that a codename is more fun and memorable than a number; it's something to hang marketing and such off of. Presumably this is why MacOS versions have always had public codenames since 2002. But I think the practical reasons are valid as well.
BUILD_ID=rolling
but I'm an Arch user so what do I know? :)Is there a reason to prefer this method, where installation, usage, and removal all happen in one RUN, vs. using a multi-stage build? I tend to prefer the latter but am not aware of tradeoffs beyond the readability of the Dockerfile.
I'm thinking the "build step" would be done in an earlier stage; it could be the exact same RUN statement, or it could be split into multiple for readability, and wouldn't bother removing any installed packages, since they won't carry over to the final stage. Then the big RUN in the final stage would be replaced with something like `COPY --from=builder ...`.
If you are doing multi stage builds it only matters to combined as many statements as possible in the last layer.
I agree that for clarity it is nice not to optimize layers in the build stage - those will be thrown away anyway.
I vastly prefer multi stage builds over having to chain install and cleanup statements
Example: I usually want to use the python:3-slim image, but this doesn't have the tools to compile certain python libraries with C extensions. Generally I will use the python:3 image for my build stage to do my "pip install -r requirements.txt" and then copy the libraries over to my final stage based on the python:3-slim image
Of course I could install and uninstall GCC and other tools in a single stage.. but that actually takes longer to do and is messier in my opinion.
Example on how to do that, please.
Specifically, multi-stage builds let you get better caching and therefore faster rebuilds, since you can cache the pre-installed gcc etc. layer, while still getting the small image.
So if you have a human being waiting on frequent Docker build results, yes, multi-stage is better.
In this case, the builds are automated, no one waits for them, so it doesn't really matter (except for burning some extra CPU cycles).
Is there no downside to multi-stage? Even aside from caching behavior I prefer multi-stage builds, as I'd much rather read & maintain a bunch of RUN lines which do one specific thing, rather than dozens joined with &&.
See here for why and how to fix it: https://pythonspeed.com/articles/faster-multi-stage-builds/
So the deal is that the system-supplied Python can get deps from apt packages, or from pip (into ~/.local, or as root into /usr/local), but either way will install them into a dist-packages directory. This keeps them separate from the packages which the built-from-source Python installs into site-packages.
Upstream Python has put a bunch of pieces in place to address this natively, for example with the "magic" tags that keep compiled assets separate in a directory of packages being potentially used by multiple different interpreters (see https://www.python.org/dev/peps/pep-3147/#proposal). And obviously the ideal solution where possible is to simply use a virtualenv and be totally isolated from the system python. But there are situations where that isn't possible or desirable, such as in this docker container, and so here we are.
Why would it not be possible or desirable to use a venv in the Docker container? The tools are available and work fine (`venv` is built into modern Python 3, and they don’t break it unlike Debian). venv will give you separation from anything unusual on the system, and let you safely install your specifically pinned requirements.txt, with zero downsides.
> Using Tini has several benefits:
> - It protects you from software that accidentally creates zombie processes, which can (over time!) starve your entire system for PIDs (and make it unusable).
> - It ensures that the default signal handlers work for the software you run in your Docker image. For example, with Tini, SIGTERM properly terminates your process even if you didn't explicitly install a signal handler for it.
> - It does so completely transparently! Docker images that work without Tini will work with Tini without any changes.
[...]
> NOTE: If you are using Docker 1.13 or greater, Tini is included in Docker itself. This includes all versions of Docker CE. To enable Tini, just pass the `--init` flag to docker run.
You can use docker (or, less efficiently, a full-blown VM) for any purpose, but it looks like the killer app that has emerged appears to be devops. I guess same goes with k8s, as well as Chef/Puppet/Salt/Ansible/whathaveyou.
However, I'm noticing there is a new way of doing s/w dev emerging, let's call it modern software development, which utilizes some of these tools to maximize s/w dev productivity. I'm just not clear what is the best way to approach this.
The core issue is obviously dependency hell, and install-reinstall-reconfigure hell.
I guess a useful way to think is, what if I want do to web-dev, and android-dev, and, iOS dev, but I don't want these dev environments to interfere with each other, and all these dev environments should be available accessible on a single workstatation or powerful laptop.
I guess I could have docker for web-dev, docker for android-dev, and so on. I came across docker compose, and then I heard it's known to be cumbersome for dev-environments, and someone created binci to address those problems (though it's not a well-known tool).
So dev environment is on the cloud? and all the data (working dir, data files, pdfs, whatever is involved during the dev phase) is also on the cloud storage?
That would be expensive, especially if you're a indie dev, and a tremendous waste of local compute/storage capabilities.
Again, I'm really not sure what near-future looks like in this space.
edit: ok reading the article further, there are handwavy explanations. Don't call it a deep dive if you're gonna say
"There’s a lot in there, but the basic outcome is:..."
and then not explain anything beyond that.
But if you want Python 3.8, the soon-to-be-released Python 3.9, or other versions, you don't have that.
That's surprisingly new even. I remember Debian being very behind when it came to Python3.
Is that really a common mistake?
(And there's likely what to clean up in the search query to make it filter more irrelevant results).
Seems like though "yes" is the answer.
There are many packages that have Python as a dependency these days. For example, on my Ubuntu system:
> ~$ apt-cache rdepends python|wc -l
> 4649
I think the best illustration of how this can happen is installing postgres libraries needed to build the psycopg2 PG client. If you know to install `libpq-dev` then you're great. But if you do something that on the surface feels totally reasonable, like installing the `postgresql-client` package... guess what? You just installed another Python interpreter.
edit: formatting
[1] https://github.com/ContinuumIO/docker-images/blob/master/min...
[2] https://github.com/ContinuumIO/docker-images/blob/master/min...
If you build manylinux wheels with auditwheel [3], they should install without needing compilation for {CentOS, Debian, Ubuntu, and Alpine}; though standard Alpine images have MUSL instead of glibc by default, this [4] may work:
echo "manylinux1_compatible = True" > $PYTHON_PATH/_manylinux.py
[3] https://github.com/pypa/auditwheel[4] https://github.com/docker-library/docs/issues/904#issuecomme...
The miniforge docker images aren't yet [5][6] multi-arch, which means it's not as easy to take advantage of all of the ARM64 / aarch64 packages that conda-forge builds now.
[5] https://github.com/conda-forge/docker-images/issues/102#issu...
[6] https://github.com/conda-forge/miniforge/issues/20
There are i686 and x86-64 docker containers for building manylinux wheels that work with many distros: https://github.com/pypa/manylinux/tree/master/docker
A multi-stage Dockerfile build can produce a wheel in the first stage and install that wheel (with `COPY --from=0`) in a later stage; leaving build dependencies out of the production environment for security and performance: https://docs.docker.com/develop/develop-images/multistage-bu...
I assume the main benefit of using these images would be if you are installing from conda repos instead of pip? Otherwise just using the official python images would be as good if not better
Edit: I guess if you needed multiple python versions in a single container this would be a good solution for that as well
- Already-compiled packages (where there may not be binary wheels) instead of requiring reinstallation and subsequent removal of e.g. build-essentials for every install
- Support for R, Julia, NodeJS, Qt, ROS, CUDA, MKL, etc.
- Here's what the Kaggle docker-python Dockerfile installs with conda and with pip: https://github.com/Kaggle/docker-python/blob/master/Dockerfi...
- Build matrix in one container with conda envs
Disadvantages of the official python images as compared with conda+pip:
- Necessary to (re)install build dependencies and a compiler for every build (if there's not a bdist or a wheel for the given architecture) and then uninstall all unnecessary transitive dependencies. This is where a [multi-stage] build of a manylinux wheel may be the best approach.
- No LSM (AppArmor, SELinux, ) for one or more processes in the container (which may have read access to /etc or environment variables and/or --privileged)
- Necessary to build basically everything on non x86[-64] architectures for every container build
Disadvantages of conda / conda+pip:
- Different package repo infrastructure to mirror
- Users complaining that they don't need conda who then proceed to re-download and re-build wheels locally multiple times a day
Additional attributes for comparison:
- The new pip solver (which is slower than the traditional iterative non-solver), conda, and mamba
- repo2docker (and thus BinderHub) can build an up-to-date container from requirements.txt, environment.yml, install.R, postBuild and any of the other dependency specification formats supported by REES: Reproducible Environment Execution Standard; which may be helpful as Docker Hub images will soon be deleted if they're not retrieved at least once every 6 months (possibly with a GitHub Actions cron task)
If you’re always on Linux you may never appreciate it but some pip packages are a nightmare to get working properly on Windows.
If you look through the source of the conda repos, you’ll see all kinds of small patches to fix weird and breaking edge cases, particularly in libs with significant C back ends.
It includes patches just like distro packages often do.