Alpine makes Python Docker builds slower, and images larger
pythonspeed.com
pythonspeed.com
First "the Dockerfiles in this article are not examples of best practices"
Well, that's a big mistake. Of course if you don't follow best practices you won't get the best results. In these examples the author doesn't even follow the basic recommendations from the Docker Alpine image page. Ex, use "apk add --no-cache PACKAGE". When you're caching apt & apk, of course the image is going to be a ton larger. On the flip side he does basically exactly that to clean up ubuntus apt cache.
The real article should have been "should you use alpine for every python/docker project?" and the answer is "No". If you're doing something complicated that requires a lot of system libs, like say machine learning or imagine manipulation - don't use Alpine. It's a pain. On the flip side if all you need is a small flask app, Alpine is a great solution.
Also, build times and sizes don't matter too much in the grand scheme of things. Unless you're changing the Dockerfile regularly, it won't matter. Why? Because Docker caches each layer of the build. So if all you do is add your app code (which changes, and is added at the end of the Dockerfile) - sure the initial build might be 10 mn, but after that it'll be a few seconds. Docker pull caches just the same, so the initial pull might be large, but then after that it's just the new layers.
In which case... why do you bother with alpine in the first place?
Never had a business driver come up for going with Alpine though.
There are cases in which it does matter. Just like anything else, it strongly depends on your use case. If you're into the microservices thing and rebuild your containers, or change requirements frequently - maybe container size (as they'll be pulled a lot, by a lot of different hosts) matter. Maybe you're making something for the public consumption and want to make sure it doesn't take up a huge amount of space. Maybe you're making an image for IOT/RPi type devices. You get the idea.
Personally, I like using Alpine where possible because it's got less stuff. Less software means less things that could potentially have a security issue needing fixing/patching/updating later.
However my default container for anything else is "miniubuntu" build as it's got all the basics, it's 85mb in size, and I can install all the things I need for the more complicated projects.
COPY requirements.txt ./
RUN pip install -r requirements.txt
If you're building the image on a CI server, docker can't cache that step because the files won't match the cache due to timestamps/permissions/etc... The same is true for other developer's machines.This is a problem if your requirements includes anything that uses C extensions, like mysql/postgresql libs or PIL.
https://docs.docker.com/develop/develop-images/dockerfile_be...
> For the ADD and COPY instructions, the contents of the file(s) in the image are examined and a checksum is calculated for each file. The last-modified and last-accessed times of the file(s) are not considered in these checksums. During the cache lookup, the checksum is compared against the checksum in the existing images. If anything has changed in the file(s), such as the contents and metadata, then the cache is invalidated.
If the CI system keeps the source tree the Dockerfile is being built from around rather than removing it all after every build, it caches stuff as normal.
1. Using poetry which keeps a version lock file so all changes are reflected/cached, or
2. Doing a similar thing yourself by committing `pip freeze` and building images from that instead of requirements.txt.
It does if you practice continuous deployment, or even if you use Docker in your local dev setup and you want to use a sane workflow (like `docker-compose build && docker-compose up` or something). Unfortunately, the standard docker tools are really poorly thought out, beginning with the Dockerfile build system (assumes a linear dependency tree, no abstraction whatsoever, standard tools have no idea how to build the base images they depend on, etc). It's absolute madness. Never mind that Docker for Mac or whatever it's called these days will grind your $1500 MacBook Pro to a halt if you have a container idling in the background (with a volume mount?). Hopefully you don't also need to run Slack or any other Electron app at the same time.
As for the build cache, it often fails in surprising ways. This is probably something on our end (and for our CI issues, on CircleCI's end [as far as anyone can tell, their build cache is completely broken for us and their support engineers couldn't figure it out and eventually gave up]), but when this happens it's a big effort to figure out what the specific problem is.
This stuff is hard, but a few obvious things could be improved--Dockerfiles need to be able to express the full dependency graph (like Bazel or similar) and not assume linearity. Dockerfiles should also allow you to depend on or include another Dockerfile (note the differences between including another Dockerfile and including a base image). Coupled with build args, this would probably allow enough abstraction to be useful in the general case (albeit a real, expression-based configuration language is almost certainly the ideal state). Beyond that the standard tooling should understand how to build base images (maybe this is a byproduct of the include-other-Dockerfiles work above) so you can use a sane development workflow. And lastly, local dev performance issues should be addressed or at least allow for better debugging.
This is a bit of a word soup, so I'll point you to https://buildpacks.io/ for more.
I've been considering writing a merge tool to support fork/merge semantics, focussing more on development and debugging than build-time optimization.
Just in case anyone is wondering, this is a great exaggeration. Idling containers are close to idling (~2%cpu currently), and slack got pretty small last year. These work just fine, I wish the trope died already.
> This is because there is no native support on Mac for docker. Everything has to run in a virtualized environment that's basically a slightly more efficient version of VirtualBox. When people say docker is lightweight and they're running it on a Mac they don't quiet understand what they're saying. Docker is lightweight on baremetal linux, it's not lightweight on other platforms because the necessary kernel features don't exist anywhere except linux.
"grind to a halt" and "is not lightweight" are not even close to being synonymous.
So "grinds" is somewhat accurate, if you have long running containers doing very little, or you are constantly rebuilding, even if the machine does not look like it's consuming CPU.
Note that a little Googling reveals that this is a pretty common problem.
I've started experimenting with coding on a remote docker host using vscode's remote connection feature.
I'd be interested to know if anyone else had gone down this path?
this is at least one use case for 'docker-machine'
Mac Virtualbox runs linux with host only network vbox0.
Docker runs in vbox linux. (now it gets ugly)
Vbox linux brctl's docker0 (set to match vbox0 ip space) into vbox0.
Docker container is reachable by IP from mac host. All is fast and good.
Project nautilus is a pretty interesting approach to running a container on mac. In theory it should be more efficient than docker for Mac.
Disclaimer I work for VMware, but on a different team.
I solved the problem by installing inotify into the container, which django will use if present, which reduced cpu from 140% to 10%. This is a couple of months ago.
Docker with a few idle containers will burn 100% of CPU. https://stackoverflow.com/questions/58277794/diagnosing-high...
Here's the main bug on Docker for Mac consuming excessive CPU. https://github.com/docker/for-mac/issues/3499
See https://github.com/moby/buildkit. You can enable it today with `DOCKER_BUILDKIT=1 docker build ...`
There is also buildx which is an experimental tool to replace `docker build` with a new CLI: https://github.com/docker/buildx
If you have a command `do_foo` that depends on do_bar and do_baz (but do_bar and do_baz are independent) and you do something like:
RUN do_bar # line 1
RUN do_baz # line 2
RUN do_foo # line 3
I'm guessing the buildkit dep graph will look like `line_3 -> line_2 -> line_1` (linear). Unless there is some new way of expressing to Docker that do_foo depends on do_bar and do_baz but that the latter two are independent.EDIT: clarified example.
So `COPY --from=<some stage>` and `FROM <other stage>`
Also, a Dockerfile is just a frontend for buildkit. The heart of buildkit is "LLB" (sort of like LLVM IR in usage) which is what the Dockerfile compiles into. Buildkit just executes LLB, doesn't have to be from a Dockerfile.
For that matter you can have a "Dockerfile" (in name only) that is not even really dockerfile since the format lets you specify frontend to use (which would be a container image reference) to process it.
There's even a buildkit frontend to build buildpacks: https://github.com/tonistiigi/buildkit-pack Works with any buildkit enabled builder, even `docker build`.
If you want to create docker image of your own app, you probably would use:
https://nixos.org/nixpkgs/manual/#sec-pkgs-dockerTools
This will produce an exported docker image as tar file, which then you can import it using either docker, or tool like skopeo[1] (which is also included in nixpkgs).
The nix-shell functionality is also quite nice, because it allows you to create common development environment, with all tooling available that one might need to work.
Nixery can be pointed at your own package set, in fact I do this for deployments of my personal services[0].
This doesn't interfere with any of the local Nix functionality. I find it makes for a pleasing CI loop, where CI builds populate my Nix cache[1] and deployment manifests just need to be updated with the most recent git commit hash[2].
(I'm the author of Nixery)
[0]: https://git.tazj.in/tree/ops/infra/kubernetes/nixery [1]: https://git.tazj.in/tree/ops/sync-gcsr/manifest.yaml#n17 [2]: https://git.tazj.in/tree/ops/infra/kubernetes/tazblog/config...
A) Changing a Dockerfile is rare B) Typically the lines that change (adding your code) are near the end of the Dockerfile, and the long part with installing libraries is at the beginning
We have been using Google Cloud Build in production for over an year and Docker caching [1] works great. And Cloud Build is way cheaper than CircleCI.
I recommend it, and I'm not getting paid anything for it.
[1] https://cloud.google.com/cloud-build/docs/speeding-up-builds...
https://stackoverflow.com/questions/46221063/what-is-build-d...
This is why multi-stage builds are a thing, which the author advocates against doing.
Author just doesn't know better. That's what happen you never build things from source yourself.
I have an Alpine docker image which was 185MB and after I added the above, it was 186MB. I was definitely hoping for more, given your strongly worded advice.
- https://github.com/insightfulsystems/alpine-python
- https://github.com/rcarmo/ubuntu-python
The first uses stock Alpine packages, and the second builds Python from scratch (with some optimizations) atop Ubuntu LTS, and they serve two different use cases, but maintaining both made me learn a few things and there are a few factual errors in the article.
For starters, yes, you can run manylinux wheels on Alpine. Here’s how to do it:
https://github.com/insightfulsystems/alpine-python/blob/mast...
So no, you don’t need to recompile every single package for an Alpine Python runtime.
As to runtime differences, yes, they exist, but are _extremely_ dependent on your use cases. I have not encountered any of the bugs - I did have a few crashes with GPU libraries and Tensorflow (which is why I also maintain the Ubuntu-based version), but the author points out third-party accounts, one of which seems to be attributable to locale settings (which you should always set anyway).
Performance differences are negligible on Intel (at least for web apps using Sanic and asyncio - I don’t do much Django these days), and the inclusion of a link to buy a “production-ready template” mid-article is just... iffy.
Give us data and working code - I’d like to see an objective benchmark, for instance, and might even set up one with my own images.
If that's all that is needed then it makes me wonder why pip is only downloading wheels on glibc distros currently.
If you just need pandas, then you can install it with apk along with many other packages and it'll be much quicker: https://pkgs.alpinelinux.org/packages?page=1&branch=edge&nam...
Alpine are doing what they can, by maintaining common python packages on apk. It is pip that is tied to glibc
I'm not saying that's the best practice. Personally I like to use distro-packaged stuff as much as possible because it's less volatile and tends to go through some more review than things on PyPI. But the virtualenv/pip exclusive method does seem to prevail in the field in my experience, so the advice to just use Alpine packages isn't useful in many cases.
In one sentence: "Installing the latest aws-cli on an image is preventing new AMIs based on that image from booting due to the above issue with urllib3."
While this specific issue wouldn't affect docker images since you normally don't run cloud-init on them, it's just luck that it wasn't some other utility affected instead. Next time it can affect docker images too.
But tlrd is: if you "apt install some-utility", then "pip install something else", you may have upgraded packages that some-utility relies on but is not compatible with the new version anymore.
I'm not saying never install globally, just let's keep in mind that this can lead to real issues which may be very surprising / hard to debug once in production. Unless you understand exactly what and how is delivered with every change, defaulting to venv is a safer option.
If you can prove it doesn't apply, then sure, why not install globally.
Now you still want to run and install pip packages as a non-root user of course, but you don't need a virtualenv in docker.
In the article, they are only using pandas and not using a requirements.txt file. So it's a pretty perfect case to just use apk. In cases where you don't need many packages, you can still use Alpine if you would like without most of the downsides mentioned in the article.
And so again need to rely on something "non-standard".
> Alpine are doing what they can, by maintaining common python packages on apk. It is pip that is tied to glibc
The "manylinux" tag means glibc by necessity, because nobody has driven a PEP for tagging, detecting and managing musl-based wheels to completion[0]. The issue is not restricted to the choice of libc, or even linux alone, either. But again, that requires people actually put in the work to chip at and fix the issue[1].
The point is though, that this is a python packaging issue. Not an Alpine issue and not an issue that Alpine can do much about. As you linked, pypa have a big task to make packages work on every possible platform. Unfortunately, this comes as an unexpected surprise for anyone who wants to try out Alpine - maybe pip should show a warning when wheels aren't available because of musl.
This looks like some sort of binary Python packaging thingy. I think @dkarp is saying that there exists pre-existing binary builds for glibc for many things, but not for musl libc so they are forced to be built explicitly when using pip to manage Python dependencies.
I could not figure out where this repository of binary packages is or who maintains it. Any corrections would be appreciated.
As others have mentioned, Alpine's repository does include Alpine-compatible binary versions of many popular Python packages. So you can often save yourself a lot of trouble by using apk instead of pip to get those.
That’s a false statement (depending on interpretation). Whether a PyPI package has wheels is entirely at the discretion of the package maintainer, who is responsible for compiling all source and binary distributions; PyPI does not compile anything and only accepts uploads, verbatim. If the maintainer doesn’t upload wheels or doesn’t upload wheels for your platform, then no wheels for you (not from PyPI). In general compiling statically linked wheels (when you have C extensions and external dependencies) is a fairly involved process.
At the moment there are six platform tags for Linux wheels supported by PyPI: manylinux{1,2010,2014}_{x86_64,i686}. Each is based on a CentOS release with very old glibc. See
https://github.com/pypa/manylinux
PEP 513 - maynylinux1 https://www.python.org/dev/peps/pep-0513/
PEP 571 - manylinux2010 https://www.python.org/dev/peps/pep-0571/
PEP 599 - manylinux2014 https://www.python.org/dev/peps/pep-0599/
This is a problem I have with Python: everything is assumed to be a standard GNU/Linux/desktop environment, you are assumed to always have to have the latest versions of everything (or to have "up to" a certain version, which then breaks everything else that relies on the latest), and all becomes sub-optimal if you deviate a bit from the expectations.
Now this is fine for some people but you can't blame the ones that have different environments. This is why most people doing actual embedded cringe when someone suggests using Python, because the experience is not usually pleasant.
For many applications, particularly those that aren't written with low-level performance details in mind, this is likely to cause significant slowdowns (obviously not the full 50% as many things will be CPU bound, but still likely noticeable).
More generally, it seems that Alpine optimises for build/deployment, but glibc-based distros are going to be faster at runtime. While a fast build/deploy is nice, it's probably not worth it if it comes at a cost of significant runtime performance issues.
I work on a Python application where we don't care about memory allocations (we care about the number of database queries for example). Using musl may make memory allocations expensive enough (as they are 50% slower) that we may have to start paying attention to this in some areas of our code.
E.g. something like:
FROM python:3.8-alpine
RUN apk add --no-cache --virtual build-dependencies gcc build-base freetype-dev libpng-dev openblas-dev && \
pip install --no-cache-dir matplotlib pandas && \
apk del build-dependencies
This way you won't keep the build dependencies in the layer and the final image.
Maybe not all of the packages are build-dependencies, but at least there is no need to keep the gcc and build-base package around.
Long story short, wrap your "pip install", "bundle install", "yarn install" around some "apk add .. && apk del" within one RUN.Yes, it would still compile for quite some time, but the final image size is not "851MB" but "489MB". Still larger than the "363MB" of the "python-slim" version. Guess "pip install" will keep some build artifacts around?
>-t, --virtual NAME
>Instead of adding all the packages to 'world', create a new virtual package with the listed dependencies and add that to 'world'; the actions of the command are easily reverted by deleting the virtual package
`pip install --no-cache-dir -r requirements.txt`
also you dont need clean up the builder, since the unused files are not used in the final image.
also add a new file you can just append it to the builder. (not rebuild required of the previous installs required since you don't care about the size of the builder image )
FROM python:3.8-alpine as builder
RUN apk add --virtual build-dependencies gcc build-base freetype-dev libpng-dev openblas-dev
RUN pip install matplotlib pandas
FROM python:3.8-alpine
# a couple of copy commands to copy binaries and py files from builder to the final image.
COPY --from=builder /usr/bin /usr/bin
COPY --from=builder /usr/local/lib/ /usr/local/lib/I would never want any build tools in the final Alpine image.
https://github.com/dmuth/splunk-twint/blob/master/Dockerfile...
To this day, I'm amazed it worked. But it took build time from 15 minutes down to 30 seconds.
I will have to revisit this, however, because I would prefer to have just one container that builds quickly instead of having to do a workaround like that.
If you include the dev dependencies and compiler toolchains in the final image. You can (and should) cut all that out of the final image.
> The image size can be made smaller, for example with multi-stage builds, but that means even more work.
...which I really don't find compelling, honestly. If you want a smaller image, then yes, you need to do one of the major things that produces smaller images. Don't install a whole dev chain, leave it in the image, and then complain that your image is big.
I am guessing with Python the difficulty is in enumerating what files you need to copy, because there's more than one. But I am sure that can be figured out exactly once and you can save yourself some time and money on unnecessary data transfer.
> Alpine doesn’t support wheels
You mean wheels do not support musl, or to be more precise they don't conform to POSIX 2008 or C11 standards.
> Alpine has a smaller default stack size for threads, which can lead to Python crashes.
You're linking to a Python bug. That's right, Python. This is a bug in Python. Not Alpine, nor musl.
The whole point of using Alpine-based Docker images is that musl provides a lean, efficient and standards-compliant implementation of libc which should in turn result in faster runtime execution of the software running in your container, and other secondary benefits such as smaller binaries and faster build times. If your software (in this case, Python) is glibc-dependent and is not standards-compliant, the bug is in your software, not musl. Don't expect it to work. The bottom line is that if you don't understand this, you shouldn't be using Alpine. Please don't trash-talk something you don't understand.
I think we will eventually reach consensus that package managers requiring a compiler is a misfeature.
I think I’d be just fine having two base images, one for runtime, and a superset image for Continuous Integration that layers the compiler toolchain on top.
But Node has this same behavior and it drives me nuts.
What's your alternative solution? For simple native libraries there's always libffi, but things that need numerical processing will want to actually plug into the underlying value representation directly to not sacrifice speed.
This isn't a new problem if you use docker. To address it you have to either use multistage builds or more traditionally you install the dev tools, build and install your artifacts, then remove your dev tools all in a single step. It's annoying, but it's a common pattern.
If you need more than a few packages installed using pip, then Alpine simply is not a good choice for python.
Unfortunately, this requires building a custom binary linking jemalloc and calling into Python C APIs to configure the embedded interpreter and that's probably prohibitively too much effort for many users, who view the Python interpreter as a pre-built black box and therefore see musl's allocator as an unavoidable constant. Fortunately, tools like PyOxidizer exist to make this easier. (PyOxidizer supports using jemalloc as the allocator with a 1 line change and jemalloc can deliver even more performance than glib'c allocator.)
Which memory allocator to use for the PYMEM_DOMAIN_RAW allocator. Values can be jemalloc, rust, or system.
Important: the rust crate is not recommended because it introduces performance overhead.
From the docucmentation
Home Assistant maintains Alpine wheels for a ton of packages for their 5 supported platforms: https://wheels.home-assistant.io/
Repository of our wheels builder can be found at https://github.com/home-assistant/hassio-wheels
Pillow PyXB cffi cryptography gevent lxml msgpack-python psycopg2 pycrypto uWSGI
It's 2020, and yet we're still falling victim to textbook examples of FUD.
I just run a multistage build where I first install all the packages, building wheels were necessary. Then I copy them to a new clean image from the first container. Another benefit is that I don't need to have a compiler installed in my final image. The Dockerfile is much cleaner than having one huge RUN which installs and then uninstalls build dependencies.
The author should use a multi-stage build where one stage is used for building the missing wheels. Those should then be copied over to the final stage.
No need to install any build tools in the final Alpine (or Debian) image. The Alpine image will definitely be much smaller.
This way when we go to build our final app images, there is no need to compile any python packages or install a toolchain. Our builds are generally under a minute or two as its just downloading and installing packages via Poetry/Pip.
If you're not satisfied with Alpine and need slim builds, answer is python3.8:slim-buster (from other article).
From empirical experience, slim-buster has been a good choice for our use case of ML services.
Just my two cents.
Matplotlib? Pandas? Those are not typical except for the datascience crowd.
For your regular web stack? It works fine. And yes, installing the apk will save time in a lot of occasions.
(Assuming your data sciency workload uses those functions; if not, this doesn't apply)
In my use cases the image size download eclipsed the build time for the most part
if it's looking for glibc /lib/ld-$foo.so you could symlink musl and see if it works and if python touches some glibc specific junk just install gcompat - https://pkgs.alpinelinux.org/package/edge/community/x86/gcom...
If I wanted maximum efficiency I'd be coding in ASM not deploying dockers.