Docker-slim: Minify your Docker container image without changing anything
github.com
github.com
It's a base image with binaries from Debian deb package, and with necessary stuff like ca-certificates, and absolutely nothing else, while still glibc-based (unlike Alpine base images).
Example images I built with the base image - C binary, <10MB https://hub.docker.com/r/yegle/stubby-dns
- Python binary, <50MB https://hub.docker.com/r/yegle/fava
- Go binary, 5MB https://hub.docker.com/r/yegle/dns-over-https
Another trick I use is to use https://github.com/wagoodman/dive to find the deltas between layers and manually remove it in my Dockerfile.
I expect someone will leave a comment saying "But you shouldn't be entering containers, you should be using Ansible/Kubernetes". Yes, that is how I manage changes but sometimes you just have to log in and see what is going on with htop/etc
docker exec -it --augment=ubuntu my_container bash
Would start bash in the container, but also layer into the filesystem all the rest of a standard ubuntu image only for my tools, but not affecting the application in the running container.I'm pretty sure that's possible with current linux kernel mount namespace/overlayfs infrastructure used by docker - all that's needed is the command line tool to support it.
(1) If it's the bash version from the standard ubuntu image, you will need to specify where to mount your application's filesystem inside the ubuntu filesystem.
(2) If it's the bash version from your application, then it's the other way around: you will need to specify where to mount the ubuntu filesystem inside your container.
Option (1) seems more practical. My point is that you will need to specify a mountpoint either way, and your commands will need to take this mountpoint into account.
For docker and k8s, there are two helpful tools which implement what I said with simple and intuitive UI:
* https://github.com/zeromake/docker-debug
* https://github.com/aylei/kubectl-debug
Edit: Add links for the helper tools.
[0] https://github.com/docker-slim/docker-slim#debugging-minifie...
One common workaround floating around the internets is to use --cap-add SYS_PTRACE. This has the side effect of permitting the ptrace syscall, but it also gives you the ability to ptrace processes owned by other users etc. That's more than you need and it's kind of dangerous in a production-ish container.
(One can still attach after everything's running, but that's not always good enough.)
For simple shell access, use the :debug variant of distroless images which include a shell.
For more complex troubleshooting, I think other people has recommended many ways. I haven't had the need to do such troubleshooting but if I need to I would mount an image with necessary binaries into the container. This is where distroless becomes handy: I can mount a Debian image and don't worry about ABI compatibility.
https://github.com/docker-slim/docker-slim#debugging-minifie...
Generally you build a special debug image that has busybox or whatever. In case of distroless, debug image has busybox and everything that comes with it.
Also, what are you trying to see with top/htop? In ideal world you will see a single process pid 1 that is your entry point. There shouldn't be more than one process running in it.
You can get resource consumption of a container without logging into the container just like you can get running processes without getting into container.
There is nothing else you can do without dragging wholelot of dependencies:
- Anything java related will require a JDK
- Debugging any native code will require a whole debugger
- Debugging python/ruby will either work or will require dev dependencies
Sidenote: Who the fuck uses ansible to debug containers?
I just wish there was a way to do the basics:
1. Look at files within my running container (maybe even modify them, without needing vim or nano installed inside it).
2. Ping/ICMP something from within the container (again, without ping being in the container itself)
3. DNS lookups from within the container
4. Connect to a port on an IP or DNS name from within the container
5. Inspect the contents of a dead container that won't start without having to commit it first.
I did a post a while back on how I feel about debuggin within containers, and I should probably write another one because I don't think I cover those 5 things:https://battlepenguin.com/tech/my-love-hate-relationship-wit...
You obviously won’t get the same operational state but if you want to poke around a container you’ve built and see what’s in it, you can just extend it.
My rule of thumb is: as soon as I have to `docker exec` into a running container because something's wrong, this container needs to be stopped and a VM should be used instead.
Curious to know what the benefit conferred by it being 'glibc-based' is exactly?
[1] Example- https://www.redhat.com/security/data/oval/
[2] This is the closest you can get- https://github.com/alpinelinux/alpine-secdb
It works well until it doesn't. And screws up your whole deployment pipeline.
In my limited experience, some things don't compile targetting musl. If everything works alpine is sup. Else, it may be fairly difficult (and mostly unjustified) to fix.
this: https://github.com/grpc/grpc/issues/18150#issuecomment-47999...
And this: https://github.com/pypa/manylinux/issues/37
I use Python. Anytime you need to compile the source to build a pip package and that the upstream package developer decided that they don't support Alpine's libc implementation (aka musl) then you will have a big problem, unless you can control your dependencies and include as few pip packages that requires compilation or find binary builds.
For example:
from python2.7:distroless - 60.7MB => 18.3MB (minified by 3.32X)
FROM scratchscratch < distroless ~= alpine < debian-slim
Distroless nodejs images is...10mb bigger than the same on alpine.
Main purpose of distroless is less attack surface rather than size. Without package manager mutating container is PITA. Not having `ps` or `cat` makes it hard to read secrets that you injected into container one way or another.
1. Start with a small base image, e.g. for Python there's "python:3.7-slim". For Python I'm not a fan of Alpine, but for Go that gives you an extra small base image (see https://pythonspeed.com/articles/base-image-python-docker-im...).
2. Don't install unnecessary system packages (https://pythonspeed.com/articles/system-packages-docker/).
3. Multi-stage builds (in Python context, https://pythonspeed.com/articles/smaller-python-docker-image...).
You can find similar guides for non-Python as well. Basic idea being "don't install unnecessary stuff, and in final image only include the final build artifacts".
I think what's more important is layering the Dockerfile in the right order. You should be putting most of your large, infrequently changing assets in the lowest layers. Then put smaller, more frequently changing assets in the top layer. If you have a 4GB image, but only change the top 10MB layer, then it only requires caching 10MB of new data when you update the container. But if you change a lower layer, then it requires re-building and re-caching everything above it.
I don't think the concern is 'how long does deployment take' but 'how fast can we iterate?'; Building and loading the images on a local dev machine to test a 2 second change would take much longer with larger images. Getting feedback of a PR merge from a CI build agent would take minutes longer.
I don't think the importance has been oversold.
And for iteration, the only thing that matters is the size of the top layers that are being iterated on. Not the overall image size itself.
You can put `RUN apt-get [kitchen sink]` at the beginning of the Dockerfile and it pretty much won't matter. When you change anything in the project repository, that bigass giant base layer doesn't get re-pulled because it doesn't change.
To validate that a layer, Docker daemon just compares the hashes. So, for unmodified layers, docker pull is constant time with image size. The Docker daemon only downloads from the registry starting with the bottom-most modified layer relative to its cache.
If anything, throwing the kitchen sink in the base layer is better for fast iteration. When you're being parsimonious about third-party libraries and packages, you'll frequently have to rebuild the base image. If you `RUN apt-get [everything]`, then you'll hardly ever never need to rebuild/re-push/re-pull that layer, because you'll always have whatever you need already available.
A couple of gigs is not bad one time. But it's multiplicative, even without updating our application at all, we're transferring: [avg size] * [# of images] * [# of security updates] * [# of nodes]
Given that we are usually able to get small images just by using slim base images, it's a no-brainier for us. No, we don't pull our hair out trying to save a meg, but we're not inheriting Ubuntu as a base image for a 1MB microservice.
By using slimmer images we have fewer security updates to apply as well. Since we usually give security updates some human attention, that's fewer man-hours we have to spend on this too, which is way more expensive than bandwidth.
I learned a lot about how to keep the container image to "only" 25 GB. I had to download files into the container, start Minio (the object store) at build time, upload the files to Minio to generate some needed metadata, and then delete the downloaded copy of the file. All of this on a build server that had 80gb of storage. I tell myself I have embedded storage at scale :)
What exactly did the 25GB include?
Often images that large are the result of including build toolchains, which you can omit with multi-stage builds (e.g. for Python https://pythonspeed.com/articles/multi-stage-docker-python/.
There are very complex web applications that download and start in < 10s.
It's nice when things are fast not slow.
A tiny image for running your app, such as an Alpine Linux base with a static executable in it, is also much more secure since your attack surface is significantly smaller.
Having said that as an Australian with sub par internet at home, I do appreciate not making images gratuitously large. Just because SDK images won’t ever be deployed to prod doesn’t mean you shouldn’t strip out unnecessary junk.
Interesting concept. I wonder how it expected to cover 100% of the app usage if certain things aren't triggered during the analysis phase.
My point is merely that this is quite a significant risk. If you fail to exercise 100% of your code paths via functional testing (so you've got to have comprehensive positive and negative functional testing, which is pretty rare in my experience), you risk producing an image with docker-slim that will break. You've got to think about exercising every single possible interaction with every other component running on an OS. That's no small feat.
Think about it. That's not just 100% of _your_ code paths, that's 100% of the code paths that you could possibly ever trigger in any library that you consume, and you have to think about what might influence those circumstances. There's all sorts of angles to consider, e.g. Does latency of DNS response matter? Does time of day matter? Does IPv4 vs IPv6 matter (answer is likely yes in this case, so you might need to think about running the functional tests coming from both address stacks).
docker-slim is a neat idea, but it seems to come with significant risk.
Yes. In practice test coverage tends to be well below 100%. This is fine if you're just running tests but if you're deciding which part of your package should be pruned based on this sort of analysis then it's very likely that this will cause problems.
It seems like the biggest FAQ item is missing from the readme: What happens is my container reads a file only every so often and this tool doesn't capture it?
Also, do I need to keep images running for a while for this tool to minimize the files in the rootfs? It seems impractical, especially in headless environments like CI/CD.
func fixPy3CacheFile(src, dst string) error {
...}
Yeahhh, that is your classic example of "things you definitely DO want an explanation comment for"And, of course, it is possible that not all artifacts will be identified. There are a couple of ways to mitigate this. First, you can create custom probes for your app/service to make sure the app container can be analyzed much better. Second, you can explicitly tell docker-slim what you want to keep in your container image (you can specify files or executables)
The future version will get more different runtime monitors and it will do much more with static analysis too. Right now its static analysis is limited to LDD-like dynamic library inspection for extra artifacts you want to keep in your image. There's a lot more that can be done there...
https://github.com/tzickel/docker-trim
Last time I've checked some of their open issues about cases theirs doesn't work, worked on mine.
Also, mine is a few lines in python, if you want to learn how to trim a docker image.
As someone who works with docker containers on a somewhat daily basis I only have a vague idea of what they do under the hood and don’t have much of a reference point when comparing this slim impl to the default one.
A link about namespaces:
http://ifeanyi.co/posts/linux-namespaces-part-1/
A nice little introduction to overlays/etc:
https://jvns.ca/blog/2019/11/18/how-containers-work--overlay...
And if you really want to learn how it all works, write your own "rubber-docker" in Python:
Ideally you choose a base image that is getting security updates on regular basis, and that doesn't have significant changes over time for the same tag.
So e.g. `fedora` is bad base image, because one day you'll jump from Fedora 30 to Fedora 31. Likewise `ubuntu` isn't great. But `ubuntu:18.04` is the Long-Term Support release and so it'll be fairly stable.
In the context of Python packaging I've written a more detailed guide to choosing a base image: https://pythonspeed.com/articles/base-image-python-docker-im...
A popular base image is Debian Buster. What's in the Dockerfile?
https://github.com/debuerreotype/docker-debian-artifacts/blo...
Wow, simple, it just adds a "rootfs" archive. What's in there?
https://github.com/debuerreotype/docker-debian-artifacts/blo...
Aha. So what we have in rootfs.xz is essentially the output of a Debian system with all of those packages installed in it, and the minimal configuration needed to tie them together. No init system, no kernel, just a big filesystem full of the usual stuff an absolutely minimal Debian install would have.
Now, when you run FROM debian:buster, what you get is a base layer with enough to be useful. `apt-get` is there with a default repository list that works fine, and you can install things to your hearts' content.
Once you have gathered a mental model, feel free to tune your images by basing them on something lighter.
I am happy with my current distroless + docker multi-stage build.
The second container went from 176MB to 4.7MB!! I've not really tested that so there's a pretty good chance things aren't going to work too well in practice (but we'll see).
If it all continues as well as it's started then we'll definitely be using it.
Considering CI/CD attacks these days I spend a while scanning if this tool was some kind of joke trying to show that developers download anything from the Internet.
PS: Thanks for this :)
By the way, if you don't feel comfortable with the minification functionality you can still use docker-slim as a Docker image inspection and profiling tool (take a look at the report files it generates). Additional package level reporting is on the todo list. Image and Dockerfile linting functionality is coming soon too.