`COPY –chmod` reduced the size of my container image by 35%
blog.vamc19.dev
blog.vamc19.dev
I wasted a lot of time in the past trying to ship with Alpine base images and statically compiling complicated software. All the performance, compatibility, package availability headaches this brings is not worth it when docker-slim does a better job of removing OS from your images while letting you use any base image you want.
Tradeoff is that you give up image layering to some extent and it might take a while to get dead-file-elimination exactly right if your software loads a lot of files dynamically (you can instruct docker-slim to include certain paths and probe your executable during build).
If docker-slim is not your thing, “distroless” base images [1] are also pretty good. You can do your build with the same distro and then in a multi stage docker image copy the artifacts into distroless base images.
The recent set of the engine enhancements reduce the need for the explicit include flags. Also new application-specific include capabilities are wip (already added a number of new flags for node.js, next.js and nuxt.js applications).
To get it Chromium under docker, I'm using the following params:
--headless --no-sandbox --disable-gpu --window-size=1920,1080 --disable-dev-shm-usageThe software required for running a Snap Store instance is proprietary [0], and there are no free software implementations as far as I know. Also, the default client code hardcodes [1] Canonical snap store, so you have to patch and maintain your own version of snapd if you want to self-host.
Snapd also hardcodes auto-updates that are also impossible to turn off without patching and maintaining your own version of snapd / blocking the outgoing connections to Canonical servers, so snapd is also horrible for server environments. To top that, the developers have this "I know what's good for you, you don't" attitude [2] that so much reminds me of You Know Who.
[0] https://www.techrepublic.com/article/why-canonical-views-the...
[1] https://www.happyassassin.net/posts/2016/06/16/on-snappy-and...
[2] https://forum.snapcraft.io/t/disabling-automatic-refresh-for...
As I said to the author of TFA a few days ago: Alpine delenda est.
And kubernetes has debug containers and the like now.
I've been using Alpine religiously for years, until the build problems became too big. Mostly long build times and removed packages on major version updates.
Now I first try with Alpine and if there is the slightest hint of a problem, I move over to debian-slim. So things like Nginx are still in Alpine for me, while anything related to Python not any longer.
At first I thought your mention of docker-slim was an error for debian-slim, but I've followed the link am glad to have learned something useful.
To debug a container, a better way is to enter the container's kernel namespaces using a tool such as nsenter [1]. Then you can use all your favourite tools, but still access the container as if you're inside it. Of course, this means accessing it on the same host that it's running.
If you're on Kubernetes, debug containers [2] are currently in beta, and should be much nicer to work with, as you can do just "kubectl debug" to start working with an existing pod.
[1] https://man7.org/linux/man-pages/man1/nsenter.1.html
[2] https://kubernetes.io/docs/tasks/debug-application-cluster/d...
Also like kyle said apt and alike(all current package managers) cant generaly deal on file per file, basis package is as is all files included, so this would require a new package manager(or old one) with packages being a single file with no dependancies for specific package. This would require new base images if we used dockerfiles and you would need to know every library every binary your program needs. While on other hand leaving in all the cruft some copy in like READMEs LICENSE files and so on.
Benefit of docker-slim is that it is in many ways just a line in your already predefined CI/CD pipeline, therefore its just another step likely just working with already preexisting technologies you might use in your pipeline.
Node.js application images:
from ubuntu:14.04 - 432MB => 14MB (minified by 30.85X)
from debian:jessie - 406MB => 25.1MB (minified by 16.21X)
from node:alpine - 66.7MB => 34.7MB (minified by 1.92X)
from node:distroless - 72.7MB => 39.7MB (minified by 1.83X)
Why are the minified Alpine images bigger than Ubuntu/Debian? Are a bunch of binaries using static linking and inflating the image? Or something else?https://www.augmentedmind.de/2022/02/06/optimize-docker-imag...
Was posted here last month: https://news.ycombinator.com/item?id=30406076
Another red flag is that you run `apt-get` after copying the binary to runtime stage(because you still want to tweak the binary there). That means any time source for binary changes, the `apt` commands need to run again and are not cached. If you just add the executable bit in your build stage you can reorder them, so the `COPY` comes after `RUN`.
The reason I did not stop with running chmod in the first stage is because this seemed like a common problem - what if I was ADDing a binary or a shell script directly from a remote source and I did not have a download stage?
I'm sure there are better ways to write that Dockerfile - I'm by no means an expert. It just so happens that I noticed this problem when the Dockerfile (it was from a different project. I was modifying it) was in this state and I had nothing better to do than ~yak shave~ investigate why the image size was a bit larger than I expected :)
The default choices are baffling in docker, it really is a worse-is-better kind of tool.
Has anyone worked on a replacement for dockerfiles? I know buildah is an alternative to docker build, but it just uses the same file format
You also have mockerfiles, being more of a proof of concept if I understand correctly https://matt-rickard.com/building-a-new-dockerfile-frontend/
But also checkout out IckFiles, an Intercal frontend for moby buildkit:
https://github.com/adamgordonbell/compiling-containers/tree/...
Nix, Guix, Bazel, Habit, and others, all solve this problem more elegantly. There are some big folks out there, quiet quietly using Nix to solve:
* reproducible builds
* shared remote/CI builds
* trivial cross-arch support
* minimal container images
* complete knowledge of all SW dependencies and what-is-live-where
* "image" signing and verification
I know docker and k8s well and it's kind of silly how much simpler the stack could be made if even 1% of the effort spent working around Docker were spent by folks investing in tools that are principally sound instead of just looking easy at first glance.
Miss me with the complaints about syntax. It's just like Rust. Any pain of learning is very quickly forgotten by the unbridled pace at which you can move. And besides, it's nothing compared to (looks at calendar) 5 years of "Top 10 Docker Pitfalls!" as everyone tries to pretend the teetering pile of Go is making their tech debt go away.
I never thought I'd come around to being someone wary of the word "container", as someone who sorta made it betting on them. There is so little care for actually managing and understanding the depth of one's software stack, well, we have this. (Pouring one out for yet another Dockerfile with apt-get commands in it.)
Bazel and company require you to clean up your ball of mud first. So your payoff is further away (and can sometimes be theoretical)
Ultimately it’s less about Docker and more about tooling supporting reproducibility (apt but with version pinning please), but in the meantime Docker does get you somewhere and solve real problems without having to mess around with stuff too much.
And of course the “now you have a single file that you can run stuff with after building the image ”. I don’t believe stuff like Nix offers that
Yes it does? Also, any nix expression can trivially be built into a much more space efficient docker container.
Nix's image building is pretty neat. You can control how many layers you want, which I currently maximize so that docker pulls from AWS ECR are a lot faster
If you want a QEMU VM for a complete system rather than a set of files for a single application, use `nixos-rebuild build-vm`, though that is intended more for testing than for deployment.
The Docker bundler seems to be using more general Docker-compatible infrastructure in Nixpkgs[4].
[1] https://github.com/solidsnack/arx
[2] https://nixos.org/manual/nix/unstable/command-ref/new-cli/ni...
[3] https://github.com/NixOS/bundlers
[4] https://nixos.org/manual/nixpkgs/stable/#sec-pkgs-dockerTool...
Uhm, can't get Nix to build a crossSystem on MacBook M1, it fails compiling cross GCC. I wouldn't say it's trivial. Maybe the Nix expressions look trivial, but getting them to actually evaluate is not.
This is the phrasing I was groping around for. Thank you
Yeah and it's competing against Dockerfiles, which I suppose in this analogy is like Python or bash with fewer footguns; syntax and parts of the functional paradigm are absolutely putting nix at a usability/onboarding disadvantage to docker.
I don't understand your point. If all you want to do is set a container image by running a shell script, why don't you just run the shell script in your Dockerfile?
Or better yet, prepare your artifacts before, and then build the Docker image by just copying your files.
It sounds like you decided to take the scenic route of Docker instead of just taking the happy path.
Meanwhile, I write packer files in HCL - a saner language and a saner format - without worrying about the way files are copied. Of course, it’s not perfect but I’d choose any of the other suggestions here before going back to Dockerfiles based on your optimism and the knowledge - that I already had but virtually every author of a Dockerfile ignores - that I can RUN a script. Thanks, but no thanks.
I run Linux as I always have. Building and running are super simple.
I feel like Docker was created more or less to let Mac devs do Linux things. Wastefully. And without a lot of reason, tbh. And of course, they don't generally even understand Linux.
Things have certainly changed with the rise of kube, ecr, and the such. But in the time of doing standard deploys into static vms, it didn't make a ton of sense.
Whatever this anecdote your team told you about Mac guys, this just has nothing to do with docker's, and containers in general, rise to fame. It wouldn't be until much later when Mac users were starting to rely on tools like Vagrant for development environments where docker was seen as an alternative to that. If your team were real linux guys, they probably would have already known about lxc, as well as all the other technologies that lead up to it: jails, solaris containers, and vserver, so seeing this as "some annoying mac thing" is especially puzzling to me.
I told a personal tale about adoption(not creation), which isn't exactly fair to the creators.
It's a slightly different and perhaps jaded view when a perfectly solid workflow is upended, and when asking why get responses like 'consistent OS and dependencies', which our vms already had, and 'we can run it locally', which half of us already did.
Admittedly, there is a lot of value in a consistent and repeatable environment specification(vs bespoke everywhere), being able to do so without needing to spin up vms, and yes - running linuxy things on Mac and Win, among other things.
That command itself means that a docker container is no longer reproducible. You cannot build it (with any code changes for your service) and guaranteed to be the same since that might be in production due to changes in the packages.
Always better to go with the base image, add your packages to the base and then use that new image as the base image for your application.
It's a tradeoff between making container images reproducible, and not shipping security vulnerabilities.
People tend to prefer the latter.
Furthermore, you can exec your way into a container and check exactly which package version you installed.
You can regenerate your base images every day or more often and have consistent containers created from an image. Freshly generated image can be tested in a pipeline to avoid issues and you won't hit issues like inability to scale due to misbehaving new containers.
How exactly does that a) assure reproducibility if you use a custom unreproducible base image, b) improve your security over daily builds with container images built by running apt get upgrade?
In the end that just needlessly adds complexity for the sake of it, to arrive at a system that's neither reproducible nor equally secure.
OP's suggestion is to build a separate image with required packages, tag it with something like "mybaseimage:25032022" and use it as my base image in the Dockerfile. This way, no matter when I rebuild the Dockerfile, my application will always work. You can rebuild the base image and application's image every X days to apply security patches and such. This also means I now have to maintain two images instead of one.
Another option is to use an image tag like "ubuntu:impish-20220316" (instead of "ubuntu:21.10") as base image and pin the versions of the packages you are installing via apt.
I personally don't do this since core packages in Ubuntu's repositories rarely introduce breaking changes in the same version. Of course, this depends on package maintainers, so YYMV.
The advantage a separate base has is allowing you to continue to update your code on top of it, even while the new bases are broken.
You could still do that without it though, just by forking out of the single image at the appropriate layer. Not as easy, but how often does it happen?
To start off, if you intend to run the same container image for 10 days straight, you have far more pressing problems than reproducibility.
Personally I know of zero professional projects whose production CICD pipeline don't deploy multiple times per day, or in the very worst case weekly in very rare cases where there is zero commit.
> OP's suggestion is to build a separate image with required packages, tag it with something like "mybaseimage:25032022" and use it as my base image in the Dockerfile.
Again, that adds absolutely nothing to just pulling the latest base image, running apt-get upgrade, and tagging/adding metadata.
The smart way of doing it would be to:
1. Use the direct SHA reference to the upstream “Ubuntu” image you want.
2. Have a system (Dependabot, renovate) to update that periodically
3. When building, use “cache from” and “cache to” to push the image cache somewhere you can access
And… that’s it. You’ll be able to rebuild any image that is still cached in your cache registry. Just re-use a older upstream Ubuntu SHA reference and change some code, and the apt commands will be cached.
Testing phase involves building Docker image from fresh system image, creating container(s) from new Docker image and testing resulting systems, applications and services. If everything goes well, the system image (not Docker image) replaces previously used system image (one without current security patches).
We have somewhat dynamic and frequent Docker images creation. Subsequent builds based on the same system image are consistent and don't cause problems like inability to scale. Docker does not mess with the system prepared by Packer - doesn't run apt, download from 3rd party remote hosts but only issues commands resulting in consistent results.
This way we no longer have issues like inability to scale using new Docker images and humans are rarely bothered outside testing phase issues. No problems with containers though, as no untested stuff is pushed to registries.
That solves nothing, as it just moves the unreproducibility to a base image at the cost of extra complexity. Arguably that can even make the problem worse as you just add a delta between updates where there is none if you just run apt get upgrade.
> Freshly generated image can be tested in a pipeline to avoid issues and you won't hit issues like inability to scale due to misbehaving new containers.
You already get that from container images you build after running apt get upgrade.
When we have VM images upon which all our usual Docker images were successfully built, we trust it more than `FROM busybox/alpine/ubuntu` with following Docker builds. I've detailed the process in a neighboring comment[1] but you're right that it doesn't suit all workflows.
I write software that a billion users see every day, so maybe I’m jaded by the sheer scale and challenges of writing code at scale that I just can’t imagine these types of problems.
Do you think EC2 capacity on AWS is on average kept in high utilization? Everyone runs (non truly elastic resources) with headroom to varying degrees
Perhaps because people do their homework and just by reading the sales brochure they understand that lambdas are only cost-effective as handlers of low-frequency events, and they drag in extra costs by requiring support services to handle basic features like logging, tracing, and even handling basic http requests.
Predictability has nothing to do with it. Volume is the key factor, specially its impact on cost.
> arrogant of you to say adopters haven’t done their homework
Those who mindlessly advocate lambdas as a blanket solution quite clearly didn't even read the marketing brochure. Otherwise they would be quite aware of how absurd their suggestion is.
I mean the unupdated files in the base image, plus the copy-on-write changes in the subsequent layers.
With Debian, there are snapshot images[3] which seem like a better approach for making apt-get reproducible. You'd simply have to change the "FROM" line in the Dockerfile to something like "FROM debian/snapshot:stable-20220316" (where 20220316 is the date of the image you are trying to reproduce, helpfully given in /etc/apt/sources.list).
With the approach you describe, you would have to carefully manage the base images: tag them, record which one was used to create each application image, and keep them around in order to reproduce older application images.
I'm sure there are situations where the approach you describe is useful (e.g. with other package managers, especially ones that don't have a notion of lockfiles), but it adds complexity and I don't think it's necessarily justified in the case of apt-get (at least on Debian).
[1]: https://docs.docker.com/engine/reference/builder/#exec-form-...
[2]: https://docs.docker.com/develop/develop-images/dockerfile_be...
But that makes your base image non-reproducible. You're just shifting the issue elsewhere.
The mistake many do is seeing dockerfiles as a 1:1 mapping of a shell script with RUN prefixed on every line. It’s not, you should only split a run if you have a good reason to add a new COPY in between for layer caching reasons or switching user.
With the --bind options from buildkit you can ensure the apt cache does not get layered and you can mount any big temporary files needed from host instead of copying them first.
I've been using Dockerfiles extensively for years and I'm yet to find anything that fits the definition of a OS quirk.
The quirkiest thing I've noticed in Dockerfiles is the ADD vs COPY thing.
> I've run into so many strange things just converting a simple predictable shell script into a Dockerfile.
What exactly are you trying to do setting up a Dockerfile that requires a full blown shell script?
A Dockerfile should have little more beyond updating/installing system packages with a package manager, and copying files into the container image. First you run a build to get your artifacts ready for packaging, and afterwards you package those artifacts by running your Dockerfile.
CMD ["/usr/local/bin/node", "server.js"]
BEHAVIOR CONFIGURATION
SOURCE MAP SUPPORT
CONTAINER INTEGRATION
WATCH & RELOAD
LOG MANAGEMENT
MONITORING
MODULE SYSTEM
MAX MEMORY RELOAD
CLUSTER MODE
HOT RELOAD
DEVELOPMENT WORKFLOW
STARTUP SCRIPTS
DEPLOYMENT WORKFLOW
PAAS COMPATIBLE
KEYMETRICS MONITORING
API
To be fair, it was a straight up conversion of an old VM in vagrant and Docker was looked at as a one to one replacement before learning otherwise.
They even have an explicit, “run in container” mode.
With Linux's CoW semantics, wouldn't the child share pages with the parent?
Probably not the best searchfoo, but confirmed...
https://github.com/search?q=%22CMD+%5B%22npm%22%2C+%22run%22...
- miniconda (~2GB)
- a final RUN chown -R statement (~750GB)
We reduced the image size and relative Spark cluster considerably by playing around with dependencies in order to stick with plain pip and using COPY --chown.
I also recommend [dive](https://github.com/wagoodman/dive) analyse what contributes to each layer.
(Edit: I realize most of those aren’t statically linked, better description might be “things copied straight into container, not installed”)
> Additions and Modifications are represented the same in the changeset tar archive.
[1]: https://github.com/opencontainers/image-spec/blob/02efb9a75e...
A more portable solution is to use `chmod` right after unzipping the binary. The `COPY` command will then preserve the executable permission.
Linus doesn’t break userland. A tarball is a deployment strategy if someone isn’t dicking with /usr/lib under you.
[0]: https://www.kernel.org/doc/html/latest/filesystems/overlayfs...
They believe it was Wilde who said, “If you want to tell people the truth, you’d better make them laugh or they’ll kill you.”