Minify your container
github.com
github.com
A containerizable project probably has its requirements known and well-specified? I think building on top of a base with a smaller unused surface is a better idea than using analysis that might backfire. These days I am using apko + melange for my personal images and they are super neat.
Removing those code paths would not be a good thing, but I guess if you build your apps right you could just have your container orchestration system recover by replacing the Pod.
And while I do have automated tests, they might sometimes stub system calls as I'm mostly testing my code to keep things stable and fast.
I'd rather explicitly declare my dependencies and use the same container for development, test and production to feel much more confident that it includes actually everything that's needed.
But I agree, this is just bandaid for lazy bois. Better use Bazel etc. for distroless builds
If you forget a critical code path when you build using Docker-Slim, and a resource file is not used, that resource will be stripped. The feature which depends on it will be broken in production.
If Docker-Slim is working for you in production apps, you are either getting lucky or your app is trivial enough to lack unseen code paths.
Like, maybe some log forwarder utility that only gets called for "CRIT" messages that didn't happen to get triggered by testing.
The best thing you can do is use minimal images and multi-stage builds. This should help you immensely to reduce your attack vector and do standard software bill of materials, too.
I agree that the multi-stage builds are the best option, but it can be hard to know if you've included everything that is required or if you've accidentally excluded something that is important in rare cases.
For example, even the best integration tests (for small/mid-size companies) don't always include tests that exercise weird paths around dates/times - leap years, leap seconds, daylight savings time, etc. We often trust that our datetime library or code will handle these for us, but what if the configuration is stored in a file that isn't accessed during your integration tests?
Best case scenario is you hit the error-path soon in production and your code either crashes or does something correct-enough with a fallback path, but a worse scenario is you start losing critical information and don't realize it/fix it until it's gone on for a while.
Its actually a good way to find dependencies you didn't know were there. As long as you're diligent this isn't something you should be shaking in your boots over.
an example Dockerfile likes:
FROM golang:1.18.1-alpine as builder
# RUN apk add, wget, etc, and build the binary
FROM alpine
# or FROM scratch
COPY --from=builder builder/binary /binary
ENTRYPOINT ["/binary"]Dev tools like bash, ls, grep, etc, have no place in production and only increase attack surface.
I usually just include a shell, but to each their own.
https://kubernetes.io/docs/tasks/debug/debug-application/deb...
[0] https://stackoverflow.com/questions/47722898/how-to-do-a-doc...
Dive is a good tool for the latter IME. https://github.com/wagoodman/dive
It doesn't do the work for you, but it does single out the big layers in your image.
https://twitter.com/ariadneconill/status/1506482425458798593
https://twitter.com/ariadneconill/status/1506483943352250371
Sorry – I forgot that the Twitter UI doesn't always lend itself to proper threading.
If you're running a website and the removed dependency is related to a feature that is uncommon enough that isn't covered by your automated tests, maybe .1% of your users experience a broken page.
If you're running critical infrastructure and the removed dependency has to do with leap-second handling, maybe eight months from now, everything crashes and you lose millions of dollars.
And for languages like golang (in their examples) - why/how would anyone get such huge container images in the first place? Doesn't go give a neat statically linked binary?
And that's also why you have multistage docker builds. To make sure your production container doesn't have all the unneeded files from your development container. https://docs.docker.com/develop/develop-images/multistage-bu... .
But if you go around "minifying" all your applications independently, you won't have that shared base layer. One application needs `sh` and another doesn't? Now you get two entire base layers, one with it and one without. Sure, each image's total size will be less, but the size of all your different images added up will be greater because you killed the sharing.
If for some reason the 29 megs of ubuntu minimal (or even fewer for alpine) are a problem (which they aren't on your server that already has over a hundred megs of `docker` binaries), then the right solution is to better control layer sharing. Ensure that you don't have different base layers between your applications. And then--strictly for kicks and giggles--you could minify that base layer to the minimal set of what all your images require. To save a 51K `passwd` binary (woohoo!).
hint yes it is and that could be a problem a giuant one
With something like alpine linux/ubuntu minimal, you trust the package maintainers to make sure that if you use python in your docker image it would work like it worked for them. Out here, it just says "Yes (it is safe)! Either way, you should test your Docker images.".
As a bad example, if a library used by your application uses a different "theme" requiring different files at night and different files during the day, you might still say "it worked during my tests" but things definitely broke and the only thing you can blame is this overzealous tool.
That bad example was from back when i was trying to make AppImages for an application we used. At first all we did was recursively collect all the libraries reported by ldd. Then it turned out some libraries were only being dlopen'ed by other libraries under specific circumstances and we missed them. So we manually added those libraries. Then it turned out that we missed the config files and other resources used by those libraries. Eventually we shipped all the files belonging to all the distro packages used by the libraries we used and left it at that.
in some cases i essentialy ensure my whole app remains using --include-path flags so that i get a removal of you know things that i absolutly dont need.
I’ve been using them (distroless) with great success for my Rust applications.
is about 4 MB.
It contains things like:
/usr/share/zoneinfo/<timezones>
/etc/ssl/certs/ca-certificates.crt
/etc/debian_version
... not much else really.On the other hand, scratch is a special image that contains literally nothing, 0 bytes. Docker doesn't have to talk to the network to download it, it doesn't have to be built, it's just a tarball of nothing.
scratch is thus infinitely smaller than distroless since it has no size.
It's also not suitable for rust since rust applications will commonly link against openssl for tls, against a libc implementation for things like network operations and threads, and so on.
Because of such tooling / ecosystem, it's more convenient to use scratch from Rust than from most languages (including C)
static has: ca-certificates, /etc/passwd, /tmp directory, tzdata
base has: glibc, libssl, openssl
https://github.com/GoogleContainerTools/distroless/tree/main...
And yeah libc was a pain for us even for AppImages. You'd think that something as fundamental as C library would be standardized on Unixes...
distroless is picking up a lot of interest especially with its recent uptick in its adoption in the kubernetes community.
In practice, I'm curious how error prone the result is.
It's a neat approach, but ultimately brings non-negligible amount of uncertainty as you can never be 100% sure your test set of inputs did not miss a particular edge case which will require to have a file present in the container that no other input does.
personaly i highly recommend as it works in most cases and gets rid of those vulnerablities that come with things like bash or passwd that you dont need in prod apps
Container optimization also seems to border on premature optimization. If you're not hitting issues with your container size, don't worry about it. People may scoff at a 10GB container but for something that doesn't scale or get built a lot, it may be fine.
There are also other options like pre-pulling containers on the machines running them (which you could build into a golden image) that can get you similar startup/pull speed gains for something like host auto scaling use cases
Instead of going extreme with coverage analysis, it shows places that can be manually cleaned during the build process. Maybe someone will find it useful. Smaller space gains, but gives more confidence.
It's nice because they're inverses - if you shrink by 30x, then grow by 30x, you're back where you started, whereas a 97% decrease in size followed by a 97% increase in size leaves you at ~6% of the original size.
If I had starting capital of 100 usd and it increased by 20x then that means I gained 2000 usd which is 20x of my starting capital.
EDIT: correction 30x is same as 3000%.
X is input
Y is output
X / 30 = Y
All you need to know, no minor required, as its taught on elementary school (age 11/12 or so?).
But... doesn't shrinking by 1/2 also mean dividing by 2?
Therefore, 1/2 == 2x ??
I feel like my elementary school math is letting me down somewhere.
We're clear on 'x is two times larger than y', right?
x = 2*y
An equivalent statement is 'y is two times smaller than x', but it conveys a construction more like: y = x/2
Which, since we're speaking English sentences, might change the emphasis/implication.It's kinda like tree shaking in JS. Super standard practice and hugely deployed everywhere. I don't care that zip or tar are removed. I care that my app runs, passes QA and gets deployed to prod.
I don't know what other containers people are running but if you have Lotus Domino running in a container, fine, don't minimize that, I'm super happy to minimize my node process that collects HTML files from other sites.
This is great.
So we’re promoting secrets being saved within a container image artifact? Ummmm?
Seems like all past attempts have stalled and/or are dependent upon FreeBSD creating a standard for what’s in a minimal userspace .
Just curious (please don't take my comments as being negative)
One thing I'd like to mention right away is that docker-slim doesn't do function level code elimination and that simplifies what it needs to do in a significant way.
Would you like to clarify any of the misconceptions?
(Also, even primarily to me, but less relevantly, it's great for gitting configuration & its reasons, and syncing across machines.)
Docker's trademark guidelines say that products, services and technology that are not their own shouldn't use the Docker name so it seems it's a matter of time before they get a nice letter from some lawyer.
Does Docker Inc promise to overlook trademark infringements if you're a "loose" partner?