Is this standard practice? I've never done this but it makes a lot of sense. Is this basically just adding `RUN apt-get update && apt-get -y upgrade` at the top of the Dockerfile? (Assuming a Debian image.)
Is this standard practice? I've never done this but it makes a lot of sense. Is this basically just adding `RUN apt-get update && apt-get -y upgrade` at the top of the Dockerfile? (Assuming a Debian image.)
In practice, it will eventually bite you. Someday some package that you depend on will have a security issue that causes a necessary config change that has a default you don't like.
Also, note that you have just changed to a non-reproducible container, where running it on Monday is not guaranteed to be the same as running it next Friday.
I pull new updates on a set schedule and only push the ones that are tested in nonprod. There are any number of tools that do this, Uyani is a good free one.
Precisely, and that may be what stops many from just updating their images. I think ideally we'd have centralised updates, perhaps just one a week, and then use tag for that week.
You could do all this in-house, but it's a lot of work for small teams. Redundant work.
Maybe the organization wants to deploy the latest trunk on every commit. There are probably some situations where that is reasonable. Those situations probably shouldn't involve people's money, privacy or safety.
Setting up a two-stage local repository isn't very hard. The intake side gets updates from upstream, and the deployment side gets updates from the intake side when the packages have been reviewed and hopefully tested. Do this for everything with an external upstream -- Ruby gems, Python modules, JS libraries, whatever -- and you have insulated yourself against supply-chain attacks. As a bonus, if your leftpad function just goes missing upstream, you still have a full copy in your deployment repository, and will until you decide it's time to implement your own.
Yes and no, I don't disagree, but remember classic hosting, semi-managed infrastructure still exists. We often get container images delivered and are responsible for the operational side of things, but we have no control of what is actually inside the containers.
Sure we can make requests, or inform a customer that we believe what they are doing isn't safe. Ideally we could reject a request to run a container, but in reality that's not really an option. That shifts the responsibility of security more in the direction of the developers and they often do not have regular patch management as a priority.
This is, incidentally, the antithesis of "devops".
What's the alternative?
Additionally, if you don't "apt update", some of the packages you try to install from mirrors will 404.
Unless you're using something with deterministic builds, reproducibility is a myth anyway. Update your OS and test the image and save the artifact somewhere with "docker save | zstd".
You have to do an “apt update” to ensure that the packages you install will be fetchable, because the apt indices inside the image are out of date and “apt-get install -y whatever” may fail with a 404. That means you aren’t guaranteed to get the same version installed from an “apt-get install -y whatever” after Monday’s “apt-get update” as you would after Tuesday’s “apt-get update”, as the current version of “whatever” may have changed in the interim, even if you don’t run an “apt upgrade”.
In any case, a lot of the files on disk are generated dynamically at install time for certain types of things, and include things derived from the state of the system (which frequently depends on remote network resources, as described above), so issuing the exact same dockerfile FROM+RUN+RUN+ADD etc lines will not result in an identical image result when run on different days: it’s nondeterministic. A deterministic build is one where the same build always results in a byte-for-byte identical build artifact.
There is effort being put in to make the building of the backing .debs deterministic, but AFAIK no Debian or Debian-like is trying to make apt itself work in a deterministic manner when installing packages. There are still postinstall scripts, for example, that are system-state dependent.
Really, you need to be saving your build artifacts when you do Docker builds. Saving the Dockerfile and expecting to be able to rebuild the image at any time later isn’t a good bet. You might, sometimes, be able to rebuild a mostly-compatible image, but there is no chance whatsoever you will be able to build a byte-for-byte compatible image, and it’s entirely possible that your build might just fail entirely (e.g. if you are pinning specific package versions that fall out of date and are no longer fetchable from the mirrors).
Then, in the worst case scenario, you can always load your original working/saved image back in, replace/patch/modify specific files on disk to address issues (either with vendor tools or manually) via a new build that pulls FROM the saved artifact image, and make a derived one.
You can use Stable or https://snapshot.debian.org/
I ways always confused about this, thinking I didn't understand what people were doing here... but it turns out maybe there is no good solution and most people are ignoring it?
Would you expect this to result in more attention at some point, after it results in more exploits?
I don't know. I have been able to tolerate "apt-get dist-upgrade" on my legacy non-container system: upgrading from one Ubuntu LTS to the next is hard, but within an LTS we've never encountered a problem over 15 years (I think? more than a decade at least) this has been done.
A newer project uses NixOS, but it is so new it has no users. NixOS allows me to mention the particular commit of nixpkgs - that means the versions of the packages in the package manager - into a lockfile. This means I can run daily updates, run my tests against them, and deploy them. And because the lockfile is checked in, if things suddenly start going wrong on the umpteenth of Octember, I can see what changes happened on that day: was it the file I commited? no? oh I see, libtwiddle was upgraded. Yes, it was libtwiddle that broke things.
As I say, this is a brand new project. It may not work as perfectly as all that. And swallowing Nix requires a certain amount of koolaid: it is user friendly but it's very picky about who its friends are.
> Would you expect this to result in more attention at some point, after it results in more exploits?
I think that we will switch to paid platforms that offer managed runtimes. These will probably offer targeted, well communicated updates of dependencies (everyone will hear Microsoft Python is releasing ms-py-http-3: here's what you need to know), and have fewer libraries that offer more code. In a sense these will be switch back to distributions. We will say "why would you manage your own dependency?" And the cycle will repeat with a new flavor.
Conventional would be to ignore it.
This gives you the ability to move the pieces independently from one another so when you release it's actually (sw_version, platform_version) and lets you track down bugs caused by platform updates easier.
That way if something in the base changes and causes a regression you can pin to the last known good until you can fix the bug.