Simple Dockerfile examples are often broken by default
pythonspeed.com
pythonspeed.com
There are two basic approaches to take with dependency management.
The first version is to lock down every dependency as tightly as you can to avoid accidentally breaking something. Which inevitably leads down the road to everything being locked to something archaic that can't be upgraded easily, and is incompatible with everything else. But with no idea what will break, or how to upgrade. I currently work at a company that went down that path and is now suffering for it.
The second version is upgrade early, upgrade often. This will occasionally lead to problems, but they tend to be temporary and easily fixed. And in the long run, your system will age better. Google is an excellent example of a company that does this.
The post assumes that the first version should be your model. But having seen both up close and personal, my sympathies actually lie with the second.
This is not to say that I'm against reproducible builds. I'm not. But if you want to lock down version numbers for a specific release, have an automated tool supply the right ones for you. And make it trivial to upgrade early, and upgrade often.
No, it doesn't. It just assumes that you want explicit control over when you upgrade. You can always change your Dockerfile or your requirements.txt and build again when you've tested your software against a new Python version or a new version of a package. You can do that as often as you like, so this is perfectly consistent with "upgrade early, upgrade often". But not specifying exact versions in those files means they can get upgraded automatically when you haven't tested your software with them, which can break something.
It might work if there's a dedicated team whose mission is upgrade dependencies for everyone in time, but I haven't seen one in action so I'm not sure how well it might work out. (Well, unless you count Google as one such example. But Google does Google things.)
You get pinned versions that get updated when needed
At my last company (a smaller startup) we used to have a Jenkins job which would open a pull request with all of the requirements.txt updated to the latest available pypi version. That worked pretty well, you always had a pull request open were you could review what was available, it would run the test suite, you could check it out and try it, hit merge if everything looked good and roll it back if it caused an issue somewhere. It made it easy to trace where things changed but not as 'cowboy' as accepting all changes without any review or traceability.
I agree. But that doesn't contradict what I was saying. I was not saying that explicit version control always works. I was only saying that it is perfectly compatible with "upgrade early, upgrade often", since the post I was responding to claimed the contrary.
Also, if an organization can't reliably accomplish timely explicit upgrades, I doubt it's going to deal very well with unexpected breakage resulting from an automatic upgrade either.
So, the alternative is that it suddenly stops working, but caused by the update being available instead of by any explicit action on your part. You'll have more time to react to the problem in this scenario than the other?
Beyond that, we have a daily job that runs the integration tests of all applications with the upstream repository, and if all integration tests end up green, the current set of upstream dependencies gets pushed into the private repository.
It is work to get good enough integration tests working, and at times it can be annoying if a flaky new test in the integration test suite breaks fetching new versions. But on the other hand, it's a pretty safe way to go fast. Usually, this will pull in daily updates and they get distributed over time.
And yes, sometimes it is necessary to set a maximum version constraint due to breaking changes in upstream dependencies. Our workflow requires the creation of a priority ticket when doing that.
This is misleading. My understanding of Google's internal build systems is that they ruthlessly lock down the version of every single dependency, up to and including the compiler binary itself. They then provide tooling on top of that to make it easier to upgrade those locked down versions regularly.
The core problem is that when your codebase gets to the kind of scale that Google's has, if you can't reproduce the entire universe of your dependencies, there is no way any historical commit of anything will ever build. That makes it difficult to do basic things like maintain release branches or bisect bugs.
> if you want to lock down version numbers for a specific release, have an automated tool supply the right ones for you. And make it trivial to upgrade early, and upgrade often.
This part sounds like a more accurate description of what Google and others do, yes.
Any tiny open source project benefits from a reproducible build (when you come back to it months later) and also new versions (with fixed vulnerabilities, and compatibility with the new thing you're trying to do).
A certain amount of reproducibility - a container, pinned dependencies - gives such large reward for how easy it is to achieve that it absolutely is worth it for a tiny open source project.
Worrying about the possibility of unavailable package registries and revoked signing keys, on the other hand, probably isn't.
It's a trade-off. But you certainly don't need to be Google-scale for some of it to be very worth your while.
For sure there are tradeoffs for big projects that don't make sense for small ones. But there are also times where big projects need a tool that's "just better" than what small projects need, and once that tool has been built it can make sense for everyone to use it. I think good, strong, convenient version pinning is an example of the latter, when the tools are available. That was the inspiration for the peru tool (https://github.com/buildinspace/peru).
We use it to do exactly that: pin down every dependency to an exact version, but automatically build and test with newly released versions of each one. (And then merge the upgrade, after fixing any issue.)
Logically, the next step is supporting such infra for containers. Automate all the mundane regression/security/functionality testing while driving dependency upgrades forward.
Actually, it goes further, `bundle update` doesn't update just to "latest version", but to latest version allowed by your direct or transitive version restrictions.
I believe `yarn` ends up working similar in JS?
To me, this is definitely the best practice pattern for dependency management. You definitely need to ruthlessly lock down the exact versions used, in a file that's checked into the repo -- so all builds will use the exact same versions, whether deployment builds or CI builds or whatever. But you also need tooling that lets you easily update the versions, and change the file recording the exact versions that's in the repo.
I'm not sure how/if you can do that reliably and easily with the sorts of dependencies discussed in the OP or in Dockerfiles in general... but it seems clear to me it's the goal.
/s
You can't do that with external FOSS libraries. The closest thing we have is deprecation log messages and blog posts with migration guides.
Rust has crater, which can at least build/test/notify over a large chunk of the rust FOSS ecosystem. It won't pick up every project, granted, and I haven't heard of anyone really using it outside of compiler/stdlib development itself, but it's an example of something a bit closer to what google has.
(Yeah they use different version control terminology since their monorepo doesn't use git, but I've translated.)
The package-lock.json or yarn.lock or similar specifies the pinned dependencies.
But neither the package-lock.json or the yarn.lock file is part of what you get when you create an angular project using the angular cli, meaning that the versions aren't pinned from googles side.
B. If you really want to know what Angular is being tested with, see https://github.com/angular/angular/blob/master/yarn.lock
Downloading and installing system packages lists, etc.
For this reason, Google doesn't use Docker at all.
It writes the OCI images more or less directly. https://github.com/bazelbuild/rules_docker
Your second point is absolutely correct - we strip timestamps from everything which tends to confuse folks :)
yeah I missed that one. but basically it was still a pain. But yeah google's tools are actually awesome. I mean I even own a "fork" (or basically a plugin) for sbt which brings jib to sbt: https://github.com/schmitch/sbt-jib It's just so much easier to build java/scala images with jib than it is with plain docker.
Which part of Google would that be? My impression is the complete opposite, dependencies are not only locked down and sometimes even maintained internally.
I still think the overall advice is good. We depend on node in our Dockerfile like this:
FROM node:11
If we went further into the version, of course we'd be even better off probably, but there's a tiny point to make here. We don't build any docker images for deployments from dev to production. In fact the last time a docker build is run is for the development environment. After that it's just carrying the image from dev to qa to stg to prod, and we simply change the configuration file along the way.
This makes it so that we're not re-building again and possibly getting a different set of binaries that were not tested in any of those other environments.
Node follows semver and rarely has breaking changes within major versions, so this makes sense to do. The article recommends pinning a minor version of Python because it doesn't follow semver and sometimes has breaking changes within minor versions.
If you use a system like nix or guix, this concern is largely obviated.
Test are ESSENTIAL. You should be able to bump all your versions, run your tests and fix the errors. If something gets through broken, then you know where to add a test (before you fix it).
You should pin versions for your sanity. You should also have a process (a weekly process) to deal with updates to dependencies. Dependency Rot will catch up with you!
This is the reason I love archlinux. Most of the time, updates are no big deal. Sometimes, they break the system. Rolling release distros force you to deal with each change as it happens, usually with a warning that breakage is about to happen, and a guide for how to quickly deal with it. Once the system is up and running, basic periodic maintenance will keep it that way. In the past, I've used arch machines continuously for 5+ years and they work great and stay up to date.
Compare to intermittent release distros like Ubuntu. Every time I need to update an ubuntu machine, I end up reinstalling from scratch and configuring from the ground up. There are too many things that need tweaking or simply break between when releases are 6-24 months apart. And I'm not convinced that locking down dependencies actually solve anything. Wait six months after an LTS release, when you need to get the latest version of some package. Suddenly, you are rummaging through random blog posts and repos trying to find the updated package. PPAs, Flatpacks, Snaps, oh my! Intermittent distros offload a lot of their responsibility onto users by pretending like package update problems don't exist.
Found out a few days later that the official release of the JVM broke the Microsoft SQL Server drivers, and Oracle had to ship a new version out asap. Meanwhile, we lost days of work.
Of course, that was also the bad old days of bad old configuration management. But I'd never do something like put an arbitrary version of a language driver in a Dockerfile, not for production.
edit: Of course, the main reason we get scared to upgrade is because we often can't easily back out the change. Docker fixes a lot of that.
Docker by itself isn't enough. But Docker in concert with Kubernetes (or Openshift, in my world) is very, very powerful.
Software can be hard to downgrade. Sometimes dependencies change. Sometimes data models are migrated one way only. Nobody takes the time to properly test them. Among other things.
How Docker, or any other container packaging format for that matter, could possibly help with that I do not understand. It is not the first time I've heard something like this, but I have never been in a situation where the application packaging was part of this particular problem.
Surely starting the an old version of some software isn't neither harder nor easier with Docker than any other way.
Sometimes people don't have good deployment processes that automatically back up whatever they deployed. They might even have installed stuff manually, so they don't know how they did it last time. In that case Docker helps. The build might not be reproducible, but at least you have the binary.
Something that's even more difficult is dealing with upstream changes. What do you do when `ubuntu:18:04` updates? It's easiest if the upstream is released with a predictable cadence (ex: every Wednesday), but none are AFAIK. That way you could plan a routine where you regularly promote an update through QC.
I'm not sure what to think about event driven release engineering like auto-builds (repo links) on Docker Hub. I think that might be an ok solution for development builds or rebuilds of base containers, but it seems to be abused. I bet there are maintainers of popular images on Docker Hub that are effectively triggering new deployments for downstream projects every time they publish a new image.
You can then run the second continuously and warn when it fails and handle whatever happened manually, without having a broken prod.
annoying and wouldn't have happened if we were running pinned versions, that said getting stuck on old software would be worse. however nothing can ever test something like that fully, just too many combinations :(
Vendor your dependencies if you have to, or maintain a cache, but don’t make your Dockerfile redo all of that work.
The difference here is the amount of labor. In your first example, you propose no labor investment.. and in the 2nd, regular investment of labor. Of course that results in the 2nd system being better. If you invested the same labor required by your 2nd example in the 1st, you would likely have an equally usable and up to date system. Similarly, if you didn't invest any labor in the 2nd, you would have a broken unusable system (since the bugs would never get fixed).
Another way to think about this is that as systems age (bugs are found, exploits, etc), they create technical debt. You need to invest the time to address that debt, or you will suffer later down the road.. same as most tech debt.
Chromium is still using Python 2 in its build system.
Only if your dependencies do the same, otherwise you have a new version of something and your dependency wants something old. Especially hard when older dependencies are no longer maintained.
For different environments and languages the problem might be bigger or smaller depending of the culture, strong/weak type coupling, cross compilation or runtime VMs.
This is the reason we've removed all Scala dependencies and now only depend on Java dependencies.
What we need is:
1. Tools that will automatically upgrade for us and report back when there are failures.
2. Better automated tools on where our code is missing test coverage (automated upgrades should break our code when something changes)
3. Better CI/CD rejection support (Something breaks the build, force a revert of the code)
And yes, Nix fixes some of the problems of building a production-ready image, but only a subset.
For example:
1. Signal handling (only one bit of https://hynek.me/articles/docker-signals/ is Dockerfile specific, the rest still applies.)
2. Configuring servers to run correctly in Docker environments (e.g. Gunicorn is broken by default, and some of these issues go beyond Gunicorn: https://pythonspeed.com/articles/gunicorn-in-docker/).
3. Not running as root, and dropping capabilities.
4. Building pinned dependencies for Python that you can feed to Nix.
5. Having processes (human and automated) in place to ensure security updates happen.
6. Knowing how to write shell scripts that aren't completely broken (either by not writing them at all and using better language, or by using bash strict mode: http://redsymbol.net/articles/unofficial-bash-strict-mode/)
etc.
Can you expand on what's missing? I've successfully used nix to cross-compile a pretty substantial python application (+ native extensions, hence the cross compilation), for embedded purposes, and it pretty much worked out of the box. Adding extra dependencies was straightforwards.
I think you can use pypi2nix for pinned dependencies, and you can run it periodically for security updates.
If it works, though, that's great!
The more general point though is that in my experience no tool is perfect, or completely done, or without problems. E.g. the cited https://grahamc.com/blog/nix-and-layered-docker-images suggests you need to spend some time manually thinking about how to create layers for caching? Again, very preliminary research—I know people are using it, I'm just skeptical it's a magic bullet because nothing tends to be a magic bullet.
I agree. I had to do a lot to cajole nix to cross-compile some python extensions.
However, I've done this before manually and using various build systems, and the advantage of Nix is that (1) equivalent builds are cached (reducing compile time), (2) the dependency graph is assured to be clean, (3) the entire state is pure (I can send my nix expressions to a hydra and be guaranteed a successful build), and (4) reuse -- once I modified the higher-level python combinators to build cross extensions, I can add new modules easily.
[0]: https://github.com/Etimo/photo-garden/blob/f597b95c0c488abad...
On one project, we made some example code available, and people would copy and paste it into their project, change a few lines, and launch it into production. Then they were surprised it didn't handle this or that situation or deal with this or that detail.
Yeah, no shit it doesn't handle those things, it's example code! You're supposed to read this code along with the documentation so you can get a gist of what the API is like. It's a learning aid, not a software deliverable. Your real code is going to be more complicated. This simplified code exists to get you past the "how the hell does all this fit together at a high level?" hurdle faster. Once you're over that hurdle, you can start on the real implementation.
I hated Java's checked exceptions, but it did let you have some idea what kind of errors were expected to be possible. Most of the time I'm excavating deep layers of source code, monitoring and logging to see what the errors might be, and tossing in catch-all handlers to keep the stack from unwinding too far if at all in doubt.
Unfortunately, a lot of people hate writing documentation and just want to get it over with as quickly as possible. (Ability to create working software doesn't necessarily correlate with ability or desire to share knowledge.)
Also, if you look at it in terms of incentives, it's tricky to reward people for making sure the documentation of gotchas is reasonably complete. By definition, these are things that people would probably not see until they're pointed out. So it's hard to tell if someone has really done it. (And when people are busy, things that aren't rewarded are unlikely to happen.)
https://gist.github.com/tuco86/67d84dfb27268b1faf05d2dbb1acb...
Ok, I kind of cheated and added the user just now. Sue me. Also posted this in the other Docker related news. Sue me again.
If you are building the container in each stage of your CI/CD pipeline, you are doing it wrong.
While I may not agree with absolutely everything in the article, this final point is paramount. Please don't blindly use technology because you managed to find a copypasta config that runs. Running != good.
using namespace std;
is just staggering. Sure, it works in a toy example posted to stackoverflow, but it will cause problems in larger projects. I think globally there needs to be better emphasis on using best-practices in tutorials and examples; I remember this particular pet-peeve of mine also being present in college textbooks. Especially for content aimed at newbies, it should be frowned upon to show the wrong way to do things, since then it gets harder to show how to do it the right way.
I've had people who were surprised to find out that they could type: using std::chrono::duration;
using std::cout;
instead of pulling in the entire std namespace; simply because they'd only ever seen examples that did it the lazy way.edit: lack of semicolons strikes again!
But in general the notion that we want to isolate "everything" into a namespace is a net loss. Clear and simple abstractions have real value, and short undeclared names are an important part of being clear and simple.
The modern convention of separately importing every symbol you use gets really out of hand, when most of the time it really is appropriate that you just declare "my code is using this API" and expect things to work without having to link your program by hand with a giant shipping manifest of symbols at the top of your source files.
I've had to look at some C# web service code recently, and the amount of magic it relied on made it impossible for me to find what I was looking for, even using grep.
from tk import *
essentially equivalent to my C++ example? They're bad practices in both languages. Thankfully python examples seem to be better in this regard, since I don't recall seeing wildcard imports in any of the tutorials or references I've used.On the other hand, import * is nice to have in a repl.
Seriously, when was the last time a C programmer needed to figure out where identifiers like "strlen" or "fread" come from? The problem you posit need not exist for big chunks of commonly used APIs, and the inability of modern programmers (trained, it seems, on C# web service code) to see that is frustrating.
using namespace std;
for std specifically, giving the reasoning that:> sometimes a namespace is so fundamental and prevalent in a code base, that consistent qualification would be verbose and distracting.
I also work mainly in C++, and personally I prefer using it, together with -Wshadow to catch possible issues.
0: https://github.com/isocpp/CppCoreGuidelines/blob/master/CppC...
I could probably safely make it "python:3.7-alpine3.9" (instead of pinning to Python 3.7.3), since the issue was the Alpine version, but at this point I'm starting to really buy into the whole reproducible build thing.
I've seen docker images that do a git clone from the master head to get the source, so basically if their Github account gets hacked. You're f'd.
I wonder how many use docker after they learn ?
I’m not saying “go run your all your Docker images as root”, but this is clearly FUD.
Non-privileged containers are still having "root", just with way fewer capabilities (See Docker [0] docs).
I’m not an expert, but I guess depending what you are doing the most problematic capability might be AUDIT_WRITE, because it is not namespaced and could be abused for DOSing syslog. But you might require it for things like sshd, sudo, adduser, passwd, …
Depending on how you are holding it the NET_BIND_SERVICE and NET_RAW can be an issue (depends on how your docker network looks like), but the others appear not to be a security issue per-se.
This page [1] gives a good overview on default capabilities, though they are also confusing to the reader with "better disable this".
I've created an issue, not sure if I have resources to fix their page though. [2]
[0]: https://docs.docker.com/engine/reference/run/#runtime-privil...
[1]: https://www.redhat.com/en/blog/secure-your-containers-one-we...
[2]: https://github.com/mhausenblas/canihaznonprivilegedcontainer...
Real world example: CVE from February 2019 which allowed escalation to root on host. It's preventable by (among other things) "a low privileged user inside the container".
See https://blog.dragonsector.pl/2019/02/cve-2019-5736-escape-fr...
Thank you for this link, I've only seen the initial CVE announcement.
> This is not FUD [...]
The site is practicing FUD, it accomplishes communicating a message in an untruthful fashion by mixing two different things into one. (Just check out their stack overflow links, it is not clear if they are talking about root or `--privileged`)
People are confused wether `docker root == host system root` and this site doesn't help them to get a better understanding whether or not it is the case. (It isn't) Plus it misses what its main goal should be, running a Secure Docker environment.
You are talking about a previous exploit, not a permanent issue. Keeping your host system up to date and additional hardenings is always going to be necessary in exposed environments.
> Use Docker containers with SELinux enabled (--selinux-enabled). This prevents processes inside the container from overwriting the host docker-runc binary.
Authors recommendation is also using SELinux, this also helped with outer Docker/Kernel related vulnerabilities in the past. Why isn't the page even mentioning this?
---
I think it is important to give a proper outlook on how problematic things are and not to confuse people with super high expectations. You often end up running containers that you have only little control about.
1. Avoiding root in self-built containers is definitively the way to go, since it reduces (unnecessary) attack surface, but
1. It requires some glue code
2. Might slow down your builds (`Dockerfile` multistage `cp --from=0 /app /app` loses permissions, requires chown afterwards)
2. Avoiding root in CI/CD is nearly impossible
1. many package managers won't work 2. some capabilities to test things (sshd for testing ansible scripts for example)
3. can you use kaniko for building Docker images from within Docker without root?
3. Harden your Docker host 1. Use SELinux
2. Use monitoring
3. Drop capabilities that aren't necessary (NET_BIND_SERVICE, NET_RAW, ...)
4. Use docker network separation
5. Frequent system updates
4. Keep yourself up-to-date, especially if you are running an exposed environmentJust googled and this is rather more helpful:
- https://dev.to/petermbenjamin/docker-security-best-practices...
- https://blog.aquasec.com/docker-security-best-practices
- https://sysdig.com/blog/7-docker-security-vulnerabilities/
You can't assume all containers have same access and trust-level as a user because at some point you also have to process your requests and store your data - and that processing is most likely also be done in another container (even if your actual data-volume is mounted outside the container). For such containers same security measures as for a physicial machine applies.
But the reproducible build aspect of the critic seems unnecessary to me: Isn't that more a concern of the packaging system? (no python scripter)
If your packaging systems supports version selection/locking, then use your packaging system right. If your packaging system cannot pin a version, how should docker solve this?
Or they don't talk about need to run as not-root.
Or they suggest base images that are often broken in subtle ways (Alpine Linux).
Or they talk about multi-stage builds for small images, and neglect to explain that you've just destroyed your caching (this is fixable, but you need to know to expect it and how to fix it.)
Etc.
Rarely is this done though. And definitely alpine+musl doesn't always do what you might expect, and it's often language dependent as to whether or not you'll encounter something strange (not to mention you forfeiting bash)
Did you understand the point i tried to make nonetheless or do i need to detail it?
That's exactly what the article is recommending in point 2. The original Dockerfile author was using pip in a way that's only intended for in development. Having a requirements.txt file is the correct way to use pip when distributing a project.
You could of course blame users for not making sure that all the commands they use in their Dockerfiles are actually reproducible but many/most examples even in the official documention are clearly not reproducible.
Therefore you end up with what is in my opinion a semi-broken system - building images seems to be reproducible (and fast) until you lose your layer cache or you spin up a new CI build agent or a new dev joins the team and tries to build the same image.
Not that I can think of an clean and performant solution to this problem.
The use of layers at the build stage adds a lot of needless complexity with very little benefits and users really need to step back and question the value they are getting from the use of layers. [2]
Words like 'immutability', 'declarative' and 'reproduciblity' are often used in ways that can lead to user misunderstanding and can be accomplished with simpler workflows. For instance immutability, reuse, composition do not require layers. There needs to be a lot more technical scrutiny to avoid confusion.
[1] https://www.flockport.com/docs/containers#builds
[2] https://www.flockport.com/guides/say-yes-to-containers.html
Best part is that was brought up as an issue in the article only to do the same thing in an example
This seems to be the happy medium for me. I don't have very strong opinions on requirements.txt always being the pinned output from a pip freeze, and it seems like pipenv may actually die in a few years, and poetry will evolve to take the mantle, but I do lots of things with conda anyway.
The idea is that you should take responsibility for your containers and verify fixes and test your application.
> Do you want it to run stable with vulnerabilities or to run secure and broken?
If these are your two choices, you have a staffing or a workflow problem.
CMD [ "python", "./yourscript.py" ]
This breaks if you want to debug yourscript.py on startup. Better use a sh to wrap it.E.g., for me, the following Dockerfile:
FROM python:3
RUN pip3 install ipdb
COPY test.py /test.py
CMD ["python", "test.py"]
where test.py is: import ipdb; ipdb.set_trace()
print('Hello, World.')
run as `docker run -ti --rm $IMAGE_ID` works as expected: » docker run -ti --rm 52e98c118dc3
> /test.py(2)<module>()
1 import ipdb; ipdb.set_trace()
----> 2 print('Hello, World.')
ipdb> p globals()
{'__name__': '__main__', '__doc__': None, '__package__': None, '__loader__': <_frozen_importlib_external.SourceFileLoader object at 0x7f809a8c8278>, '__spec__': None, '__annotations__': {}, '__builtins__': <module 'builtins' (built-in)>, '__file__': 'test.py', '__cached__': None, 'ipdb': <module 'ipdb' from '/usr/local/lib/python3.7/site-packages/ipdb/__init__.py'>}
ipdb> ^D
Exiting Debugger.
»`docker exec -e foo=bar -it ....`