Docker Compose best practices for dev and prod
prod.releasehub.com
prod.releasehub.com
https://docs.docker.com/compose/profiles/
Prior to that feature, I had a hand-rolled approach using bash scripts and using overrides. Worked well enough, but felt like a dirty dependency.
Ever since the V2 migration, I've been banging my head against multiple bugs in Docker Compose itself, which has severely damaged my trust in it as a production-ready utility. I no longer feel like I can safely update Compose without running extensive testing on each release.
For example, my journey since migrating to V2 has been:
* Originally on 1.29.2
- Stable and I did not encounter any issues.
* Swapped to 2.3.3. - As part of the migration from V1, docker-compose lost track of several networks and started labeling them as external. Not a huge deal.
- docker-compose stopped updating containers when the underlying image or configuration changed (#9291). This broke prod for a while until I noticed nothing was getting updated.
* Updated to 2.7.0, which was the latest release at the time. - This introduced a subtle bug with env_file parsing (#9636), which again broke prod in subtle ways that took a while to figure out.
* Updated to 2.8.0, which was the latest release at the time and before the warning was put up. * This broke a lot of things as it renamed containers and generated duplicate services, etc.
* Finally swapped to 2.9.0. No issues yet, but I'm keeping an eye out.1. A `docker-compose.yaml` for basic configuration common among all environments
2. A `docker-compose.dev.yaml` for configuration overrides specific to development
3. A `docker-compose.prod.yaml` for configuration overrides specific to production
Then I have `COMPOSE_FILE=docker-compose.yaml:docker-compose.dev.yaml` in my `.env` file in the root of the project so I can just run `docker compose up` in dev, and `COMPOSE_FILE=docker-compose.yaml:docker-compose.prod.yaml` in a `prod.env` file so I can run `docker compose --env-file=prod.env up` in production. Sometimes also a `staging.env` file which is identical to `prod.env` but with a few different environment variables.
It's been working pretty well for me so far.
In your example, "dev" should probably be the default configuration, which "prod" overrides. It's a tiny detail but I think it pays over time.
A similar situation is when you have lots of almost-identical settings between many (micro)services.
Having prod be the default is a problem, as it’s inevitable someone will accidentally break production/write junk to it when they didn’t realize they were talking to prod.
Having dev in it inevitably results in broken prod pushes because no one has ever tried it outside of a synthetic dev environment.
Similarly, in dev, you can set the email config to output to a file and be done with it. I don't want to see all the email server settings in my dev config: the host, the port, the username, the password, the protocol, etc.
Why would you want two different instances of the same project to have the same name?
Why on earth would you want to have lots of different directories that all point to the same instance? That's simply idiotic.
volumes:
- ./data:/use/lib/data
I don't know why anyone would expect this to work.It means that if I run multiple instances of a compose setup, they are completely isolated from each other because they live in separate namespaces. It's a very handy feature (for example, if I'm working on 3 branches of the code, just check them out in 3 different directories etc).
Of course it has the consequence that if you move it around it's going to break things, but that's actually the idea. They should break! Because to compose they are meant to be segregated, and if they interacted it would actually be a bug.
Similarly yaml anchors are cool and can easily reduce duplication in your files. However I've found the majority of Devs and users find them confusing and unexpected in a compose file. As a result I would recommend not using them and just sticking with plain yaml for simplicity's sake and not being afraid of duplication.
One other alternative to yaml anchors is instead using the env_file: property on your services (https://docs.docker.com/compose/environment-variables/#the-e...) . This way you can place all of your env variables in a single file and share this same of env variables between many services without having to have large duplicate env blocks in your compose file.
docker compose --env-file ${INSTANCE}.env -f docker-compose.yml up
In the `*.env` file: CONF_A_VALUE="some_value_a"
CONF_B_VALUE="some_value_b"
CONF_C_VALUE="some_value_c"
...
In the compose file: services:
service-1:
# you need to be more explicit here
environment:
SOME_NAME_FOR_A: "${CONF_A_VALUE:?err}"
SOME_NAME_FOR_B: "${CONF_B_VALUE:?err}"
# but it gives you simple templating in other places
labels:
com.company.name.label.a: ${CONF_D_VALUE:?err}
com.company.name.label.b: ${CONF_E_VALUE:?err}
com.company.name.label.c: ${CONF_F_VALUE:?err}
ports:
- "${CONF_G_VALUE:?err}:8020"
service-2:
# also gives services more freedom to diverge from one another (if required)
environment:
MAYBE_ANOTHER_NAME_FOR_A_BECAUSE_COWBOY_TEAM_REASONS: "${CONF_A_VALUE:?err}"
SOME_NAME_FOR_C: "${CONF_C_VALUE:?err}"
Where `?err` forces the variable to be set and non-empty: https://docs.docker.com/compose/environment-variables/#subst...It's just a shame that `docker stack` doesn't have an `--env-file` argument yet for swarm deploys.
unfortunately the ecosystem is on kubernetes. you want to integrate with spot instance bidding on AWS...u need kubernetes. you want monitoring tools...u need kubernetes, etc
the Compose spec is now open and standardised - https://www.compose-spec.io/
a kubernetes distro that can be managed entirely using compose files will be a massively impactful project. Not like Kompose which converts Compose files to k8s yml files....but entirely on Compose files.
I really hope someone does a startup here.
e.g. datadog, etc https://www.datadoghq.com/blog/monitoring-kubernetes-with-da...
that said, Grafana's own hosted solution "Grafana Cloud" gives Kubernetes operators...but not for anything else. https://grafana.com/docs/grafana-cloud/kubernetes-monitoring...
the entire ecosystem is supporting k8s. there's not much choice here.
Do you mean from the your perspective, you set up your docker-compose and this startup ends up deploying pods into Kubernetes based on your compose file? Basically cutting out the middleman of you having to deal with Kubernetes yaml files?
Compose would be awesome.
We basically do this at ReleaseHub, however there is a level of indirection we introduced. We basically have our own version of Kompose which generates our own YAML file, which we call an Application Template[1]. The Application Template then gets parsed into Kubernetes YAML and deployed, however as an end user, you never have you deal with the Kubernetes YAML.
The main piece missing is that we set everything up to parse the docker-compose file once and then expect people to interact with our YAML. However, having seen the level of interest about compose in this thread I'm wondering if maybe there is a feature to be built where we remove the need for interacting with our YAML, an end user can push changes to their compose and through GitOps we update everything and deploy new changes to Kubernetes.
The compose spec has come a long way since we started (we were attending the spec meetings back in 2020 to see how the project was going to kick off) and is in an even better place to support this direct Compose -> Kubernetes idea.
If you're curious, feel free to reach out jeremy@releasehub.com.
[1] - https://docs.releasehub.com/reference-documentation/applicat...
At this point, you are sitting in a landscape which has already argued around Borg vs Kubernetes and consequently YAML vs Jsonnet vs Cue (e.g. https://github.com/cue-lang/cue/discussions/669). We have original Borg and k8s architects who have spilled lot of ink on this.
It is going to be super hard to digest another custom markup. This is really your battleground. People will ask - why not the standrdised Compose specification or standardised jsonnet/cue. In this context, Compose is kind of loved by all (though not the technically superior choice here)
the good news is - ur asking the same question probably and thats a good direction. You also have the steps to build up on (k3s, k11s, k9s - yes they are all different things and very cool). Best wishes if you go down this path.
Docker Compose is also unnecessary for that. I'm talking to the demographic of users who want to run Compose/K8s/Rancher by themselves. For most people, i dont recommend it - you should use Fargate.
Update: reading comprehension, it's a thing! :) Thanks chrsig for pointing out that this is under an appropriate sub-heading. D'oh!
> Docker Compose Best Practices for Development
It's also a lot easier if there's a bad deploy to roll back the update by reverting the image tag in the compose file and restarting rather than checking out specific older commits and risk getting into a funky state with detached heads and the like.
Did you know that compose accepts stdin for `-f` files? Anything that outputs yaml is a valid "compose provider". Self plug, more on that, https://eskerda.com/complex-docker-compose-templates-using-b...
Try having 40 microservices and quickly iterate between fixing bugs in them. Some you want mounted, others you want running unmounted. I understand the premise is already broken, you do not want 40 microservices, but that's not a precondition you can fix without a time machine.
[[* EDIT: by env vars on the global scope, I mean setting things on the global scope by using an env var, say, you want a certain volume named after an environment variable. You can only use ENVs as values.
# this works
foo: "${SOME_ENV}"
# this does not
${SOME_ENV}: "foobar"Edit: Also, I don't see the need for a turing complete language for something like docker compose, if you need something really complex you can always script a docker-compose.yml generator with all the logic and complexity you need.
* docker compose could be better
* docker compose does not improve that much
* by mixing docker compose with templates you can get complex stuff going, but is not a much advertised feature.
Most of the "cool dev tool env" that appear are based on kubernetes, or something that is _not_ docker compose, which makes them difficult to adopt. I would like docker compose to improve on a direction that makes development easier, that's all.
services:
redis:
...
{{ if some_condition }}
volumes:
bla bla
{{ end if }}
{{ if some_other_condition }}
another_service:
...
{{ end if }}
volumes:
{{ for volume in volumes }}
- ...
{{ end for }}This feels like it's a very incorrect solution to a problem that you're making far more complicated than it needs to be.
The following is just a contrived example. Why would I do that? Because I can
# Bring up services without a command
#
# usage:
# nocmd.sh up -d service_name
function _compose {
docker-compose -f docker-compose.yml $@
}
function nocmd {
cat << EOF
services:
$1:
command: tail -f /dev/null
EOF
}
function nocmd_compose {
local service=${@: -1}
_compose -f <(nocmd "$service") $@
}
nocmd_compose $@Kids these days don’t know anything about Tmux and port 433. Exposing localhost to the internet and getting your internet monitored by your ISP and getting letters in the mail; for hosting movies and Storage services.
Lol.
https://en.wikipedia.org/wiki/Network_News_Transfer_Protocol
> The Network News Transfer Protocol (NNTP) is an application protocol used for transporting Usenet news articles (netnews) between news servers, and for reading/posting articles by the end user client applications.
> Well-known TCP port 433 (NNSP) may be used when doing a bulk transfer of articles from one server to another.
The way ISPs catch you is by seeing that you're connecting to IPs that are known to host content illegally. There's now way for your ISP to know what's going _out_ of your port 443 without breaking TLS.
They care a LOT about what you're serving to other people from their network though as there might be some legal issues directed back at them.
https://arstechnica.com/information-technology/2015/03/atts-...
This was a step up for novice sysadmins/developers from "FTP to the webroot" without having to learn real deployment tools, how to write init scripts, etc. Of course it wasn't good enough for anything business critical, but I'm sure it happened there too.
Also, if you then wanted to access that service from your college network, who may have gone out of their way to block gaming, a common workaround was to host it on port 443. By 2010-2015, colleges were using DPI sophisticated enough to see that "that's not http traffic" on port 80, but tended to treat 443 as "that's encrypted, must be a black box". That's how I ran mine in ~2013.
Has anyone else had this issue with Windows? The last time I looked into the issue was about 8 months ago, so it's possible the issues have been addressed since then.
If you find yourself spending a ton of time on a one-size-fits-all-disappoints-everybody solution, maybe its time to build two effective solutions instead.
Containers are not meant to make anything OS-agnostic. Containers are just a way of running Linux
Something hilarious to me is that "multi-arch" images are a thing, and can result in the same dockerfile building a Windows image on a Windows PC, and a Linux image on a Linux PC!
dockerd for windows isn’t even free software, last I looked.
Docker Desktop volume performance on Windows is great if you're using WSL 2 as long as your source code is in the WSL 2 file system. It also works great with WSL 1. I've been using it for a long time for full time dev.
In fact, volume performance on Windows tends to be near native Linux speeds. It's macOS where volume speeds are really slow (even on a new M1 using Virtiofs). Some of our test suites run in 30 seconds on Windows but 3 minutes on macOS.
But the Docker Compose config is the same in both. If you want, I currently maintain example apps for Flask, Rails, Django, Node and Phoenix at https://github.com/nickjj?tab=repositories&q=docker-*-exampl..., I use some of these exact example apps at work and for contract work. The same files are used on Windows, macOS and Linux.
The only time issues arise is when developers on macOS forget that macOS' file system is case insensitive where as Linux is not. Docker volumes take on properties of their host so this sometimes ends up being a "but it works on my machine!" issue specific to macOS. CI running in Linux always catches this tho.
My experience actually was the exact opposite. This one project had horrible IO performance on Windows, when running a bunch of PHP containers, with the WSL2 integration in Docker Desktop set to enabled.
It was bad to the point of the app taking half a minute to just load a CRUD page with some tables, whereas after switching to Hyper-V back end for running Docker, things sped up to where the page load was around 3 seconds.
I have no idea why that was, but after switching to something other than WSL2, things did indeed improve by an order of magnitude.
No `services` block above the `api` and `web` services, on the same level as `x-app:`.
So the `<<: *default-app` anchor won't be recognised as the block that defines it is never terminated prior to use?
(I've not tested as currently on the windows gaming box. this is purely from reading it, so happy to be corrected).
Like with `lando init` and then `lando start`, it'll fetch a LEMP "recipe" (say, for Wordpress or Drupal), configure all the Docker containers for you, set up all the networking between them, get PHP xdebug and such working automatically, and leave you with an IP and port to have your app ready at. It supports a bunch of other common services too, like redis, memcached, varnish, postgres, mongo, tomcat, elasticsearch, solr, and only very rarely do you need to drop down to editing raw docker compose or config files.
Having had to use the Docker ecosystem for a few years, Lando made my life way easier. Although these days I just tend to avoid Docker altogether whenever possible (opting for simpler/more abstracted stacks, often maintained by a vendor like Vercel, Netlify, or Gatsby).
> If you happen to be unlucky enough to need Docker for LAMP/LEMP stacks
We run Docker LAMP and have a new version underway that's LEMP, haven't run into any problems. xdebug works fine, I even got it working recently in ECS. I'm not sure what's different about LAMP/LEMP in Docker vs any other stack.
The "unlucky" part is just having to work with a LEMP stack at all, and that's just my personal bias leaking through my post, sorry. Having grown up with that stuff and used it until just last year, I am so so grateful that I was able to finally move into a frontend job where I don't have to manage the stack anymore. A new generation of abstracted backend vendors (headless CMSes coupled with Jamstack hosts) makes it so that there is a sub-industry of web devs who never have to touch VMs directly anymore. I, for one, couldn't be happier about that.
(But of course there will always be other use cases that require a fuller/closer to the metal stack, and also backend and ops people who love that work. I don't fault them in the least, I greatly respect them, I'm just glad I don't have to do that.)
Ex:
docker-compose.yml:
hostname: ${HOSTNAME}
.env:
HOSTNAME=foo.bar.com
Made a video of using this pattern with Blazor WASM, SQLite and Litestream here [1].
Docker Compose is fine for running smoke tests/unit tests, but if Dev and Prod run differently, there's really no point to using Docker Compose other than developing without internet access.
(It goes without saying that Docker Compose only works on a single host, so that won't work for Prod if you need more than one host, but single-host-everything seems to be HN's current fetish)
I then find out that no-one bothered to run the compose file that matches what gets put into prod on their local machines.
My point is: your view might become more black and white on this matter after the N-th "but it works fine on MY machine" comment.
EDIT: Where N = your personal tolerance level of bullshit.
There's always one in the crowd that refuses to give up local development.
But even so, how you can go from that to:
> [...] there's really no point to using Docker Compose other than developing without internet access
boggles my mind.
EDIT: apologies, mixup on my end and quote is from someone else.
EDIT: I usually try to build, test, inspect, push and deploy from the same compose file. You can deploy to ECS direct from compose using an AWS context: https://aws.amazon.com/blogs/containers/deploy-applications-...
Straight up 100% hard agree.
> It goes without saying that Docker Compose only works on a single host, so that won't work for Prod if you need more than one host, but single-host-everything seems to be HN's current fetish)
technically this is only correct for compose file versions up to and including 3.8.
Part of the push with the new mainline compose V2 plugin has included the updated compose specification [0] which seems to have unify quite a lot of differences between the "compose" file for compose deployments and the "stack" file for swarm deployments.
FYI You can switch over to the new version with the compose V2 plugin by remvoing any `version` key from your compose file.
Although the `deploy: restart_policy:` values are currently not unified, much to my sadness.
I've heard of and worked with multiple dev teams that develop with compose and deploy a totally different way (k8s, AWS auto scaling groups, heroku, etc), to great success.
Devs know about compose and understand it. As long as the containers behave as planned, the DevOps people can figure out how to deploy them on whatever production environment.
>If you don't use the same Docker Compose file for Production, don't use it for Development. Your two different systems will diverge in behavior, leading you to troubleshoot two separate sets of problems, and testing being unreliable, defeating the whole "it just runs everywhere" premise.
Anyone ever developing a react project on docker compose would never run a build step for development.
In fact for most setups, I probably wouldn't even run react on docker-compose for development, and just use the dev server straight up on my machine.
Furthermore, working with any python http servers is substantially easier when working with a dev server in development, rather than running your server through gunicorn or other wsgi servers - not to mention the ability to hot reload for development.
I'm sure that there are many other cases where it's necessary and convenient to separate production and development docker compose files.
I don't think it's fair to say that there is no point in having separate docker-compose files as the lion's share of dependencies that need to be consistent is inside each container, not on the docker-compose configuration level.
"If dev and prod run differently" — Why is there an "if" here at all? Isn't it obvious to verify with two eyes in the real world? Is there any company in the world that runs production from somebody's laptop at home? Or distributes server racks for people to take home? It looks like an extremely disillusioned question out of touch with reality.
1. For development, engineers want to run a local throwaway mysql/redis. In production, engineer want to use a proper managed mysql/redis. The compose file, and envvars/links will be different. This is extremely normal, it makes sense.
2. Development laptops will always be underpowered than servers. The cpu/memory requirements will be different. Again makes sense.
3. For development, engineers will build and run the docker image locally. In production, these are separate, you would build and publish the image separately in CI, and run only published images in prod. Again, makes sense.
These are all perfectly logical things that make practical sense.
Trying to share/ reuse/ resemble is fine. But saying "dev===prod" is just a disillusion out of touch with reality.
Edit: For context, yes I do docker for a living for last 10 years, since v0.4 (lxc) days. I've been to and presented in meetups. If someone says they're trying to get dev as close functionally to prod, by doing X, Y, Z and sharing practical tips — you know you can listen to them. If someone says "dev===prod" and when you ask how they talk abstract principles, you quietly get out of there.
Bad idea. You develop against a completely different local system, then you push to prod, and "oh no it doesn't work". Yeah - because you're developing against a completely different system! It behaves differently. It causes different bugs. You take a ton of time setting up both different systems differently, duplicating your effort, duplicating your bugs, having to fix things twice for no reason, and causing production issues when expectations from local dev don't match prod reality.
Every experienced systems engineer has been repeating this for decades. Listen to them. They all say exactly. the. same. thing. "Make all your systems as identical as possible." That doesn't mean make all your systems different whenever it is more convenient for you. It means that if it's possible for you to be developing using the same tech you use in prod, you should be doing that, unless it is impossible. If you think it's not possible or practical, you simply haven't thought hard enough about it.
The whole selling point of the cloud is that you can spin up environments and spin them down in minutes and pay pennies for it. Not a shared environment where everyone's work is constantly conflicting, but dedicated, ephemeral environments that actually match production. Literally the only example I know of where this doesn't work is with physical devices where you have a lab of 10 experimental pieces of test gear and you have to reserve them to test code on them. That can make remote development difficult. Still not impossible though.
> The cpu/memory requirements will be different.
Unless you're developing using the same servers as prod. There is no reason that you have to use your laptop. If you think "it's more convenient!", that's only because you have done zero effort to make Prod-like ephemeral environments more convenient.
> For development, engineers will build and run the docker image locally. In production, these are separate, you would build and publish the image separately in CI, and run only published images in prod. Again, makes sense.
That is the antithesis of what Docker was created for. The point of containers was to say "it works on my machine", and use that exact same image in production. Not a different image in CI that nobody developed their code against. It's just going to result in bugs that the developer didn't see locally, so they will need to spend more time to get it working. Your laptop can do the same build CI can on the same branch and push the same image that CI can.
What I'm saying is nothing new. Look into cloud-based development environments and the dozen different solutions made solely for rendering your local code in a remote K8s pod as soon as you write to a local file. Imagine a world where you don't waste your time testing something twice and fixing bugs twice.
We haven't achieved it (yet?) but that's the direction I was aspiring to!
Managing all these different configurations and the matrix of interactions between them is why we're trying to integrate the SDLC into one configuration at Coherence (withcoherence.com). [disclosure... I'm a cofounder]
I am now using this setup[0] for development, but it doesn't really match production 1:1.
In my company we only set CPU requests + Memory req+limits. I read somewhere that CPU limits are kinda broken and are not reliable. In some experiments CPU limits tend to cause unnecessary throttling, which is why some people don't recommend them.
However, I've also found some cases where I can't make the app to use only some of the CPUs, some frameworks tend to read the number of total available CPUs on the node and try to use them.
[wsl2]
memory=10GB
processors=4
swap=2GB
By limiting wsl2 I also limit the docker containers that run inside wsl2.I think CPU limits (and also, choosing to set memory limits equal to the requests or higher) depends on your goal. If you want to get the best use out of your hardware, set low requests and high limits and you'll be able to minimize "wasted" scaling. But if your goal is consistent performance and availability, setting high requests and high limits (usually equal to each other) save you from a whole host of problems.
Just my take. I have a fair bit of experience but am far from an expert on the topic. And my experience is only on the scale of dozens of applications/dozens of nodes...people with higher scale experience may feel differently.
docker-compose.yml (build current dockerfile) docker-compose.image.yml (fetch stable image)
To avoid duplication, the second file only specifies the image.
It lays out every single hard lesson I discovered when working with docker compose.
Great minds think alike! (Or I just got lucky)
who would do this??