Best practices around creating production ready web apps with Docker Compose
nickjanetakis.com
nickjanetakis.com
For anything more serious, Terraform makes much more sense IMHO
Shame really, I find compose simple and neat and k8s+terraform an unsettling mess.
I haven't used either of them in production though
Personal or small/medium business sites, research software, free web utilities
What's considered lean in startupland can get pretty beefy compared to the shoestring a lot of projects are compelled by necessity to get by on
Compose can be great if you just need a web server and maybe a database and background worker and you don't have anybody suing you or losing millions for the occasional maintenance hour
It is an argument I've lost with managers more times than I've won though.
1) create a VM, set it up with docker compose to auto start -- have it reference code on a network drive.
2) take a an image of 1)
3) create a load balancer, balancing that image. GCP load balancers/instance groups even a way to rollout images.
That is %80 of k8s right there (endpoint, scaling, containers, rollouts).
K8s was a little too complex and docker swarm a little too simple. Rolling our own solution was a last resort, but it worked out well in the end - more control, deeper understanding and no nasty surprises in behaviour while scaling etc.
My home setup is now k3s.
A lot of my clients are happily using Docker Compose in production. Some even run a million dollar / year businesses on a single $40 / month server. It doesn't get handed off to Swarm or another service, it's literally docker-compose up --remove-orphans -d with a few lines of shell scripting to do things like pull the new image beforehand.
You can still use Terraform to set up your infrastructure, Docker Compose ends up being something you run on your server once it's been provisioned. Whether or not you decide to use a managed database service is up to you, both scenarios work. The blog post and video even cover how you can use a local DB in development but a managed DB in production using the same docker-compose.yml file, in the end all you have to do is configure 1 environment variable to configure your database at the app level with any web framework.
I rather would use kompose+kustomize to convert it to k8s and than use k3s with ingress-nginx with hostPath and backup local volumes with velero.
Backups? You only need backups or the Db, and that have zero impact in the use of docker or not*
*ie: I know what very complex setups the micro-service crow use for serve their smalls-size disruptive app. I still think cron + pg_dump + rsync is more than enough...
No downtime deployments: I don’t bother for these projects. A minute of downtime in off-peak hours for a restart is fine.
Data backup: just like you would backup any other local data: cron, rsync, pg_dump. And there isn’t much local data anyway as most services are stateless, using S3 or a database for long term storage.
I choose docker compose for production when I want to keep complexity to a minimum: All I need is a cloud vm, an ansible playbook to install docker and configure backups. Typically in these scenarios I run the database as an OS service and only use docker for nginx/redis and application services.
Using it in docker would allow easier local development, and shorter setup time.
For production some of these projects run alongside each other on the same host and share the database service. Sometimes the database is on another host.
A typical host like this would be a 50/month Hetzner dedicated machine with only docker daemon and postgres installed natively. And then multiple projects deployed with docker compose. Sometimes nginx is native as well with vhosts pointing to the docker projects.
Edit: s/ghosts/vhosts/
Ex:
- blue/green stateless nodes
- DNS/load balancer pointer flips
- Do the DB as a managed saas in the cloud and only ^^^ for app servers
- Backups: We've been simply syncing (rsync, pg_dump, ..) w/ basic backup testing as part of our launch flow (stressed as part of blue/green flow), and cloud services automate blob backups. We're moving to DBaaS for doing that better, which involves taking the DB out of our app for our SaaS users
In practice, our service availability has been largely degraded by external infra factors, like Azure's DNS going down, than by not using k8s/swarms for internal infra. Likewise, our own logic bugs causing QoS issues that don't show up in our 9's reports.
It doesn't seem you bothered looking into alternatives.
Take for example docker swarm. It already uses docker compose files (well, there are some differences) , they support blue-green deployments of stacks, and they even support multi-node deployments and horizontal scaling.
you'd rather use six complex things instead of one simple one?
> we do not care about no downtime or stateless nodes and a lb in front
and b:
> backup with the standard tooling or stateless nodes
so basically they just use it mostly to "package" the application and everything else is probably done with different tooling.
By that measuring stick, using Docker Swarm instead of docker compose is even simpler as you only need to use docker alone instead of having to install a separate script.
Things are way simpler with Docker Swarm. You just run a single ingress controller as a separate stack and route all traffic internally.
If you use Traefik as your ingress controller, all you need to do is set labels in your docker compose files and everything just works.
There is absolutely no reason or excuse to use docker compose in anything resembling production.
I work at a university where I make an app meant for maybe a couple of concurrent researchers. Am I seriously supposed to set up Docker Swarm for this or is it OK with you if I run my little setup using Docker-compose on the tiny server that the university has provisioned for me?
Seriously, why do so many people on HN have to act like the entire world of software developers is busy making the next Facebook? I'm sure many (most?) people are making tiny things meant for tiny audiences.
I agree generally, but playing devil's advocate for GP, what I think they meant is that swarm can be run on a single node with little downside/complexity and a lot of upside from docker compose.
See for example:
https://jarredkenny.com/single-node-docker-swarm/
https://devopstuto-docker.readthedocs.io/en/latest/docker_sw...
Manage deployments as stacks, which support history and backtrack deployments.
Blue-green deployments.
Frankly, it boggles the mind how people put to production devtools that are barely managed to be secure when deployed locally, when alternatives are both simpler to deploy and more capable.
I'm sure you're great at setting up servers. That is probably the part of my job I dislike the most, but it's something I have to do, so I do it.
https://docs.docker.com/cloud/aci-compose-features/ https://docs.docker.com/cloud/ecs-compose-features/
I take the few second downtime hit on each deploy. I know, it's iNsAniTy but it's really not. A 5 second blip a few times a week is no biggie for most SAAS apps or services.
To me that's well worth the trade offs of not having to deal with a load balancer, a container orchestrator tool, dealing with database migrations that need to be compatible for 2 versions of your app as it's being rolling restart and all of the other fun complexities of running a zero down time service.
You can still architect your app in a flexible way (env vars, keeping your web servers stateless, uploading user content to S3 or another object store, etc.) so that moving to something like Kubernetes is less painful if / when the time comes. It's just for me, that time has never come.
> how do you backup data?
A background job / cron job runs every few hours and does a pg_dump to a directory in a block storage device. Could easily change that to be S3 too. Just comes down to preference. Personally I found it easier to drop it into block storage since it's just copying the file to a directory.
Very modern stack but totally overengineered.
It's common to have a load balancer, that routes traffic to nodes that are up.
That part is mostly straightforward. It could be the application's responsibility or the load balancer's. The tricky bit is maintenance on the load balancing parts.
Then in dev the restart policy is set to no. This is handled with an environment variable.
https://gitlab.com/stavros/harbormaster
You give it a config file with a few repos and it pulls/restarts whenever one changes.
Use case is when a repository requires compilation to build the docker image, and I don't want that happening on the production server.
It doesn't currently check upstream Docker registries for changes like Watchtower does, mainly because I think that's not a deterministic enough way to do deployments (but that can certainly change).
Why hand it off to anything?
> When would you choose this path over a separate system to provision your production resources
When you have a simple system with modest availability requirements and you don't like unnecessary complexity.
I see a stark difference between infra management tool (like Terraform), and deploy tooling. Infra I prefer to describe "declarative", and Terraform helps a lot with that. Deplyments are "imperative" in nature to me: turn on maintenance page, remove cluster X, spin up cluster Y, run migrations, etc.
So I find Terraform a bad fit for deployment tooling (please explain if you disagree, I'd like to learn).
https://docs.docker.com/cloud/aci-compose-features/
https://docs.docker.com/cloud/ecs-compose-features/
I haven't (successfully) used ACI to deploy my compose app. I'm hoping somebody else on this thread might have.
However, I prefer (at least for my needs) AWS CDK because it also supports pushing of the image to remote registry, and spawning other AWS resources, such as databases.
If your app fits on a single VM, and you don't have k8s, this is incredibly straightforward.
It's even "easier" to just use kubernetes to deploy something - but that's if you already have kubernetes! As my org is adopting k8s we are favoring that over docker-compose where applicable, but there's a huge complexity gap between docker-compose and this isn't an overnight transition.
I guess you could view docker-compose as a much more accessible stopgap as your team learns how to run the more complex orchestration strategies, or a preferred solution if you have a small number of VMs you want to explicitly configure to run specific services on (where you don't need the scheduler etc).
I'm fairly surprised by this. You can deploy a fairly similar compose file into a Docker swarm cluster, but if you're doing a single-server deployment of everything, that is awfully brittle for production, to say nothing of considerations brought up in other comments where you may be using managed services from a cloud provider for things like database and caching in production but you're probably just spinning up a local postgres server on your laptop for development.
Still, someone already mentioned docker-desktop having some tricks for deploying straight to aws and [there is definitely a way to do this with ecs](https://aws.amazon.com/blogs/containers/deploy-applications-...). There's also tools like kompose which translate configs automatically and are suitable for use in a pipeline, so the only thing you need is a compose file in version control.
So yes, you can still basically use compose in production, even on a non-swarm cluster, and it's fine. A lot of people that push back against this are perhaps just invested in their own mad kubectl'ing and endless wrangling with esoteric templates and want to push all this on their teammates. From what I've seen, that's often to the detriment of local development experience and budgeting, because if even your dev environment always requires EKS, that gets expensive. (Using external/managed databases is always a good idea, but beside the main point here I think. You can do that or not with or without docker-compose or helm packages or whatever, and you can even do that for otherwise totally local development on a shared db cluster if you design things for multi-tenant)
At this point I'll face facts that the simple myth here (docker-compose is merely a toy) is winning out over the reality (docker-compose is partly a tool but isn't a platform, and it's just a format/description language). But consider.. pure k8s, k8s-helm, ECS cloudformation, k8s-terraforming over EKS, and docker-compose all have pretty stable schemas that require almost all the same data and any one of them could pretty reasonably be considered as a lingua-franca that you could build the other specs from (even programmatically).
From this point of view there's an argument that for lots of simple yet serious projects, docker-compose should win by default because it is at the bottom of the complexity ladder. It's almost exactly the minimal usable subset of the abstract description language we're working with, and one that's easy to onboard with and requires the least dependencies to actually run. For example: even without kompose it's trivial to automate pulling data out of the canonical docker-compose yaml and injecting that into terraform as part of CD pipeline for your containers on EKS; then you keep options for local-developer experience open and you're maintaining a central source of truth for common config so that your config/platform is not diverging more than it has to.
I'm an architect who works closely with ops, and in many ways not a huge fan of docker-compose. But I like self-service and you-ship-it-you-run it kinds of things that are essential for scaling orgs. So for simple stuff I'd rather just use compose as the main single-source-of-truth than answer endless bootstrappy questions about k8s or ECS if I'm working with others who don't have my depth of knowledge. (Obviously compose has been popular for a reason, which is that kubernetes really is still too complicated for a lot of people and use-cases.) Don't like these ready-made options for compose-on-ECS, or compose-to-k8s via kompose? Ok, just give me your working docker-compose and I'll find a way to deploy it to any other new weird platform, and if I need some pull-values/place-values/render-template song and dance with one more weird DSL for one more weird target deployment platform, then so be it. I've often found the alternative here is a lot of junior devs deciding that deployment/dev-bootstrap is just too confusing, their team doesn't help them and pushes them to an external cloud-engineering team who doesn't want to explain this again because there's docs and 10 examples that went unfollowed, so then junior devs just code without testing it all until they have to when QA is broken. Sometimes the whole org is junior devs in the sense that they have zero existing familiarity with docker, much less kubernetes! Keep things as simple as possible, no simpler.
Seen this argument a million times, and no doubt platform choices are important but even pivoting on platforms is surprisingly easy these days. When you consider that compose is not itself a platform but just basically a subset of a wider description language, this all starts to seem a bit like a json vs yaml debate. If you need comments and anchors, then you want yaml. If you need serious packaging/dependencies of a bunch of related microservices, and a bunch of nontrivial JIT value lookup/rendering, then you want helm. But beyond org/situation specific considerations like this, the difference doesn't matter much. My main take-away lately is that leadership needs to actually decide/enforce where the org will stand on topics like "local development workflows"; it's crazy to have a team divided where half is saying "we develop on laptops with docker-compose" and half is saying "we expect to deploy to EKS in the dev environment". In that circumstance you just double your footprint of junk to support and because everyone wants to be perfectly pleased, everyone is annoyed.
That of course is way close to the "production" in "production ready".
Different users run different containers, and no user can see other user's containers.
Like, a staging cluster? A pre-prod cluster?
But like I say, dev still is far from close to that. So you end up doing a lot of development in dev, and then fixing a ton of things in staging because you weren't developing against staging. And your tests aren't fantastic, so some bugs make it through both staging and prod, and it turns out they're because your design was making assumptions based on dev. Hence, make dev and prod as close as possible, and then you don't need to have "staging", because you have confidence that what worked in dev will also work in prod.
Edit: RDS is great for the DB in this setup!
Minikube is your personal kubernetes, and it's very to use, and is very close to the production kubernetes, but on a single node.
I don't see enough praise for minikube on HN, but I think it's deserved.
Albeit, you cannot do a "docker-compose up --build", but instead "docker-compose build && kub apply -f <template.yaml>", which is equally fast IMO.
https://docs.docker.com/cloud/aci-compose-features/ https://docs.docker.com/cloud/ecs-compose-features/
There’s a low chance of this scenario happening, but lord knows there’s a Unix greybeard somewhere that doesn’t want Docker on his precious pets, and will like to know what permissions your app actually needs.
If not then you’re playing a dangerous game because now your container is running as root on the host system which is almost certainly not what you want. The fundamental security boundary on Linux is users and containers are not a sandboxing solution. You can use seccomp, SELinux, gvisor, or honestly virtualization to get real sandboxing and docker can use those tools to create a sandbox but Linux namespaces alone “containers” aren’t that. You can use these tools without any namespacing and vice versa. It’s not a question of “not trusting” process isolation but defense in depth. If you wouldn’t be comfortable running every process on your host system as root and put full trust into SELinux, AppArmor, and seccomp then you shouldn’t start just because you added a high-tech chroot.
Now, there are other alternatives with less downtime, more scalability and a lot more or complexity? Sure there are... but for my personal box I won't be touching any of that, thanks
Edit/Add: ... and also VPN, which I added later on in like, 10m or so. Adding OpenVPN in my previous box was not as trivial, I tell you that.
Me too. I have about half a dozen personal projects running in Docker Swarm on Hetzner. I have an Ansible script that gets a freshly created barebones VM to serve live traffic in less than a minute, and I am free to just delete all VMs and reinstall from scratch without any problem at all.
I even split the Ansible playbooks into stack-specific playbooks so that I can redeploy specific stacks if I feel like it.
I firmly believe that Docker Swarm is a killer application that is not given the attention it deserves because everyone is stuck in a mindset where they purposely want to make their own lives harder by doing CV-oriented development while running after the devops version of doing enterprise versions of hello world.
The health check should only include the health of the service it resides within, not its dependencies.
Kubernetes using Compose would be a killer product. People would pay for it. Its funny why nobody built it.
Interestingly, now there is Cue lang + Buildkit - https://docker.events.cube365.net/dockercon-live/2021/conten...
e.g. suppose you have your postgres db url in an environment variable - with the password in the string - isn't that a big no-no? If someone found an exploit in your web app container then they would also in theory get access to valuable secret info - the db data is by far most valuable part of your app imo.
Is there a more managed way to handle this kind of secret?
A good pattern here is to reverse proxy all requests to the application through something like nginx. My applications tend to have a back-end application that is not accessible to the internet, with an nginx instance that proxies all API requests itself. Only port 80 is public facing. If someone can get console access to an nginx container and then use that to springboard to another container and get root access there (again, where the only open ports are ports 80, maybe 8000?) to get envvars, they should get access to it all.
If you are worried about secrets, check out the Docker Compose 3.9 documentation: https://docs.docker.com/compose/compose-file/compose-file-v3...
Environment variables are typically how secrets are used. Anything else is going to add further complexity and have another set of caveats. Env vars are simple and all devs (and programming languages) know how they work.
If you want to limit the blast radius, you will need to find ways to limit what parts of the database the web app has access to.
It's mentioned on video and briefly in the blog post. It's so you don't have to duplicate a bunch of services between 2 files while you slightly change a bit of configuring for those services depending on the environment you're in, or even running different containers all together (which is the problem the override file solves).
> <<: default-app
> ...
> worker:
> <<: default-app
This is always somewhat of a code smell to me. This makes me want to dig deeper and see if the two use cases shouldn't have been separated into separate Docker files or even separate apps.
And that's the general flaw in using Compose outside of local development. It's simply not designed as a general purpose (and thus large-scale) deployment orchestrator. If you get it working locally, you probably need a different setup for remote. That affects things like testing, which means it's not really the same in dev as in prod. And even if it were, Compose isn't persistent, meaning it's just a script to start jobs on some other thing.
Every time I've worked on a team that tried to do Compose in production (or even just for CI) it ended up a weird snowflake or with duplicate processes, and we replaced it with something more standard. My standard advice now is to model both production and local development to be as close as possible and then pick the simplest method to support that use case. In general this means going from "Compose on one host" to Swarm/Nomad/ECS on multiple hosts to K8s. And of course, try to make your CI use the same exact methods you use for dev and prod, so that using CI doesn't give you different results than in dev or prod.
Tutorial: https://docs.servicestack.net/do-github-action-mix-deploymen...
Video: https://youtu.be/0PvzcnxlBvc
> Or perhaps you want to use a managed PostgreSQL database in production but run PostgreSQL locally in a container for development.
It’s mentioned in passing with no discussion of why you should not run your database in docker compose - a very common use case - so the reasons I mentioned were not explicitly discussed.