Running Docker in production for 6 months
racknole.com
racknole.com
At Convox we have been running Docker in prod for 18 months successfully.
The secrets?
1. Don't DIY. Building a custom deployment system with any tech (Docker, Kubernetes, Ansible, Packer, etc) is a challenge. All the small problems add up to one big burden on you. 6 months later you look back at a lot of wasted time...
2. Don't use all of Docker. Images, containers and the logging drivers are all simple great. Volumes, networks and orchestration are complex.
3. Use services. Using VPC is far simpler than Docker networking. Using ECS is much easier than maintaining your own etcd or Swarm cluster. Using Cloudwatch Logs is cheaper and more reliable than deploying a logging contraption into your cluster. Use a DB service like RDS is far far easier than building your own reliable data layer.
Again thanks for sharing your experience as a cautionary tale.
If you are starting a new business you should not take on building a deployment system as part of the challenge.
Use a well-built and peer reviewed platform like Heroku, Elastic Beanstalk or Convox.
As nzoschke recommends, we rely heavily on ECS, RDS and other managed services. We are very careful about using exotic new Docker features until somebody else has successfully used them in production. We use DNS, load balancers and regular networking for discovering containers and communicating between them.
And it all basically just works. The worst problem we've encountered is that twice a year, we deploy a new version of ecs-agent to our staging cluster and need to revert it because of issues.
For us, the biggest challenge has been setting up a good local workflow for developing Docker apps with multiple services and multiple underlying git repositories. The docker-compose tool is great but it doesn't go far enough and doesn't provide enough structure. We've open sourced some our internal Docker dev tools here http://blog.faraday.io/announcing-cage-develop-and-deploy-co... but we think there's a lot more which could happen to make it easy to develop complicated apps.
I never liked Cloudwatch Logs and recently switched it for an ELK stack (hosted by logit.io).
I suppose it's much relevant for those of us who are not on AWS.
If the problem is Docker complexity that needs to be solved. If the problem is difficulties of running stateless apps, or distributed storage and networking these need to be highlighted and widely understood.
Using Convox or any other platform will not magically make them disappear. You still need to troubleshoot and understand what you are running. Another layer to hide the complexity can hardly help.
I strongly agree! I'm sensitive to self-serving, but while Convox is another layer, it does make problems disappear by disallowing them.
Use Convox and try to boot or deploy a docker-compose.yml that uses networking or uses volumes incorrectly. You are blocked from doing so with a nice reason why.
> If the problem is difficulties of running stateless apps
This is not the problem. In fact this is a solved problem.
> or distributed storage and networking
These are hard problems.
Networking is largely solved with AWS VPC. Other providers do this very well too. This lends to your earlier point, adding the Docker networking layer generally doesn't help the problem. I take it for granted that any new networking stack shouldn't be trusted.
Distributed storage is barely solved anywhere.
Then proceeds to name every problem kubernetes fixes
Now add container management - understanding the split between images and containers, then how to publish and download images (which IMHO should be a completely separate project), how caching works so you could write a decent Dockerfile... And then there are features like Docker Swarm which I never touched and seem particularly complex to me.
It's way too many projects stuffed into one thing.
It felt like every other link led back to the various sales teams for the enterprise solutions, and there existed no page with a straightforward overview of all the components and what puppet / chef actually did.
Ansible was a breath of fresh air - as well as saltstack (but at the time it had too many holes in its security/transport story).
I'm a NixOS fan, but have to use Docker at work occasionally and I have the feeling they are both difficult, but in "their" ways.
But could you elaborate on the "Nix-env conflicting with system profile" thing? :)
I support the idea of declarative configuration and dependencies, but I also realize that this is a very very hard thing to do, when each language have its own preferred package manager and they all expect a mutable environment.
(I tried NixOS for 2 months and switched back to Ubuntu. The issues are too numerous so far, despite heroic work of NixOS dev team)
I really don't like docker.
I the configuration.nix of NixOS is a nice, small way to get your dev machines configured and installed with the stuff you like. Just copy it on your new machine and off you go :)
I don't know if I would ever use Nix instead of the more language specific packet managers.
It was a very frustrating experience. A frustration which was led by the fact that i could tell how powerful Nixos was - if only i could grok it.
I love nix, my biggest gripe would be that the Nix language is dynamically typed...
This article is a great example where in using docker for the first time they chose to attempt to run the state layer(database) in a container. State layers tend to be always available and difficult to scale horizontally anyway. This cancels out a decent chunk of the benefits of containers in production. If that team had more experience with service based architecture they would have known to use an third party database service provider(RDS) or host their own on a persistent server. Especially on their first experience with containers.
One question I've had as I evaluate moving a Rails/Postgres app to GKE is whether I should abandon the thought and go with AWS/RDS, or containerize Postgres on GKE, or some third option. Has anyone else been part of a similar migration onto GKE, and how did it go?
not a fan.
They are the last to adopt new things as they tend to be the lowest on the talent chain as meeting estimates, often via client coercion or moving goalposts, is much more important than successful solutions over the mid-ling term.
For anyone reading this that is part of that group, leave stuff like containers to outside organizations with stronger engineering talent and focus on what ultimate makes you money which is client networking and sales.
I don't know where did you get that impression from. The article clearly states we ran docker in PRODUCTION for 6 months. The article wasn't written after a docker trial over a weekend.
>> How much you actually read the documentation will be immediately and plainly obvious and inversely proportional to the amount of problems you run into.
Pretty much everyone who's commented here and have had docker experience in production, agrees that a docker documentation sucks. Maintainers of the dockers project have already conceded that this is an issue so I don't why you're saying this.
You can be forgiven to have that perception. Sorry, but that's not true.
>> They are the last to adopt new things
Are you actually complaining about this? Well it just shows your immaturity as an engineer and your lack of understanding of how tech startups are run. This is something my very first manager drilled into me on my first job a decade ago, as should yours have - not to use any v1.0 software for ANY CLIENT PROJECT, but only for your hobby projects. Wait atleast for a v1.1
>> as they tend to be the lowest on the talent chain as meeting estimates
Please don't embarrass yourself. There are comments on this thread from the maintainers of both docker and kubernetes projects as well as people who have been running docker in production for much longer than us. None of them think that this article is stupid. Now, if you think you're more talented than us and all of them combined, just tell us why?
Your comments would be much more useful if you could say something like - "They are so stupid because they couldn't figure out X which was as easy as doing Y."
I'm sure we'll end up finding another way that's similar (provides the same benefits for us) but without such crazy image sizes.
I've deployed a global unit into our coreos cluster that regularly executes:
#!/bin/sh
while :; do
# remove stopped containers w/ volumes
docker ps -a -q -f status=exited | \
xargs -r docker rm -v
# remove dangling images
docker images -f "dangling=true" -q | \
xargs -r docker rmi
sleep 6h
done
You need to be sure that you don't lose important data when running something like this in your setup, but it works nicely to remove old images. This script is deployed to non-coreos servers as well.I want to put my application, any application, in a nice tidy box, ship it to a server, any server, an be confident that it'll work. That's the promise of Docker. To me, Docker doesn't deliver on that promise if I have to copy&paste a 12 line bash script from HN to be able to do that or else face out-of-space problems at surely the worst possible moment. If I have to jump through hoops like that, then what's the gain?
Might as well just install an Ubuntu VM, apt-get everything I need by hand or with ansible and add my app to the startup script. I'll have similar complexity in a more mature and well-understood environment.
Honestly didn't even realise this was such a big deal for so many.
Check it out at https://coreos.com/rkt/docs/latest/subcommands/gc.html
`kubelet --container-runtime=rkt` and be on your way.
I've talked about the relative immaturity of Docker as a used system (outside of dev) [3] and am struck often by how rarely people understand that it's still a work in progress, albeit one that can massively transform your business. The hype works.
That said, Docker can work fantastically in production, but you need to understand its limits and start small.
[1] https://www.amazon.com/Docker-Practice-Ian-Miell/dp/16172927... - working on 2nd edition, if anyone has any suggestions @ianmiell
[2] Blog: https://medium.com/@zwischenzugs
As someone starting considering Docker (and possibly Swarm), these seem to be pretty serious criticisms. Any experiences to corroborate / counter these two posts? Going by what's written here it would be suicide to use Docker, but many people are...
I would suggest starting off by stepping back from picking specific technologies and architecting your system properly first. Where do you need load balancing? What are the different parts of the system? How can you break those parts up so that different teams can work independantly from each other? What microservices do you really need? Do you actually need microservices? How will they communicate?
If you start off by having to design everything around Docker (or any other implementation detail) then you're going to have a very brittle system that's going to cause you pain in the long run. After designing things fully you may realise that you've managed to eliminate most of the initial perceived complexity and can actually work just fine with more boring tools.
Docker and swarm don't actually solve any of feature set X. You have to do that yourself, and while they may help facilitate a solution, they're not going to really do anything for you. Lots of people have jumped onto Docker without really understanding what it's doing for them, which is why it seems like everybodies using it.
If you don't know which problems you have the Docker is a good fit for solving, don't use it yet. If you can't fit it into your process at a later date, it probably wasn't a good solution for you in the first place, so you'll have saved yourself a headache.
> failover, load balancing, local integration testing, mixed language platforms,
Docker does none of that.
It's only a packaging and deployment system. You package the app as a docker image, then you can call a docker command on any system to grab that image and start it.
Without docker:
1) You'd make a zip/deb/rpm of your application.
2) Download the zip to some servers
3) Update the dependencies & systems stuff
4) Start the app
With docker:
1) You'd make a docker image [basically: run a script to install the app and the dependencies, as in the previous steps]
2) Save that image to the container registry
3) Deploy & Start the image on some systems
Jails (from bsd) and chroots are a great idea. But a way is needed to manage the file system of a) the chroot (the c libraries, the configuration files the application code). This is docker image; b) persistent data (database, images and binary user data etc). As far as I can tell docker doesn't really come with a compelling story here - something that's easier to manage and gives high performance (say something that competes with iscsi for database files, and a solid out-of-the-box clustered filsystem).
Now, docker gets (justified) hype for pushing the jail/chroot (aka "container") idea. But I think a lot of people (possibly including docker Inc) think that docker does much more (and do those things well) beyond being a nice-ish set of tools for building and managing self-contained chroot file systems for applications ("images").
Docker Inc certainly is working on "everything else" - but I think moat would be well-served to look at Lxd/lxc if what you want is "lightweight Linux vms", or kubernetes if what you want is to move towards a "container/chroot (micro) service paradigm".
Kubernetes might seem a bit complex, but that is because it tries to solve a complex set of problems.
Docker is more like "yo! Synchronise your /etc/passwd file across systems so you can log in to all your machines", while kubernetes is more like LDAP+dns+kerberos. More complex, but more sane. And built not just to get started, but continue to work as your system evolves.
And I've been quite happy playing with Ubuntu and Lxd + zfs on the other end - the simple light weight Linux vm end.
When you have multiple containers, it is imperative that they are managed as a group, with clearly defined dependencies. This is why Docker requires an orchestration tool. And when the number of your production nodes is greater than one, Docker on its own becomes inadequate for that.
Can't the above be orchestrated via docker-compose?
(I've been using docker with docker-compose for my side projects but not been using it excessively in customer facing production)
For anything more than that, native Docker tooling is inadequate.
If you have multiple devs...how does docker-compose fall short? (I'm asking because I wanted to suggest using docker for my team and wanted to know it's short comings before I suggested it).
From my understanding: you can use compose to orchestrate databases, cache, app all on one machine quite easily (and pass environment variables, get them linked etc.)
- Secret distribution
- Managing persistent volumes
- Monitoring containers for failure, restarting according to policy
- Service discovery and DNS integration
- Integration with load balancers, setting up routes, etc.
- Managing affinity / non-affinity for containers
- Sharing resources on a cluster via namespaces
This alone warrants an orchestration solution, even on a single machine.
At the container level, Docker containers are the basis for the container specification from the Open Container Institute. The kertuffle is over orchestration. Swarm is a feature from Docker-the-Company that trys to make 'Hello World' container orchestration Ruby-on-Rails easy. Right now the alternative orchestration layers are more toward the "Apache server man page" end of the spectrum.
Essentially, Docker-the-Company and the Docker critics are focused on different contexts. Docker-the-Company thinks container orchestration on a Raspberry Pi is worth pursuing. The Docker critics are coming from a world where CentOS 5 is still relevant (metaphorically).
Could you elaborate on this? Did you settle on an orchestrator, and if so, which one?
Here are some things that might help:
* minikube: http://kubernetes.io/docs/getting-started-guides/minikube/
* the new tutorial: http://kubernetes.io/docs/tutorials/kubernetes-basics/
* deploying in 2 steps with kubeadm: http://kubernetes.io/docs/getting-started-guides/kubeadm/
I'd love any honest feedback on how we can either (a) fix the learning curve, or (b) remove the perception that it's steep.
Overview
asics
through of the basics of the Kubernetes cluster
module contains some background information
...There's no way to zoom out, and there's 2 hamburger menu : one white on black on the left and one blue on white on the right, both are broken The 2 other links work well
If so, please raise this here (with a screenshot perhaps)
The system has matured incredibly over the last year. There are also third party tools like kompose/compose2kube which will convert docker-compose deployments into kubernetes manifests.
I think that a problem that leads to a lot of these articles being written is that the motivation of the authors to use Docker is unclear.
Why bother to put your database inside a container if the way you ran your databases before worked just fine?
You shouldn't just rush to put all your processes in containers because it is cool and Containerization Is A Good Thing. You should use technology that makes sense given the problems you are trying to solve at a give moment.
For the same reason you'd bother putting anything in a container at all. The purpose of docker -as advertised by docker- is to containerize EVERY service. Unless you can point me to a single mention in official documentation that databases are an exception. Till date, I haven't found any and thus this blog.
Don't get me wrong, docker is one of the most frustrating technologies I've used (partly because it shows such promise), but a lot of the problems he describes can be sorted with the most cursory Google.
My experience with Docker is that Google searches often turn up with configurations that other Google searches will say are a bad idea. There don't seem to be any kind of well-documented emergent best practices.
So, something like docker run -it -v /path/to/source:/app my_image /bin/bash
The article _specifically_ refers to production environments, and, by extension, staging servers:
“Good practices dictate that you don’t mount your source code directory in the docker container in production. Which means you also have to rebuild the image on test/staging server every time you make a single line of code change.”
Then it also states in the “logging” section that, contrary to the development environment “On production, since your source code directory isn’t mounted in container […];” so they clearly _are_ mounting the source code in the development containers.
Also something that can technically be done with Docker, just it's really not a good idea (and that has nothing to do with Docker).
Unless you change your system packages or add a new line to your requirements.txt, in which case the cache would be invalidated and the build takes longer.
But... it's done by your CI server anyway, so...
We didn't. Please read the relevant sections again carefully.
What's the point then?
Moreover. Docker is locking away processes and files through its abstraction, they are unreachable as if they didn’t exist. It prevents from doing any sort of recovery if something goes wrong"
"A crash would destroy the database and affect all systems connecting to it. It is an erratic bug, triggered more frequently under intensive usage. A database is the ultimate IO intensive load, that’s a guaranteed kernel panic. Plus, there is another bug that can corrupt the docker mount (destroying all data) and possibly the system filesystem as well (if they’re on the same disk)."
1. Containers are not ephemeral. They have a lifecycle. Data written in the container is persisted to disk and available after the container is stopped and then started again.
2. Processes/files/etc are not locked away as if they don't exist. See `ps aux` on the host. You will see all the processes running. You can inspect the filesystems for each container, etc. There is no magic here.
3. A database crash could cause data corruption inside a container or not. This has nothing to do with the container, and chances of a database crash are not made worse by being in a container.
That said, I would let a volume driver manage persistent storage rather than manually managing this through the host fs... but that's my preference.
--- EDIT --- Disclaimer: I work at Docker Inc, and am a maintainer on the Docker project.
Even if you don't mount any folder from the disk onto the container? Are you sure? Then everything I know about containers is just wrong.
How this happens is dependent on the storage driver used. The `aufs` driver (default when available), as well as `overlay(2)` and `vfs` drivers just sit on top of the existing filesystem at `/var/lib/docker` (or the defined docker root). BTRFS, ZFS, and devicemapper must be pre-configured to even use and depends on how you configure these, but still generally would be on an actual disk.
I've found a good place to visualize how the filesystems work was this blog: http://merrigrove.blogspot.com/2015/10/visualizing-docker-co...
To harden this even further, you can run clustered DB nodes in Docker (+<your_preferred_orchestration_tool>) quite easily. So with persisted data, multiple node replication, and server snapshots I'd be interested to know as well.
0. The lifecycle of docker containers is an extremely complex topic with limited documentation. It's safe to assume that it's out of reach for 9X% of readers here. One needs to fully understand the lifecycle of their containers to attempt to run databases in Docker, that's a huge barrier to entry. Advising 100% of people to run production (i.e. permanent, long lived) databases in Docker is terrible advise.
1. The entire concept of containers is based on being ephemeral. They do have a storage (in /var/lib/docker/<cryptic-structure>) and they should be started with -rm to make sure that everything they did is cleaned up automatically after they exit. If you want to keep the data and make something around that, good look with that!
2. Wrong. There is a truckload of magic going on here from filesystems to networking. Docker is hell to debug. A fucked database hidden away in Docker will be close to impossible to debug. If you're a sysadmin, you do not want to be in that position, trust me.
3. The odds of a database issues are at lest 3 orders of magnitudes higher if running within Docker. The docker ecosystem is notoriously unstable and the filesystems are unreliable. (Plus Databases are IO intensive which is gonna trigger all the rare bugs and race conditions).
Seriously. If you got a brain cell at Docker Corp. PLEASE STOP overselling your product and advising it for absolutely everything without considerations for what people are doing.
Every time one of you guys advise to run databases in Docker, you're objecting to everything that docker stands for (i.e. statelessness). Not only it is confusing the hell out of people but it's putting them on a guaranteed path for future catastrophic failures.
Running production databases inside docker. Just because it's not strictly impossible, doesn't mean it's possible.
[See RFC1925 https://tools.ietf.org/html/rfc1925 ]
(3) With sufficient thrust, pigs fly just fine. However, this is
not necessarily a good idea. It is hard to be sure where they
are going to land, and it could be dangerous sitting under them
as they fly overhead.1. This is simply not true. Your understanding is that they are based on being ephemeral, but this is not inherent in any sort of design of containers.
2. Magic is not really magic when you understand what's happening. Cgroups apply resource limits on a process, namespaces limit what a process can see. These come together to make containers. The host still has full visibility on these processes just like any other process on the system.
3. Do you have data to back this up? A container is just a process that is namespaced and resource limited. If you are writing to the copy-on-write filesystem provided for the container with a database, then you are doing it wrong (in 99% of cases). For that matter, you can even use ZFS for the container FS, which has been in use in production scenarios for quite some time... performance may not be great with ZFS here but integrity will be (not that I'm advocating for writing directly to the container FS... not at all, really).
There is nothing about Docker and statelessness. It can sure make cleaning up after a process a bit simpler but this doesn't mean that docker equates to statelessness.
Storage is hard whether you are in a container or not. Process isolation does not affect this.
1. The stateless & The ephemeralness & The tooling. It all goes together. Just because its not enforced all the time at every level doesn't mean that it's a good idea to diverge from it.
2. What about the networking? the DNS magic? the storage? the filesystems? the lifecycle of data across containers & images and containers & further containers? the log management? the logging drivers? It would take multiple books to cover these topics.
3. Again the filesystem and storage issue should cover an entire book. There are many blog posts and issues talking about that. ZFS only became available very recently and exclusively to Ubuntu, it's ridiculous to consider that as a real world scenario.
Docker equals stateleness. That's the only thing it's supposed to do and could do well. Maybe you should consider focusing on one use case that Docker does well (i.e. packaging & deploying stateless applications). That would make up for better documentations and explanations and goals ;)
(IMO. After reading your comments, it seems that you have no clue whatsoever about systems internals [or maybe we just don't communicate well on that]. That's scary if Docker itself doesn't have a clue about what it is nor what it should be.)
It's not a very good fit for a production database if "performance may not be great"?
> (not that I'm advocating for writing directly to the container FS... not at all, really).
> There is nothing about Docker and statelessness.
You just recommended against storing state in the container FS on the previous line. What kind of state are you advocating a container should keep (that is different from what is captured the docker file and any separate data volumes)?
> Storage is hard whether you are in a container or not. Process isolation does not affect this.
But abstraction does. Normally for a database, you'd have a mirrored set of ssds, lots of ram, spread over a couple of physical nodes. Maybe with a loadbalancer thrown in.
Or maybe you'd run your nodes as a vm, with iscsi or some other nas/das. I can't recall seeing reasonable advice on how to set up such a production system with docker (but I haven't looked all that hard!).
Last time i checked, I couldn't find any suggestions for high-performance, well-tested container storage?
Why would a container keep from using mirrored sets of SSDS, RAM, or an LB?
The absolute worst case you can set these up manually on your host and map the directories into the container.
A better scenario, the various storage systems (EMC, NetApp, Ceph, name it) out there have volume plugins integrating with Docker, Kub, etc.
How to handle storage in the container depends on your needs, just like as if it was VM or a physical machine... and ultimately the setup is in the worst of cases no different.
I use kibana to monitor logs over 60 containers. Sure, the initial dashboard is a pain to setup, but once you get going going its very useful
A max_parallel_queries setting would be great.
>> Try using "ENV PYTHONUNBUFFERED 1" in your dockerfile
I would like to share some pointers on how I and my team deals with the issues that you mentioned in this post.
1. Orchestration :-
For us, Swarm never became a choice, as we started our journey adopting Docker in early 2015. That time there was no Swarm. We resolves to Mesos and Marathon for our orchestration. Both worked out well for us in the long run. We have production systems running this setup for last few months. Swarm is mature now, and with Docker 1.12 its been made more easier to use. The good thing with Swarm is that you could avoid adding another new system in your infrastructure like Mesos, Kubernetes etc. if you have reasonably simple requirements. We found that Orchestration also established service discovery and routing capabilities for us. We use HAProxy and Mesos DNS for our routing and discovery needs. Marathon-lb project is used to allow us to reconfigure HAProxy everytime a new Docker Container is deployed by the CD Pipeline. Marathon manages our service ports across the cluster, and every new Docker container gets its own unique service port. This service port is then informed to HAProxy, and reload happens. This setup worked good for us, although we had some initial trouble. We also practice Zero downtime deployment with our stateless services using the ZDD script inside the Marathon-lb project.
2. Running out of disk space :-
This is a common problem especially with the idea of rebuilding and deploying disposable containers with the CD pipeline. In our case, we use Monit to gather system wide metrics at all times. We use a Garbage collection script that we developed in house to remove the old Docker images and Containers periodically whenever Monit detects file system usage beyond the set thresholds. We do continuous production deployments as often as we need, so this allows for our Docker image diff to be minimal. We avoid big bang releases so that the latency for docker push on the Build server and docker pull on the cluster is minimal. The Spotify Docker-GC project is a good choice according to me.
3. Docker registry :-
We use Docker Registry container that runs on the Mesos cluster via Marathon. The Docker Registry is backed by a shared volume on the Docker hosts. We share the same volume on all Docker hosts in our cluster. So, if the registry crashes on one host, Marathon is able to redeploy the Registry Container on another host which has the access to the shared registry volume. We tried moving our Registry backed to S3, but never in production. For the systems we manage, we need the Docker images in house due to compliance requirements. Therefore, we could not use Gitlab Registry or Docker hub for our production deployments. But I heard good things about Gitlab registry.
4. Logging :- We use Logspout on each Docker hosts. It forwards the logs to our managed Logstash and further to Elasticsearch service. We use Kibana for log dashboard. Logs are rolled over on each Docker container, so that we avoid storing the logs on the host for long time. However, any distributed logging introduces log ordering and latency issues. So, we are tackling them as of now through some optimisations.
5. Dependency and Base Images :-
We use hierarchical model of managing Base Images : One top-level Registry (Global), and isolated docker registry for each project.
Every Base image gets into our Top-level Docker registry which is curated, and the associated Dockerfile for that Base image is checked into our Git Repository. We insist using these Base Images from our registry for all projects.
Each project can then inherit the base image, and customize to the local needs of the project. We follow CI and CD for our Base Images as well. Each project gets a notification when the Base image changes. They are free to opt in or opt out. This model works for us, but many not be that interesting for smaller setups. I had written about it here:- http://thenewstack.io/bakery-foundation-container-images-mic...
6. DB and Persistence :-
We avoid running stateful services in production on Docker Container, as we have qualms about the persistence support in Docker. However, we do use Docker volumes for all purposes in non-prod environments including CI, Elasticsearch and other services. We are very interested to pursue this further with ClusterHQ and Flocker based offerings in the Docker ecosystem. I may blog about it in the coming days on this. But so far, I don't have any production experience with the Database in Docker. But I am optimistic that this will happen soon.
7. Longer build times :-
This is correct as per your assessment, but widely varies across deployments. As said earlier, we want the team to have faster build times, so we build and release as often as possible.This allows to not have to deal with build latency. We use lightweight Base images like Alpine, and prevent the use of Configuration and Package mangers like Puppet inside the container. In our base images, we prevent bloating by avoid installing irrelevant packages. This has backfired some times in production, but we have found ways to go around it most times.
Overall, I am aware of the challenges that you had, and can connect with all of them. Docker is not the panacea for all the infrastructure woes, but its certainly gives the taste to me and our team on how software development and delivery will change for good in the coming days.
A lot of the issue I see him describe in production are fixed by Kubernetes. Compose works fine for local dev orchestration but pattern doesn’t work for deployment. The ideal world would be that I can run my compose file on my cloud provider, but Swarm isn’t there yet. I have to rewrite my compose file using kubernetes configs — it’s not a 1:1 mapping but the high level connection are there if you think of Cabernets Pods as Docker Containers. He mentions orchestration across a cluster with Swarm is nasty, but it’s elegant w/ Kubernetes.
Docker Registry:
Obviously, there is no constraint permitting him to use a 3rd party service. Why to let Google Container Engine (GKE) or AWS ECR handle it for you?
Longer build times:
I think this is really where he is missing the mark. It sounds like he has a fundamental misunderstand that if you mount the source code in dev you have to do it in prod too and that you have to have 1 container. Not true: You can mount the source code in dev using compose, so you don’t have to rebuild every time you change a line. Also, I think it’s a pattern in docker to try to keep your containers as atomic units of your app architecture. It sounds like they are trying to bake all competent of their app into 1 container (app + db + service, etc). Just break them up into containers, link them up w/ compose. This architecture them translates cleanly to one of the cloud providers for production: GKE or AWS.
DB and Persistence:
Yes, I think it is very clear that containers are stateless. So, yes if you want to run a DB in a container you’d have to mount an external drive somewhere. There merits and risks of that are another discussion, but as he states it’s generally frowned upon to containerize a DB. (not completely sure why, some argument about stability of container and corruption of data…) I talk more about this here: https://news.ycombinator.com/item?id=12913198
Logging:
I think 12-factor style app containerized fit more smoothly into the docker compose style architecture. Accordingly if all you containers are logging to std out it, it’s all conveniently merged and printed out to the terminal if you run compose. Then on production, GKE handles it nicely too w/ the Logging system.
In conclusion, I think most of his problems would been avoided if he didn’t skip researching Kubernetes and if he didn’t make the mounting oversight. The other big oversight I think he did, at least he didn't mention, is that he has no deployment tool. I wouldn't be able to effectively deploy w/o a build tool like Jenkins. I talk about a lot of these issues and how to fix them here: https://news.ycombinator.com/item?id=12860519
The conclusion from these is not that Docker sucks, but YOU HAVE TO LEARN it. I agree that it's a very steep learning curve, but after the pieces come together, Docker solves quite a lot of problem and actually very useful.
Could you point me to good piece of documentation or success story of running Docker in production?
(I do use Docker, but it requires to learn Kubernetes or DCOS to do anything production-related. And those are separate projects).
The documentation just plain sucks. The doc parts that don't suck are outdated, and therefore useless.
As long as that is the case, you are going to have people who go through the exact same troubles with Docker time and time again (one theme is databases in docker).
In essence, it is impossible to work with an existing Docker-managed host from another computer. I often work from home, and when I wanted to manage my Docker hosts from my laptop, it turned out to be impossible and this issue was closed without resolution.
This was the day when I gave up on using native Docker.
Nothing of it is documented, of course.
I know this doesn't help with your current issue on docker-machine, but... I think the issue here is docker-machine's primary intent is as a developer tool to spin up dev environments quickly and easily (zero-to-docker as we say), as such the datastore and security model is tailored to this. We are working on production-level infrastructure management, the base-layer of which you can find here: https://github.com/docker/infrakit
(Or, how else should I provision Docker on multiple non-AWS hosts at once?)
I agree with the article that many of the official docs around use-cases (non-reference material) is outdated and difficult to keep up to date.