Docker in Production: An Update
thehftguy.com
thehftguy.com
You can just SSH into one and see for yourself.
Kubernetes was built from the ground up to orchestrate Docker. CoreOS did a lot of work to make it possible to trade rkt in for Docker's engine, and the cri (Container Runtime Interface) is now generalising that so that there is a clear abstraction between the kubelet and the engines it orchestrates. Read about it here: http://blog.kubernetes.io/2016/12/container-runtime-interfac...
If you want to do things that are different to what we provide support for on GKE as a Managed Service (tm), you're able to run your own Kubernetes clusters on GCE. (We do let you run a Kubernetes alpha version, but only on non-supported clusters that self-destruct after 30 days.
Honestly the hardest thing is keeping up with how fast Kubernetes evolves and gets better and better. The same goes for all Google services (pubsub, bigquery, etc)
We started the migration on Kubernetes 1.1 and are now live on 1.5.1
Even using it for things we probably shouldn't (old stateful applications) without a single problem. At least no Docker related problems.
I don't know, this article seems to be very presumptuous. A lot of bold claims, little backing and, as you state, some pretty false claims.
TBH I don't know why it's among the top of HN
GitHub's lmctfy README states "We are not actively developing lmctfy further and have moved our efforts to libcontainer."
With 'docker info' on a node you can see we use OverlayFS, which seems a popular choice in the community also: http://burkelibbey.s3.amazonaws.com/dockercon-fs.pdf
I assume customized kernel as well? and where does the overlay drivers come from? How much custom back-ports and custom development?
The article is not on point to say GKE replaced docker entirely then... but you are not point to deny and pretend that you are running Docker on anything remotely common.
CoreOS is a very similar idea, and given that, we don't make builds of the container images available outside GCE.
Customers who need specifics of the OS (that they can't find them by just looking at the kernel config on the node) are welcome to ping us a support ticket.
> CoreOS is an operating [system] that can only run Docker and is exclusively intended to run Docker.
No, it isn't. (Maybe if you substitute "Docker" with "containers.")
> First, the main benefit of Docker is to unify dev and production. Having a separate OS in production only for containers totally ruins this point.
No, it doesn't. Completely the opposite, in fact: Docker makes it possible to not care about OS differences between dev, test, and prod.
> Docker on Debian is major no-go
> Docker is 100% guaranteed suicide on Debian 8 and it’s been since the inception of Docker a few years ago
> Debian froze the kernel to a version that doesn’t support anything Docker needs and the few components that are present are rigged with bugs.
rubs temples It's Linux. It's Debian. You can run any kernel you want. There are plenty of repositories out there with binary kernels, including Debian's own back-ported 4.9 kernel, or you can build one from source. Because, you know, open source. I've run many instances of Jessie on Linode with their 4.8 and 4.9 kernels and it's no problem. Hundreds of companies are running Docker in production on Debian and Debian-derived systems.
> I am not aware of any serious companies than run on Ubuntu.
Just because you are not aware does not mean they do not exist. How about Netflix, Snapchat, Dropbox, Uber, and Tesla for starters?
> I cannot comment on the LTS 16 as I do not use it.
It's been out since last April. If you're serious about this survey, this is a no-brainer.
> I received quite a few comments and unfriendly emails of people saying to “just” use the latest Ubuntu beta
What? No. Just use the latest LTS.
> I am moderately confident that there is no-one on the planet using Docker seriously AND successfully AND without major hassle.
LOL
The guy works in finance, not an area where you can typically say "oh, no worries, we'll just push out this distro with a non-standard kernel." Enterprise IT, and specifically finance (and healthcare) have different requirements. Currently working with a client where not only the full technology stack has to be certified by accountable vendors, but all the internal documentation and customer facing content has to pass strict legal reviews.
"Here is Debian with a custom kernel" won't fly. Nether will "Here is Debian" - Ubuntu LTS will barely pass the line, and that will be after a whole lot of timewasting and ass-covering from many people involved.
Absolutely right.
> Docker Issue: Breaking changes and regressions
Yes, it's new, and it changes regularly. This is why you do what you do with any other piece of software in an enterprise environment: Pick a stable release (e.g. 1.7), deploy it, then upgrade your development environment to the next stable release (e.g. 1.11) and work through the breaking changes. Regressions are frustrating, I'll agree, but it's free software: report the regressions and help the community fix them.
> Docker Issue: Can’t clean old images
A built-in feature to do so was added in 1.13 (a few months after this article was published).
> The only way to clean space is to run this hack, preferably in cron every day: docker images -q -a | xargs --no-run-if-empty docker rmi
That's not a hack, that's how you do things in Unix.
> As a long-standing goal, the AUFS filesystem was finally dropped in kernel version 4.
> There is no unofficial patch to support it, there is no optional module, there is no backport whatsoever, nothing. AUFS is entirely gone.
While the first point is technically true (actually, based on some light googling, I'm not sure it was ever merged in the first place), many distributions provide it as an optional kernel module. For example, Ubuntu provides it in linux-image-extra. Yes, Virginia, you can build kernel modules from source.
> How does docker work without AUFS then? Well, it doesn’t.
It does: Btrfs, Device Mapper, ZFS...
> So, the docker guys wrote a new filesystem, called overlay. [..] Note that it’s not backported to existing distributions. Docker never cared about [backward] compatibility.
Docker supports multiple storage drivers, one benefit of which is to be able to support older systems: AUFS on older distributions, OverlayFS on newer. The container abstraction allows you to not care about the underlying storage subsystem.
> Right now. We don’t know of ANY combination that is stable
In November 2016? Are you kidding?
0. https://thehftguy.com/2016/11/01/docker-in-production-an-his...
There were two major obstacles; the file system driver and the Docker daemon. I ended up settling on 1.11.{i forget}, because virtually every other version was unusable. As sad as it may sounds, I was tracking two metrics; probability the container will start and time until the docker daemon reaches a state of irrecoverable dead-lock.
Personally I found containers (on a highly contentious system) launched at ~98% success rate with a time-to-dead-lock somewhere around 6500 container start/stop (it was a pretty fat tailed distribution). For a busy system... those numbers equate to an administrative headache.
And on Ubuntu/Debian the defaults (the thing the majority of non-specialists are probably using) were far worse. The only filesystem driver which seemed to work at all was devicemapper direct-lvm.
-
If you have developers on your team who can spend their time figuring out the one magic incantation that makes Docker work most of the time, you're fine. Everybody else should follow his advice for using service providers.
Admittedly, the latter strategy is a bit frustrating and counterintuitive on 16.04. When you install Docker it will start automatically and hang because it can't find the AUFS driver and Device Mapper isn't set up. The solution is to either 1. modify policy-rc.d to prevent services from automatically starting or 2. set up Device Mapper (install dmsetup and run `dmsetup mknodes`) before installing Docker and changing the storage driver. Unfortunately, these workarounds are not particularly well-documented.
I would be interested to see how Docker 1.12 or 1.13 with a modern kernel and OverlayFS would be able to handle your described workload.
FYI: It's because of this sort of bullshit that there are articles called "Docker in Production: An history of failure".
And there goes my hope of ever running Docker on Ubuntu with any success.
I no longer work with that company and so I can't tell how those improvements would change the reliability. Presently I am using docker with btrfs on a low-throughput workload and it seems to work just fine.
Tangentially related - I have been using OverlayFS without Docker for about a year now. It's pretty great. My read-layers can be on an NFS drive and the penalty for writes is relatively small. It basically gives me 90% of what I wanted from Docker in the first place.
wget -q https://raw.githubusercontent.com/docker/docker/master/contr... -O - | bash
There's just not a ton of substantiated content here. It's mostly really lousy anecdata like:
> Sadly, I am not aware of any serious companies than run on Ubuntu.
I am troubled by the direction of Docker and there are serious issues, some of which were raised in this post. However, whatever signal there is in this post is lost in a sea of noisy ranting.
My advice to anyone who doesn't already have strong opinions:
* Complement any reading that you do with your own research.
* Don't try to invent your own container orchestration system.
* If you aren't sure how to best do something, ask someone!
And for the love of everything holy:
* Don't run stateful systems on Docker if you can't handle failure or data loss!
Docker and orchestration systems like Kubernetes can be an excellent pairing. It's going to require research, a change in how you develop, build, test, and deploy systems, and a gradual building of operational experience. It will not be a quick process, and it's not for every org. But for some orgs and usage cases, it's an excellent way to go!
This is a patently false statement and needs to be revised, especially given:
> I will not comment on it.
P.S. I'm getting downvoted for speaking the truth, no matter how pedantic it may seem. CoreOS can run rkt containers and is not exclusively intended to run Docker. One possible reason is because Docker is the 800 lb. gorilla in the room and having options is always good.
P.P.S. Gotta love the Streisand Effect. :)
This shows an astonishing level of ignorance for someone who claims to have done their research.
> First, the main benefit of Docker is to unify dev and production. Having a separate OS in production only for containers totally ruins this point.
What? This makes no sense. Your images will be the same between dev and prod, even if the host running the containers is different - which is really the whole point - if you build and run an image in dev, it should run identically in prod.
> If you like playing with fire, it looks like that’s the OS of choice.
We spent the last year+ running containers on Centos7 with no problems from the OS. Whatever issues we did encounter were either transient bugs with Docker or our own configuration. Perhaps we got super lucky, but we were running 120+ containers on 12 hosts, so I would've expected at least some evidence of significant problems within that timeframe if it were really such a risky setup.
> It’s not possible to build a stable product on a broken core, yet both Pivotal and RedHat are trying.
We've been running OpenShift Origin since March of last year, it's been very stable during that time - the few issues we did encounter were due to our own mistakes, and were usually fixed just by changing some configuration and restarting the host.
While there are undoubtedly problems with Docker, and likely many of the issues you brought up are very real, there are many teams like mine that use it successfully, and painlessly. Docker isn't the tire fire you want to make it out to be.
I think the author is saying that while this is the point, it's not the reality. As it is right now, things can run absolutely fine in dev and then due to prod running a different operating system, things break.
it is virtually irrelevant what you use to run and operate this sort of cluster. it gets tricky when you pass 100 nodes and even trickier when you get to 1000+ nodes in a non-linear fashion. Docker certainly has issues that are pretty severe when you are in the financial space and not a big deal if you are running a popular blog like medium.com for example.
Let's multiply that a few times for dev, test and support systems. Still, no need to have hundreds of nodes.
I use Docker at work, and I 1) don't have any loyalty to their brand and 2) try as hard as I can to abstract away their specific APIs.
For example, we use Convox [0] to deploy containers to AWS. I could care less what Convox and/or AWS are doing under the hood. They could switch out Docker for rkt under my feet, and I probably wouldn't even notice.
It is kind of like POSIX to me. My apps are designed to run in a POSIX environment, not specifically CentOS or Debian. And just like it's easy for a new Linux distro to come along, give me POSIX, and give me some other shiny features I like, it will be easy for any competitor to come along and replace the Docker interfaces I use.
Though the tone of this article is very negative, the conclusion is interesting to me as someone that's considered Docker but hasn't done much with it beyond experimenting locally.
I'm assuming if there's any forum that has users that can speak to their use of Docker in production, this would be the place.
We went through the gammut of swarm etc. and were never able to get any of them to work in a meaningful way.
We've been using Kubernetes to deploy Docker containers on Google Container Engine, and while there have been a few issues due to Docker/Kubernetes (namely, getting the containers to expose localhost to each other, and to expose themselves to the world), the issues that we've had so far have been issues that we would have been bit by eventually. Namely, if we hadn't been forced to deal with the issues now, we would have been screwed later. There have been some weird bugs due to the internal environment that containers use (we had extremely slow DNS lookups that caused our request times to shoot up to 9s each). These issues have been transient though, so it's not clear that it's Dockers fault, or if we're making mistakes in our code.
Docker has made our deployment much easier. You just build & push your container, and you instantly have a versioned deployable instance of your code. Kubernetes makes it extremely easy to rollout or rollback containers, so I have nothing but good things to say about containers.
A quick skim would suggests that sits on about 8TB of Ram and 1000vCPUs spread over 100 machines. It's not the biggest stack in the world, but it's doing a good job as our little internal PaaS.
We have some non-aws versions too, but I don't have metrics on those right now.
> Google offers containers as a service, but more importantly, as confirmed by internal sources, their offering is 100% NOT Dockerized.
> Google merely exposes a Docker interface, all the containers are run on internal google containerization technologies, that cannot possibly suffer from all the Docker implementation flaws.
IMO two of the few systems that were sound were Mesos and Kubernetes (and their roots are somehow interleaved) and neither puts the Docker as the central piece, nor the container, which is just a building block to achieve the actual goals of running distributed, highly available jobs and services efficiently in a shared environment.
FWIW, there was drama with the initial rkt announcement, but we've actually seen docker adopt most of the original criticisms and that should be applauded for both CoreOS and Docker.
Google running their own containerization tech with a Docker interface seems a bit far-fetched given the level of integration of Kubernetes with Docker. That's totally possible though, I'd like to read more about it.
It's not true.
Second part TLDR - "If you are locked-in on Docker and running on AWS, your only salvation might be to let AWS handles it for you.". Use Docker-For-AWS that sets up Cloudformation + Swarm mode.
Google offers containers as a service, but more importantly, as confirmed by internal sources, their offering is 100% NOT Dockerized. That is a huge label of quality: Containers without docker. Actually, this is deliberate. Containerd. For example, rkt support on GKE is "coming soon" . https://www.mail-archive.com/google-containers@googlegroups....
I call B.S. on this. Amazon wouldn't have spent so much effort on ECS if this was true.
Basically stating that ECS was Amazon's attempt to plant a flag in the container-space and that it was half-hearted, not done with the rigor of solutions like k8s that had additional advantages of also being cloud-agnostic, and finally that ECS was not a solution that anyone they knew was recommending or would recommend.
> when ECS launched, the goal was to do a smaller number of things really well, and then listen to the community on what _they_ wanted to see from a container management platform, and grow accordingly
I'm coming from a small CoreOS cluster on bare metal, onto Kubernetes and Helm on EC2 nodes. I was a Fleet user before and I loved it! But always with the understanding that when things got better in the cluster space, I'd move from Fleet toward some resource aware scheduler.
I love to read how you're all going down a similar road, however you get there! Thanks for the blox links and I think I will be able to make immediate use of ECR with the rest of my AWS stack.
I went looking for how to enable ECR and I didn't find it, is this a feature you can only use from within ECS?
On my legacy CoreOS cluster I always used Deis components to (theoretically) manage all of the cluster things. Kubernetes offloads many of these concerns to the Cloud provider, and handles others of them using Addons. Can I get ECR as my private registry on a Kubernetes cluster running on EC2 nodes?
https://console.aws.amazon.com/ecs/home?region=us-east-1#/re...
You'll need to already be logged into AWS for that link to work, and you can select a different region.
As for your second question, the answer is absolutely yes. k8s supports this natively: https://kubernetes.io/docs/user-guide/images/#using-aws-ec2-...
You can also reach me at abbyfull AT amazon DOT com
First of all I can assure you that Amazon is 100% committed to containers. Amazon's compute strategy is aimed at three levels of abstraction: instances, containers, lambda. All three are equally important to Amazon.
With regard to ECS feature set relative to K8, the thing to understand is that AWS follows a startup-like strategy of launching an MVP and then letting customer feedback drive roadmap from there. AWS is definitely not half hearted about ECS. Rather AWS is constantly working with customers to define a roadmap for further development on ECS.
To me the most exciting thing about ECS is the open source work being done around ECS, for Blox (a framework for building out custom container scheduling logic), and the ECS agent itself:
https://github.com/aws/amazon-ecs-agent
With these ECS components you can open issues or even PR's just like any other open source project. Additionally another cool thing we are doing is sharing feature proposals for public comment. You can check out an Amazon employee's public fork of the ECS agent to see a preview of coming roadmap, and we are actively soliciting feedback on proposals such as this one:
https://github.com/aaithal/amazon-ecs-agent/blob/f4d80440db0...
If you have any questions about ECS, feedback, or concerns feel free to reach out to me directly (peckn@amazon.com) and I'd be happy to chat!
There is no payback for the hassle...
Running financial infra is quite the opposite. Long term reliable processes, not much space left for failures and mishaps, especially not in the lowest level of your infrastructure. What is the benefit for using Docker for this sort of services? Almost zero.
There's also FreeBSD jails that can be used to contain applications. I'm not sure about the timeline, but I think jails came first. Sun engineers wanted to achieve the same so they ported the same technology over to Solaris.
He will be speaking at Dockercon in April, so I suppose we'll find out soon enough.