Running 1000 containers in Docker Swarm
blog.codeship.com
blog.codeship.com
(Disclaimer: I'm on the Nomad team but wasn't at the time of the post)
Disclosure: by coincidence of market forces, we're mortal enemies. Let's send christmas cards!
Yup! Repo could definitely be clearer, but here's the code:
https://github.com/hashicorp/c1m/blob/master/schedbench/test...
Basically calls an increment in Redis and then blocks forever.
> Disclosure: by coincidence of market forces, we're mortal enemies. Let's send christmas cards!
Haha, hi mortal enemy! Christmas cards it is! If you're ever in Portland, OR I'll buy a beverage of your choice as well. :)
1.55k nodes, 250k containersed applications[0].
Mind you, it's hard to compare these as there's no real "cloud bench". For pure benchmark porn Nomad are the undisputed champs on their 1 million case.
The Cloud Foundry scaling test was intended to show a system with fully service-configured, fully-routed apps, with varying app characteristics (memory and RPS). To further stress the system, thousands of apps crashing and are relaunched on a continuous basis.
Cloud Foundry installations with >10k containers have been ordinary for a while now; the 250k thing was to ensure we had lots of headroom and shake out chokepoints in Diego.
[0] https://content.pivotal.io/blog/250k-containers-in-productio...
Disclosure: I work for Pivotal, the majority donor of engineering on Cloud Foundry.
Your apps will still be containerised, distributed and wired up the same way.
There's always a point at which it makes engineering sense to flip the switch to doing it yourself. But that frontier is never static. We (plus our peers in the Cloud Foundry Foundation) and others in this space like Red Hat OpenShift are constantly pushing back the tipping point at which it makes economic sense to DIY.
We already have very large customers with very large engineering teams, who've built platforms before. And they are switching because that effort no longer makes business sense. It's an expense they don't need for a platform they're the only maintainers of.
One of our peers at IBM wrote about DIY[0]. We have our own much more markety-businessy whitepaper, with a very detailed case, on the same topic[1].
[0] https://hackernoon.com/stop-spending-engineering-effort-solv...
[1] https://content.pivotal.io/white-papers/the-upside-down-econ...
Disclosure: I work for Pivotal, etc.
It will setup your swarm, which uses auto scaling groups for the worker nodes. You can then configure the auto scaling groups how ever you want, to scale based on your cloudwatch metrics, etc.
There is also a Docker for GCP product in beta. https://beta.docker.com but I don't know how auto scaling works for it.
Disclaimer: I work at Docker on the Docker for AWS product.
I've tried most of the Docker orchestration offerings and Container Engine seems by far the nicest. Swarm and Compose are really simple for getting up and running, but when we evaluated them there was still a missing piece required in that there was no neat way to do zero downtime deployments.
There's a tool called Kompose to convert docker-compose config to kubernetes manifests (https://github.com/kubernetes-incubator/kompose) although whilst it's nice to get you started we tend to maintain them separately now.
For instance/node level autoscaling (which is closer to what you need), I would recommend using the autoscaling features provided by AWS/Google Cloud.
It would have to be integrated with Kubernetes though -- when we push a new docker container, the container would need to be updated on any new machines created. We'll look into GCP's autoscale solution.
Even if you don't need autoscaling, I'd suggest still using autoscaling groups and setting it to a fixed number of instances, so that instances will automatically get restarted if they go down.
As for image management, it would depend on how you would like to propagate new images. With a private docker registry, you could potentially point each new instance to the registry and take care of propagating new images. I favor this approach since it keeps everything separate and easier to manage.
Theres an open issue (made ~2 years ago) on GH for the 1st example and it still hasn't been resolved.
[1] https://docs.docker.com/engine/swarm/
[2] https://github.com/docker/swarm
[3] Difference between Docker Swarm and Swarm mode: http://stackoverflow.com/questions/40039031/what-is-the-diff...
Why is worrying about a single web server going down more worrisome than some part of the Docker stack going down and causing the same issue?
They are part of the same app.
They should not have the same level of privilege.
The secrets in endpoint A's memory should not be visible to endpoint B and vice versa.
Containers increase assurance that this is so.
But if a single process has the single account on the database, how do you partition those permissions? Simply providing multiple logins won't help if you assume hostile code is in your process space.
On the other hand, if each service has its own login, then the database can enforce lowest authority for each. A compromise of one service isn't a game over scenario.
It's the difference between having a single account with the union of all permissions, or disjoint sets.