In general enterprises don't like to be told "throw away your existing system to adopt my platform".
- Setting up Swarm is trivial, setting up Kubernetes is not so much (kubeadm is still not recommended for production, and there are valid reasons for this). I guess the only thing that's probably easier on K8s is (totally unsupported) multiarch cluster - it's somewhat messy with Swarm[1]. Although my experiments with multiarch K8s had failed (got issues with CNI stuff and postponed research for later).
- Compose file format is significantly simpler and concise. Less code to write is good.
- Debugging failing Swarm is significantly easier than debugging failing Kubernetes. Well, that's probably subjective and I haven't truly deeply debugged either, but at least I believe so - on occasions I was able to find my way through moby, swarmkit & libnetwork source, and K8s feels a very different beast.
____
Update[1]: I re-checked multiarch status for Swarm and found out that now `docker service create --name test --placement-pref 'spread=node.id' --replicas 2 --no-resolve-image ubuntu sh -c 'while true; do uname -a; sleep 1; done'` just works on a freshy set up mixed x86_64+armhf Swarm cluster (no multiarch alpine images yet, though - promised to come next week). Guess, Swarm had beaten K8s here.
I work with a team of four total developers and realistically two of us handle the vast majority of "operations." As a result, what was important for us in orchestration was ease of setup, speed of initial implementation, and the lowest immediate and ongoing difficulty associated with whichever orchestration tooling we chose.
My company is currently fairly locked into another DO for misc. reasons, and as a result would be deploying to DO, where you do not have the benefit of all of the automated tooling/management provided by GKE.
What people usually don't realize is that once you're in the cloud, all your services talk TCP. A service in GKE and another one in AWS are just a network hop away. Two considerations are:
* Network egress costs, which is currently outrageously high, and serves as a lock-in device of sorts. Depending on actual workloads, it may or maynot make sense for you.
* Security. Though there are numerous VPC solutions out there, some of them supported by the cloud providers themselves.
Anecdotically, we also run a small RDS database in AWS, though we only need ~100 small queries a day.
Unfortunately our egress costs would likely be significant.
I am honestly anxious to get my employer started on the path to k8s, but until the tooling reduces the hours required to maintain it successfully, Swarm seems like the superior solution if you are locked in on a non-GKE/Azure Container Service provider.
I have heard good things about Kops helping setup/enable production-stable orchestration, but they are not ready yet for Digital Ocean either, although I think it is on the list™.
Disclosure: I work for Pivotal.
With GKE, most of your focus will be around (1) tooling for generating manifests, such as Helm or (forgot the name of it). I had written something called Matsuri, but that is only useful if your team is a Ruby shop. (2) what to put into the containers and how they link up.
So yeah, I agree with your assessment.
A big part of that is because I have had experience running Docker (not Swarm), ECS, K8S, and building developer tooling (Vagrant), in addition to being a regular developer.
The flip side: I had to opportunity to try out these different tech in production and saw where the pain points are. The overhead of K8S exists to solve those pain points, though that is probably not that obvious to a small team without prior knowledge.
For example, I had set up a prod k8s by hand. I will never do that again. On the other hand, I know roughly what is going on when something breaks in our Google GKE cluster.
https://rocketeer.be/blog/2015/11/kubernetes-from-the-ground...
And although I never ran through Hightower's Kubernetes the Hard Way, it is like that. https://github.com/kelseyhightower/kubernetes-the-hard-way
After running through that as a kind of kata, it was easier to infer and troubleshoot things when things go wrong. The transfer-of-learning happens only if you run yourself through these exercises.
I can share some things at a higher level though:
Label selectors are your friend. Master them. They are used everywhere.
Stateless is still easier than stateful. Start with putting stateless workloads in production before ever trying stateful.
If you have the expertise to mix your stateful pods with your stateless pods, make sure you master StatefulSet and things like persistant volume claims.
If you fake stateful pods like I did in production, then Kubernetes does not know how to cleanly shut them down. Automated maintenance involving kubectl cordon and drain no longer function well. You end up having to hand migrate stateful pods from node to node.
(edit: or maybe I'm thinking of CRI-O)
Looking at moving to k8s but that was my reason for swarm at the time. And it was the right one as it just worked with minimal effort.