Cloudflare Uses HashiCorp Nomad (2020)
blog.cloudflare.com
blog.cloudflare.com
I recently looked closer into my billing data for GKE (Google managed Kubernetes) and was astounded by the overhead taken up just by the Kubernetes internals. Something like 10-20%. It might be better if using bigger node types.
how does Nomad compares on this front?
You generally want to pair it with Consul (as Cloudflare has done) to get service discovery. Consul is also quite lightweight.
GCP makes that conveniently pretty easy with their pricing model for CPU and memory, so it is possible to do this incrementally.
A consensus protocol has O(logn) behavior on any network that displays any of the 8 fallacies of distributed computing. But the larger the cluster the more fallacies you're likely to have to deal with in a given day, and O(logn) is too optimistic. If the cost is ever 'marginal', it's in little islands of stability that won't last long.
What I see over and over again is people expending huge opportunity costs trying to keep their brittle system in one of these local maxima as long as they can. I think because they fear that once they slip out of that comfy spot people will see they aren't some miracle worker, they're just slightly above average and really good at story telling.
What I'm talking about, and I think the OP I replied to is referring to, is the monetary and compute cost in CPU and memory overhead of the kubelet and it's associated processes on real-world deployments. There are plenty of other costs associated with Kubernetes, of course.
GKE charges a fixed cost to operate the consensus protocol (etcd) and control plane (kube-apiserver) on their own systems.
On the nodes the user operates, the costs then are relatively fixed, that is, there is some amount of CPU time and memory spent per node on logging, metrics, and so on, but as the node gets larger, that quantity does not increase. (It's definitely sublinear.)
Or in other words: you can more greatly increase the capacity of a GKE cluster by doubling the size (CPU/memory) than doubling the number of nodes. Scaling up will increase gross capacity by 100% while increasing overhead by a few percent at most, scaling out will have the same increase in gross capacity, but also increase the amount spent on node overhead by 100%.
We regularly call out pharmaceutical and petrochemical companies here for doing that. I don’t know why you would expect tech to get a free pass.
The most important thing is that you don’t fool yourself, and you are the easiest person to fool.
You do have to know about worst case behaviors.
There are no circumstances where best case behavior is worth even calculating for anything other than curiousity, unless you're trying to deny said behavior to a cryptographic adversary (so it's actually your worst case).
I totally understand there are use cases for which k8s is a great fit because it has so many capabilities, but I often see folk on HN advocating for using it everywhere, because it's "so simple" once you understand it. I just don't get it.
You see, there used to be no such "fixed overhead" in early versions (including post 1.0) of k8s.
This turned out to be a worse idea for less involved or experienced operators than fixed overhead, because it turned out that people would run k8s nodes way underpowered, load them with a ton of workload, then have lots of outages as they starved critical system components of resources.
Because of that a (settable) "overhead minimum" was added to calculation done about available resources, iirc originally going for 0.9 cpu core and I don't recall how much memory. This allowed to still run a 2 core node (though it's really not recommended) for experimentation, while greatly lowering chances of obscure for newbies issues. It didn't prevent them (PSA: If another team provisions your cluster, check what resources they provisioned...) but it makes it much harder to fail.
On a 16GB machine that's not a big deal. On a 2GB machine that's a significant fraction of your disk cache, for a very slow disk. On a 1GB machine that's just make or break.
ETA: And you may think that these toys don't matter, but people have to learn and buy-in somewhere, and the fact of the matter is that I have production services right now that mostly need bandwidth, not CPU or memory, and once I tune them to the most appropriate AWS instance for that workload, the Kubernetes overhead would become the plurality of resource usage on these boxes (if I were using k8s, which we are not at present). K8s only scales by getting into bin packing, where this latency-sensitive service runs with substantially more potentially noisy neighbors.
[0]: https://cloud.google.com/kubernetes-engine/pricing#cluster_m...
these are full-fledged, certified k8s distributions that run on raspberry pi as well as all the way in production.
https://www.youtube.com/results?search_query=raspberry+pi+k3...
(Disclaimer: engineer at Cloudflare, and author of blog post linked below.)
[1]: https://blog.cloudflare.com/high-availability-load-balancers...
I’ve looked into kubernetes but after I got it all going I decided it’s not worth the complexity. Same experience as installing gentoo :-)
I think Borg was DistBelief, Tensorflow is Kubertenes and I’m waiting for Keras. Helm isn’t there yet IMHO so I’m still waiting.
Yes, they had their own system before, but they replaced it with somethign else. However, it turns that I read this too long ago and mistook DC/OS for Nomad. My apologies!
I tried using Kubernetes, but that was a lot of work for a home server, and so I went back to Swarm. Is Nomad appreciably easier to set up?