1000 nodes and beyond: updates to Kubernetes performance and scalability
blog.kubernetes.io
blog.kubernetes.io
Upgrading to 1.2 has been incredible. Deployments are faster and pods get scheduled almost instantly now. Our Master nodes are down to about 1/4th of what they were normally doing in terms of CPU usage.
We're really excited to be ridin' on kubelets!
Our team is looking at using it, but we haven't found a great way to do automated deployments with our current build system (Bamboo). The best we've come up with is a series of bash scripts as the deployment step, but I'm not fully comfortable with how that would handle failed deployments yet. Basically, we need a way to handle automated deployments, and see the status of our currently deployed systems / promote environments.
If anyone is using Kubernetes in production, I'd love to hear what your deployment process looks like.
Please feel free to email me (aronchick (at) google) if you'd like to discuss this further (either P or GP post). We've seen a lot of this, and would love to help you out!
Disclaimer: I work on Compute Engine and chat with the Kubernetes folks a lot.
Can you say more? Did you just spin up too many nodes?
I was looking at the Kubernetes tutorials and couldn't even start to figure out, how much would it cost to run them. (Well, I didn't try too hard, it wasn't that important.)
Drop state=INVALID
Allow state=EATABLISHED,RELATED
Drop all
You're now just as secure as if you had a typical NAT setup, but without the decades of kludges that is NAT.Disclaimer: I co-founded Kubernetes and help to coordinate the k8s Scalability SIG, although I'm no longer at Google. I didn't see this before it was published, though.
Unfortunately I would be running it on AWS and HA still hasn't been worked out and manual setup is a bear.
Slides here: https://speakerdeck.com/philips/pushing-kubernetes-forward?s...
Cluster Federation (a form of HA) is coming in 1.3.
What is not yet in 1.2, but is planned for 1.3, is HA Master - so that failure of the zone which contains your master won't interrupt the control plane. (i.e. you will be able to update your apps even as zones are failing).
No doubt protobufs are probably much more battle tested in google scale environments, but are there any other clear benefits if the goal is to reduce spending cpu time encoding/decoding messages?
Especially in SOA deployments where many small services need to communicate with one another, I would think that the ability to quickly read any field from a message and pass it on (without first having to decode the entire message) would be a very desirable trait.
Even though FlatBuffers is technically from Google, it's from a sub-team of Android working on tools aimed at Android games. The idea was that you'd store your assets in this format. IIRC the initial release didn't do bounds checking so was totally vulnerable to malicious input (but it wasn't intended for such use cases anyhow). I doubt it is widely used on Google's servers.
Cap'n Proto is not from Google and there's simply no way they'd choose to use it. To be fair, its support for languages other than C++ remains weak, largely because Sandstorm.io doesn't currently have the resources to build it out.
FWIW the ability to read a single field from a message is less important in networking situations because sending/receiving the message is already O(n) and the messages are small-ish, so parsing in O(n) is not a huge deal. Random-access parsing really shines when the input is a massive file on disk.
(I'm the author of Cap'n Proto and also of Protobuf v2 (the first version Google open sourced).)
The latency referred to in the demo is the latency of requests from the loadbots to the nginx containers running in the cluster.
Do you mean GKE?
Google Compute Engine itself was difficult to name. There were those that were pushing for Google Compute Cluster. But I veto'd as the TLA would have been GCC or GC2. Both would have been awful.
Naming is hard.
I tried to read HA documentation on Kubernetes and it all starts with warnings like "this is fairly advanced stuff, requiring intimate knowledge of Kubernetes inner workings", and going on with pages and pages of setup process.
Basic HA is not a "fairly advanced stuff", it is a commonplace requirement in any production environment. Why do I need a 1000-node cluster if all 1000 nodes are in the same AZ, which can have an outage anytime?
This is a better explanation than I would write on this:
https://www.quora.com/What-is-the-difference-between-a-highl...
I still wonder what is the primary use case for Kubernetes (or Docker Swarm, which has similar issues) if high availability is so low on the priority list.
I seem to remember the article noting that 99.9995 was more useful.