Kubernetes – The Hard Way
github.com
github.com
It's amazing to see how many large Enterprises already selected OpenShift, either in various cloud setups or on premise scenarios.
Looks like that for many HN people there is only AWS and GCE out there.
https://docs.openshift.org/latest/getting_started/administra...
After decades of replacing RHEL where possible (and plenty agony where it wasn't possible) I can not help but chuckle at the idea of even considering the company that brought us 'yum' for a virtualization platform.
For many of us, RHEL (or CentOS) is the only distribution we would consider running at scale because we feel they make better choices and have better policies and documentation than the competition.
I think moe means "considering (the company that brought us 'yum') for a virtualization platform".
The debian stable kernel has been suffering for many months from critical bugs that results in docker crashing randomly.
Note: kernel crash = unresponsive stale system, only fixable with a hard reboot
Given the title "The Hard Way", I kind of expected it would dive deeper into this topic.
Seriously though, it is getting a bit silly how almost every single kubernetes demo involving persistent data just uses emptyDir or hostPath and we are all supposed to pretend that those are a production-ready solution that won't lose data. AFAIK the PersistentVolumeClaim and StorageClass stuff will be the solution for this?
But non of these are part of your K8s tutorial :)
Gluster is completely in userspace. How would stability issues, if any, in userspace cause kernel panics? Can you comment more about that?
Tried using it for a while (last November until ~April) on Debian Jessie, with latest versions. Lots of bugs with SSL mode, lots of desynchronization issues, crippling performance issues, …
And no matter what I ran into, the bugs were already known and filed for months at that point, with zero developer reaction.
> Gluster is completely in userspace.
Unless you try using its NFS server to get not entirely disastrous performance.
> Unless you try using its NFS server to get not entirely disastrous performance.
Wrong again, Gluster's NFS server is in userspace too. Not sure how that can cause a kernel panic.
Sharing config files. Wordpress sites. Django sites. The performance was too shitty for everything.
> How did you reach out to the developers? Usually the mailing lists are very responsive and most bugs brought up on the lists are addressed.
IRC. I ended up finding all my errors unresolved by it in the mailing list archives with zero developer attention, and since I didn't have the time to help RedHat make their product actually usable, I dropped Gluster.
> Not sure how that can cause a kernel panic.
I guess it's technically not a kernel panic if I/O just stalls so hard that you can neither mount nor unmount nor otherwise access any of your NFS targets… but you have to hard reset the whole machine anyway.
I recall reading that EBS is being [partially] implemented and tested. (Read: The documentation is full of "known bug: ..." and "please don't try that in production").
On that note, are anyone using encrypted NFSv4 in production? On paper it ticks most of the right boxes, but I'm not convinced it actually works...
http://kubernetes.io/docs/user-guide/persistent-volumes/#par...
as well as the "volume management" section:
http://kubernetes.io/docs/user-guide/volumes/#types-of-volum...
Those two should give a good overview of the available options. Basically, you can use local paths, NFS shares, vendor-specific storage (e.g. AWS EBS) and various other things.
Yeah, but which of them actually work?
And on top of that, which of them work when scaling the app out. In other words: how does each option deal with having multiple containers mount it.
edit:
From this link -- http://kubernetes.io/docs/user-guide/volumes/#types-of-volum... -- I found that the following volumes allow mounting by multiple writers: nfs, glusterfs and cephfs.
But which one of these works well under (production) load?
Everything should work, but it's harder to validate / work around every NFS server's possible set of configuration options (for example).
The other solution is object storage, which is the real way forward (IMHO). This route make a app more adherent to the 12-factor principles.
Histpath and emptydir aren't really acceptable as solutions since they add complexity to your cluster set up.
There are some cluster storage technoloies out there but there are no tutorials, or overviews detailing, the pros and cons, and their performance limitations in a k8s environment.
Networking was the hot topic this time last year but storage is the topic no-one is willing to talk about. Tectonic is attempting to build a solution designed for the container use-case[1]. But that's a long way off.
If you are building cloud-native applications on-prem and you need state you're on your own for now.
[1] https://coreos.com/blog/torus-distributed-storage-by-coreos....
Same conclusion I've come to.
I guess K8s is just one part of the solution, it needs to be paired with a storage solution to be really able to replace existing infra. When looking at open source solutions the landscape is pretty empty with Gluster, Ceph and FreeNAS. Where only Gluster and Ceph provide some level of HA.
Please see author's tweet here - https://twitter.com/kelseyhightower/status/77498377095689830...
The github repo is updated with updated DNS add-on, better examples and AWS support.
My advice: You should be using kops or kube-aws for production. (I work on kops, so I am biased to believe it is important)
Never tried it though, so I can't speak of the quality (and of course it talks about running it on CoreOS).
In some cases you'll find side-by-side labs that focus solely on AWS like this one for bootstrapping the underlying compute: https://github.com/kelseyhightower/kubernetes-the-hard-way/b...
In other parts of the tutorial you'll find sections for both GCE and AWS that highlight the different commands required for each platform - only a small fraction of the tutorial requires something different between GCE and AWS. Both cloud providers are now treated equally in the updated tutorial.
https://github.com/kelseyhightower/kubernetes-the-hard-way/b...
Apparently, there's something odd in the GitHub support for relative links in the main project README.md.
If you don't mind filing a few issues on the project I'll be happy to rework each lab to start addressing these concerns. Until then I've added a note to the README regarding production readiness, but that should not prevent people from learning.
However, the long term goal of sig-cluster-lifecycle is to bridge the gap here, to make a production configuration easy and obvious. So I'd love to see you start using the new kubeadm work, for example, so that we can meet in the middle!
I would also like to note that I'm not in competition with those automation tools. I'm just really focused on helping people learn. Some parts of the Kubernetes community really want to learn this stuff without an automation tool so they can skill up to build their own, custom, tools that match their preferences and tradeoffs.