Kubernetes 1.10 released
blog.kubernetes.io
blog.kubernetes.io
EDIT: I should have clarified, I want to self-host this on our internal VMWare cluster, rather than run it on GKE.
Not really. There are plenty of ways of getting a single node instance. None of them will give you a "production-ready" one, because they don't define it that way (and I happen to agree). You can of course do whatever you want.
Since you are using VMWare anyway, why can't you spin up more VMs (maybe smaller)? You can vmotion them away to different nodes when you are ready to actually make the cluster HA. It is a really really good idea to keep master and workers separate, even if you run a single node for each.
Failure of the worker will of course bring down your applications. When you recover it or spin up another one, K8s will recover your apps for you. Failure of the master will not adversely affect your systems only the cluster's ability to manage itself and its self-healing capabilities (which will affect uptime at some point).
Failure of the combined master/worker/etcd node should be recoverable, but frankly, at this point, should you care? I would just shoot it in the head and add some automation to provision a brand new cluster and deploy those applications again. Since you are not worried about HA and just want a place to deploy the containers, just make the k8s single-node-cluster cattle.
Took the day to setup minikube on a CentOS server and play around, however, I wasn't able to expose anything to the outside world. Looking into Ingress at the moment, however, documentation is a bit loose there, I think.
Another comment suggests to setup a single-node cluster and remove the taint on the master, maybe I will try that instead.
edit: any advise much appreciated!
> I don't have anything that needs to autoscale or fail over or do high availability
You should not be using Kubernetes.
That is: what is the “run your own thing-that-is-slightly-lower-level-than-Heroku on a fixed pool of local hardware resources” solution in 2018?
But Kubeadm will still do it. Kops if you're on AWS. GKE if you're on GCP. Just Docker would be easier to set up though, and that's what the OP means.
EDIT to add: It's assuming Ubuntu LTS as the node's OS, not sure if that fits your use case. Should be possible to adapt this to ContainerLinux or anything else without much trouble.
I haven't worked with GL's Auto DevOps yet, but I think the cluster should have everything necessary to get going with that.
Similar use case - self-hosted VMs, for low-traffic, internal tools, and no need for autoscaling.
I can't speak to how well it integrates with Gitlab's Auto DevOps, but Nomad integrates very well with Terraform[1] and I'd be surprised if there wasn't a way to plug Terraform into Gitlab's process.
A production-ready cluster has dedicated master(s), period. In order to get your single-node cluster to work (so you can schedule "worker" jobs on it) you're going to "remove the dedicated taint," which is signaling that this node is not reserved for "master" pods or kube-system pods only. That will mean that if you do your resource planning poorly with limits and requests, you will easily be able to swamp your "production" cluster and put it underwater, until a reboot.
(The default configuration of a master will ensure that worker pods don't get scheduled there, which makes it harder to accidentally swamp your cluster and break the Kube API, but also won't do anything but basic Kubernetes API stuff.)
If things go south, you're going to be running `kubeadm reset` and `kubeadm init` again because it's 100% faster than any kind of debugging you might try to do, and you're losing money while you try to figure it out. That's not a production HA disaster readiness or recovery plan.
But it 100% works. Practice it well. Jenkins with the kubernetes-plugin is awesome, and if I have a backup copy of the configuration volume and its contents, I can start from scratch and be back to exactly where I was yesterday in about 15-20 minutes of work.
My 1.5.2 cluster's SSL certificate expired a few weeks ago, on the server's birthday, and after several hours trying to reconcile the way that SSL certificate management has changed, to find the proper documentation about how to change the certificate in this ancient version, as well as making considerations that I might upgrade, and what does that mean (read: figuring out how to configure or disable RBAC, at the very least)... I conceded that it was easy to implement the "DR-plan Lite" that we had discussed, went ahead and reinstalled over the same instance "from scratch" again with v1.5.2, and got back to work in short order.
I've spoken with at least half a dozen people that said administering Jenkins servers is an immeasurable pain in the behind. I don't know if that's what you intend to do, but I can tell you that if it's a Jenkins server you want, this is the best way to do it, and you will be well prepared for the day when you decide that it really needs more worker nodes. It was easy to deploy Jenkins from the stable Helm chart.
Once you get onto 1.8+ w/ CRDs you can manage your SSL certs automatically via Jetstacks Certmanager; https://github.com/jetstack/cert-manager/tree/master/contrib...
It just hasn't been a priority. I have no need for RBAC at this point, as I am the only cluster admin, and the whole network is fairly well isolated.
I couldn't really think of a good reason to not upgrade when it came time to kubeadm init again, but then I realized I could probably save ten minutes by not upgrading, it was down, and I didn't know what the immediate consequences of adding RBAC would be for my existing Jenkins deployment and jobs.
Chances are it would have worked.
You can tell already from what little conversation we've had that "always be upgrading" is not a cultural practice here (yet.)
We have regular meetings about changing that! Had two just yesterday. Chuckle
If you make each of the pieces required parts of the whole, then yes - adding more of them will increase the chance that the whole system fails. But in kubernetes, the additional pieces (nodes) are all redundant parts of the whole, and can fail without affecting the availability of the whole system. The more nodes you add, the more redundancy you're adding, and the less chance that the system as a whole will be affected.
Mathematically:
If a component fails F% of the time, then adding N of them "in series" (all of them need to work) means your whole system fails with a (1-(1-F)^N)% chance. Iow, as N goes up, the system approaches (1-0)% => 100% chance of failure.
Otoh, if you combine the parts "in parallel", and you only need any one[1] of the components to work in order for the whole system to work, then the system has a F^N% chance of failure. As N goes up, this system approaches 0% chance of failure.
[1] Kubernetes (etcd) isn't quite this redundant, since etcd needs a majority quorum to be functional not just any single node. But the principle is similar and still gets more reliable as you add nodes.
When its done we expect to be able to scale the different parts of our application independently. This also makes it easier to detect problems (why did Gitaly suddenly autoscale to twice the number of containers).
So a lot of our effort in moving to kubernetes has required us to try and duplicate all of our previous installation work in ways that don't require root access while the containers are running.
To help, we've chosen to use Helm to manage versioning/upgrading our application in the cluster: https://helm.sh/
Journey to Cloud Native: Breaking the monolith and scaling towards tomorrow - https://docs.google.com/presentation/d/1fsgvSuGpn-MnMqKaTOoi...
Our biggest blockers have been how to separate file system dependency for our various parts. For Git content, we've implemented Gitaly (https://gitlab.com/gitlab-org/gitaly/). For various other parts, we've been implementing support for object storage across the board. Doing this while keeping up with the rapid growth of Kubernetes and Helm has been a challenge, but entirely worth the effort.
We are also bringing all of these changes to GitLab CE, as a part of our Stewardship directive (https://about.gitlab.com/stewardship/, https://about.gitlab.com/2016/01/11/being-a-good-open-source...). We don't feel that our efforts to allow for greater scalability, resiliency, and platform agnosticism belong only in the Enterprise offering, so we're actively ensuring that everyone can benefit, just as they can contribute.
We have a lot of improvements coming as well: https://about.gitlab.com/direction/#ci--cd
While it's not a single button, GitLab is getting closer to that goal with the addition of our cluster integration feature: https://docs.gitlab.com/ee/user/project/clusters/#adding-and...
If you create a trial account on GCP, you can then use GitLab to make a new GKE cluster and it will be automatically linked. You can then click a single button to deploy the Helm Tiller, a GitLab Runner, and then run your CI/CD jobs on your brand new cluster!
The GCP trial is pretty nice, you get $300 in credits and you won't be automatically billed at the end: https://cloud.google.com/free/
What made you eventually go ahead with K8s other than the main obvious reason (the community) itself?
One big reason we've chosen to implement with Kubernetes is community adoption. You can now run Kubernetes on AWS, GCP, Azure, IBM cloud, as well as a slew of other platforms both on-premise and as a PaaS. With the expanding ecosystem, our decision has only been further confirmed, as the list of available solutions and providers continues to grow rapidly (https://kubernetes.io/docs/setup/pick-right-solution/)
Another big reason is far simpler: ease of use and configuration. A permanent fixture in our work for Cloud Native GitLab has been that this method must be as easy, if not easier to provide a scalable, resilient, highly-available deployment of GitLab at scale.
We can’t say, “This is our new suggested method. By the way, it’s harder.”
What we have found is that many other solutions require a much larger initial investment of time to understand and configure GitLab as a whole solution, as compared to the combination of Kubernetes with Helm. Helm provides us templating, default value handling, component enable/disable logic, and many other extremely useful features. These allow us to provide our users with a practical, streamlined method of installation and configuration, without the need to spend countless hours reding documentation, deciding on the architecture, and making edits to YAML.
We've been looking at these technologies for two years now, with our focus being Kubernetes for the last one.
We also try to use the same deployment tools for GitLab.com that we provide to customers, and this lets us offer a scalable production-grade deployment method that can run nearly anywhere.
We also no longer have to manage servers.
All of this was possible without Kubernetes, but it is so much easier with it. Although admittedly much of the ease of use is due to not having to manage Kubernetes itself – we use Google Kubernetes Engine. I would not want to install or manage a Kuberenetes cluster/master myself.
I can recommend managed Kubernetes to anyone that runs many different apps/services.
We first deployed v1.2 nearly 2 years ago, and I can say Kubernetes has made some amazing improvements in that time – in terms of functionality, usability, scalability, and stability. This release continues that trend.
We've invested a lot in it too with things like making sure our etcd clusters backing Kubernetes are really solid [2], and we've even added some of our own features to Kubernetes like allowing CPU throttling to be more configurable to get more predictable >p99 latency from our applications.
We've been through our share of production issues with it (some of which we've posted publicly about in the hope that others can learn more about operating it too [3]), but I don't think there's any way in which we could run an infrastructure as large and complex as ours with so few people without Kubernetes. It's amazing.
[1] https://monzo.com/blog/2016/09/19/building-a-modern-bank-bac...
[2] https://monzo.com/blog/2017/11/29/very-robust-etcd/
[3] https://community.monzo.com/t/resolved-current-account-payme...
I spent a couple weeks not quite full time going through tutorials, reading the documentation, reading blog posts and searching for solutions to the problems I was having. My biggest problem was with exposing the services "to the outside world". I got a cluster up quickly and could deploy example services to it, but unless I SSH port forwarded to the cluster members I couldn't access the services. I spent a lot of time trying to get various ingress configurations working but really couldn't find anything beyond an introductory level documentation to the various options.
Kubespray and one blog post I stumbled across got me most of the way there, but at that point I had well run out of time for the proof of concept and had to get back to other work.
My impression was that Kubernetes is targeted to the large enterprise where you're going to go all in with containers and can dedicate a month or two to coming up to speed. Many of the discussions I saw talked about or gave the impression of dozens of nodes and months of setup.
Other options I'll probably look at when I have time to look at it again: Deis https://deis.com/ , Dokku http://dokku.viewdocs.io/dokku/ , Flynn https://flynn.io/ , LXC https://linuxcontainers.org/lxc/introduction/ and (though I'd been trying to avoid it) Docker Swarm https://docs.docker.com/engine/swarm/
You can use either their free tier in cloud or use the Open source OpenShift Origin for trials (there's also MiniShift, which is similar to MiniKube).
From my looking at it OpenShift comes with some of the parts that base Kubernetes leaves to plugins, so things like ingress, networking etc are installed as part of the base.
Anecdotally, I got an HA cluster running across 3 boxes in the space of about a month, with maybe 2-3 hours a day spent on it. The key for me was iterating, and probably that I have good experience with infrastructure in general. I started out with a single, insecure machine, added workers, then upgraded the workers to masters in an HA configuration.
I don't think it is really that hard to get a cluster going if you have some infrastructure and networking experience, especially if you start with low expectations and just tackle one thing at a time incrementally.
At Red Hat, we define an HA OpenShift/Kubernetes cluster as 3x3xN (3 masters, 3 infra nodes, 3 or more app nodes) [0] which means the API, etcd, the hosted local Container Registry, the Routers, and the App Nodes all provide (N-1)/2 fault tolerance.
Not to brag, since we're well practiced at this, but I can get a 3x3x3 cluster in a few hours, I've lead customer to a basic 3x3x3 install (no hands on keyboard) in less than 2 days, and our consultants are able to install a cluster in 3-5 working days about 90% of the time, even with impediments like corporate proxies, wonky DNS or AD/LDAP, not so Enterprise Load Balancers, and disconnected installs. Making a cluster read for production is about right-sizing and doing good testing.
[0] http://v1.uncontained.io/playbooks/installation/#cluster-des...
The great thing about automation is that once you have these basic tools (Prom/Graf monitoring/alerting, ELK, node pool autoscaling, CI/CD) implemented as declarative manifests, they're deployable anywhere in minutes.
Edit: especially load balacing the master servers. (that's actually the hard part of k8s, not even setting it up with/out openshift/ansible whatever) load balancing services on k8s itself is basically just running either calico network and use one or two haproxy deployments of size 1 with a ip annotation or just using https://github.com/kubernetes/contrib/tree/master/keepalived...
[0] https://github.com/openshift/openshift-ansible/blob/master/i...
I do have a lot of infrastructure and networking experience, it was mostly a matter of the ingress setup having many moving parts which were poorly documented. I could see that it had set up bridges and iptables rules and NAT and virtual interfaces, but I was never able to get a picture of how the setup was supposed to work to be able to see what parts of that picture were right or wrong.
There was no clear road-map of setting up a cluster. Most people talking about Kubernetes were doing "toy" deployments, which only had limited application to what I was doing. I only found kubespray because of a passing mention, for example.
I'd say your about right with a month. Had I given it another week or two, I probably would have gotten it going. I had really only expected it to take a couple days to have a proof of concept cluster, so at 2 weeks I was way beyond what I had slotted to spend on it.
Looking over the Getting Started Guide it looks very simple to get a test cluster set up. Which maybe set my expectations unreasonably high.
I guess that's what I'm trying to say: With the current state of documentation, it's probably a calendar month investment to get going.
Rancher 2.0 promises nice UI for K8s
(1) it's like 15 minutes to learn how it works and then maybe a morning or so playing around with it. Very low investment.
(2) if you're already using docker-compose, it's pretty much an in-place switch. You might need to deal with a few restrictions (mandatory overlay network, no custom container names, no parameterized docker-compose), in exchange you'll get zero-downtime upgrades, automated rollbacks, and of course the ability to add more machines to your swarm.
I know that people like to talk about cost savings and stuff like that, but I'd like to see if it actually lowered your app latency and increased sales/conversions/whatnot. Things that matter a lot to a growth business.
I ask because the various overlay/iptables/NAT/complicated networking setups in Kubernetes lend themselves to adding more overhead and being much slower than running on "bare metal" and talking directly to a "native" IP. I really, really wish that Kubernetes had full, built-in IPv6 support. It would remove a lot of this crud.
Our solution works around this by assigning IP's with Romana and advertising those into a full layer 3 network with bird. The pods register services into CoreDNS, and an "old fashioned" load balancer resolves service names into A records. Requests are then round robin'd across the various IP's directly. There's no overlay network. There's no central ingress controllers to deal with. There's no NAT. It's direct from LB to pod.
The nginx ingress controller is not a long term solution. It's a stop-gap measure. Someone really needs to build a proper, lightweight, programmable, and cheap, software-based load balancer that I can anycast to across several servers. That or Facebook just needs to open-source theirs.
For load balancing, have you looked at Traefik?
This tarball can be used to re-create the exact replicas of the original cluster, even on server farms not connected to the Internet, i.e. it includes container images, binaries for everything, etc.
But if the replica clusters do have internet access, they can dial back home and fetch updates from the "master". People use this to run complex SaaS applications inside their customers' data centers, on private clouds. These clusters run for months without any supervision, until the next update comes out.
Thanks to Kubernetes, you have a complete, 100% introspection into a complex cloud software, which allows for a use case like this. Basically if you're on Kubernetes you can no longer be tied to your one or two AWS regions and start selling your complex SaaS as downloadable run-it-yourself software. [1]
If the offline server farms hosted a private Docker registry (which is very simple to set up), couldn't you then simply push the container images to the registry, copy the relevant YAML files, and instantiate an identical cluster that way?
Efficiency is increased a great deal. Previously, I'd have idle instances costing money when they were unused. Now my resource consumption is amortized across my entire infrastructure.
It also drives an architecture that lends itself to better availability. A pod in k8s may be rescheduled at any moment. If it's a stateless service, one may just increase the number of replicas to prevent a service interruption. If it's a stateful service, one is forced to think about how to persist data and gracefully resume operations.
It took a while for everyone on my team to get familiar with it, but once we started to really grok it, the sense of safety went way up. I feel pretty good about the fact that if someone accidentally blew away a cluster with thousands of pods in it, that I could easily replicate the entire thing and get it back up and running in tens of minutes without a lot of hassle.
Also, my customers don't have strong requirements for network isolation. I'd be more concerned about compliance in areas like finance or healthcare.
I can't say that anything got worse. Once you (collectively) get past the initial learning curve and add automation in place, it's way better than most other deployment scenarios. Kubernetes worker nodes fail from time to time, AWS kills them and we barely notice. POds get rescheduled automatically and mostly work.
Frankly, the remaining headache-inducing things are mostly related to the software which is not running on Kubernetes (mostly stateful infrastructure, and mostly for non-technical reasons). Managing VMs is a pain.
There is one thing that one needs to be careful. If you are moving from a VM for each service, to Kubernetes, now suddenly your service is sharing a machine with other services, which wasn't the case before. So I'd suggest that proper resource limits be set so that the scheduler can work properly.
One thing that can get worse is network debugging. There are way more moving parts, so it is not as easy to just fire up wireshark.
So far zero problems.
The knowledge/documentation base just flat out sucks. It feels like documentation expects you to just know what you are doing and are really just reading it for the smaller feature switches to change how things work. That or expect you to be running on GKE.
That said, it's the best at doing what we need which is a scheduler and manager for running multi-container workloads.
I expect it to be more fulfilling/difficult when we move to more long-lasting pods but ultimately still use kubernetes.
P.S: if you’re interested in working on Kubernetes at reddit, send me a message at saurabh.sharma [at] reddit.com
We don't really have tremendous auto-scaling needs. But there are a few reasons why we did it.
We've since scaled to a small handful of services (frontend web app, backend node api, legacy stuff, custom deployments for high paying customers). We knew that standardizing on Docker was a good idea just to simplify the build process. What were previously service-specific, hard to maintain, hard to read, just weird shell scripts became a simple Dockerfile for each service, often only 5 or 6 lines long. Once you have that, setting up CI is ridiculously easy; AWS CodeBuild into ECS ECR took all of a day to implement. If we were on GCP it would have been even easier.
Comparatively, we'd spent longer than that actually maintaining the old scripts. The new ones require zero maintenance; they "just work", 100% of the time.
So we knew we wanted Docker. And that was a very good choice. No regrets. We start there. But once we had docker, our eyes turned to orchestration.
Its worth saying that there are a lot of nonobvious answers to questions in the "low/medium scale up devops" world. One big one we ran into early is "where do we store configuration?" Etcd or Consul are fine, but we're a pretty small shop; we didn't want to have to manage it ourselves. We could go with Compose or something. But, how do we get that configuration to the apps? We were in the process of writing new services and we wanted to follow 12-Factor and all that, so envvars make sense. But to get configuration from a remote source would break that, so we'd need some sort of hypervisor for the app to fetch that...
Additionally, how do you deploy? Let's say we go with a basic EC2 instance running Docker. We'd have to patch it together with shell scripts, ssh in, pull new images, GC old images (this was a huge problem with our "interim" step in the first paragraph, our EBS volumes filled up all the time, lots of manual work), restart, etc. Can you do that with zero downtime? Probably. More scripts...
Load balancing and HA? Yeah of course we can wire that up. More scripts. More CloudFormation templates or whatever...
Eventually you arrive at an inescapable conclusion: At a surprisingly low level of scale, you are reinventing Kube and Kube-like systems. So why not just use Kube? You get config management built in. You get a lot of deployment power in kubectl. You get auto-provisioning load balancers on AWS. You get everything.
And I could not be more serious when I say: It "just works". We've had two instances of random internal networking issues between our ingress controller and a service, which were resolved by restarting the ingress controller. That's... it, in a full year.
Like Docker before it, I think Kube is a foundational piece of technology that is only going to get more and more popular. Its incredible. But I do think that there is room in the market to make it easier to adopt. Best practices around internal configuration stuff is hard to come by; even something as simple as "I want a staging env, do i use two clusters or namespaces in one cluster" doesn't have a clear answer. Getting a local development environment cycle up and running is a pain in the ass. Monitoring and alerting is still pretty DIY; GCP solves some of this, but there's no turn-key alerting piece that I'm aware of. Logging is a nightmare if you are self-hosting, including on Kops; we ended up installing the Stackdriver agent and we use GCP for it, even though literally everything else in our stack is on AWS.
I literally don't believe there's a level of "production scale" your organization could be at where you wouldn't benefit from Kube. It is far far easier to set up than a bare deployment of anything if you do GKE. Connect Github to CI to Kube... good to go.
Compare it to init. Systemd is currently the dominant init for most Linux flavors. When your system starts and the kernel has done its thing, init takes over. It launches daemons, sets up networking, that sort of thing. After startup, it will relaunch things that crash, possibly do other things.
Kubernetes is sort of an init for an emerging pattern of cross-machine systems. Like systemd, it uses configuration to figure out what should run in what order, relaunches things that fail and manages networking (albeit in a very different context), that sort of thing.
There are huge differences because the problems are very different. But as a high level comparison, there are worse ones. (Especially if you're very charitable when interpreting "that sort of thing".)
You send the Master some YAML files that list what containers you want running, how many, and lots of optional details. Master makes it happen so it matches what you said should be running.
Master also constantly checks everything so it matches what you asked for, even if something changes like a server failing. That's it.
In the same way in OO languages we instantiate a class to get an object instance, containerisation seems to involve creating an instance (or several instances) of an image.
But I am really still learning all this stuff - I'm sure there's a reason behind it, and at the end of the day it's simply a set of terms one needs to absorb.
Instance of VM image => Running VM
Instance of Container image => Running Container
Of course I have come to terms with the fact that by now it's entrenched so we're using it, end of story - give me a few more years and I'll be pretty much adapted to it :-)
As others here have said, it’s a Container Orchestration Platform and is a largely pluggable architecture that also manages at varying extents Clustered Computing Resources, Application Resource Management, SDN Networking, DNS, Service Discovery, Load Balancing, and other concerns of “Cloud Native Application Development and Deployment“.
You can try it out at http://kubernetesbyexample.com/ .
Red Hat is a major contributor to Kubernetes and continually upstreams OpenShift features (like RBAC - they implemented it in OpenShift first, then upstreamed it, and then rebased OpenShift on top of it, removing the custom implementation).
I'm currently looking at migrating a large enterprise setup to Kubernetes/OpenShift.
Heck, even I am not sure what "application resource management" means, after working with Kubernetes for two years now. I could make educated guesses, but nothing certain.
Most people are familiar with memory isolation between processes. The use of further isolation mechanisms for the filesystem, I/O etc. is generally termed a container [1].
Kubernetes adds an orchestration layer on top of containers to manage processes running across different operating system instances.
[1] https://jvns.ca/blog/2016/10/10/what-even-is-a-container
Aside from that, it's a way to manage the deployment of containers across one or more hosts. Getting more traffic than usual? You have a program that detects it and sends a command to kubernetes/dockerswarm that tells it to scale up the number of web containers. Stuff like that.
The setup of the cluster is cloud specific but after that everything can just talk Kubernetes.
It chooses where to run the containers, wires up networking between them. And keeps them running.
I'm on my phone so I can't grab the link for you but a good intro is Kelsey Hightowers Kubernetes The Hard Way on GitHub.
You will use a tool called a “minikube” which is “mini” version of kubernetes that runs on your local machine.
You should focus on to get familiar with kubectl(1) first, it’s a simple CLI tool to manage kubernetes cluster.
If you have any questions & issues you can ask on Stack Overflow #kubernetes and you can also join the Slack channel.
The main benefits over the competing docker project, Docker Swarm, is that it does WAY MORE, is 100% free and open source, and has much better adoption.
I would argue that with Docker Swarm you have to bring the glue yourself, and it doesn't really solve any of the hard problems. Kubernets on the other hand is an all-in-one package that solves a LOT of hard problems for you.
Kubernetes is one solution. You use config files to specify how many of what should be running and it makes sure the system stays in that state and recovers as need be. It’s at a higher level of abstraction than running Docker containers manually on machines.
User feedback has been that that we want to keep that, but that we should also offer an alpha/beta of 1.10 much sooner, so that users that want to try out 1.10 today can do so (and so we get feedback earlier). So watch for kops 1.10 alpha very soon, and 1.11 alpha much earlier in the 1.11 cycle.
[0] https://gravitational.com/blog/kubernetes-release-cycle/
I mean without cloud services like Google cloud persistent disks.
IBM has been running kubernetes on baremetals internally for over a year and just recently announced it as a product: https://www.ibm.com/blogs/cloud-computing/2018/03/managed-ku....
I will admit that it isn't as straightforward to get it to work as you might imagine at first. But once you've got the automation humming, its been a surprisingly easy to maintain. Would highly recommend that route! (but of course I'm biased since I've worked on it)
One interesting use case is that its straightforward to access the underlying hardware on baremetal machines with `SecurityContext: priviliged` (you can do a much more fine-grained security permissions; I'm just giving an example). So for instance, you can access GPU's, TPM (trusted platform modules) this way.
Personally I wouldn't do it unless you have a RedHat contract (or already have a team that manages glusterfs), but its worth looking at.
you just use a storagecontroller, then its no work whatsoever, ceph does the rest
https://github.com/torchbox/k8s-hostpath-provisioner Proved to be useful, it allows you to use any mounted path on (all) a node(s) (hostpath method) to return satisfy a persistent volume claim and return a persistent volume backed by the mounted file system.
A similar set up could work using bare metal, are you using something like Openstack Ironic?
OpenSource: OpenEBS, Rook (based on Ceph), Rancher's Longhorn
Commercial options highly recommended if you want safety and support as storage is hard, although people seem to be running all of these options well enough. Portworx probably most highly developed with Quobyte a good option if you want shared storage with non-kubernetes servers.
IOW, they're not production ready yet.
I'm not affiliated with redhat in any way but I have enjoyed using the openshift platform.
Here's the link for anyone interested: https://www.openshift.com/pricing/index.html
I could finally run gitlab’s CI safely on kubernetes and generate containers.
The biggest problem I have is debugging all the moving parts when there are ~10+ minute async responses to config changes.
No?
So, I guess, let's dig into that a little bit. Erlang's always kind of left the health and safety of your deployed system as a whole up to you, being preoccupied with giving you tools for understanding and maintaining its internal operation. OTP provides two semi-unique things: supervision of computation with control of failure bubble-up and hot code reloading. Hot code reloading isn't used all that much in practice, outside of domains where it's _very important_ that the whole system can never be offline and load balancing techniques are not applicable. That's a specific niche and, sure, probably one that kubernetes can't service. With regard to supervision, there's no incompatibility between OTP's approach and a deployment consisting of of ephemeral nodes that live and then die by some external mechanism. Seems to me that kubernetes is no different a deployment target in this regard than is terraform/packer, hot-swapped servers in a rack or any of the other deploy methods I've seen in my career.
That said, there might be some pragmatic tradeoffs eventually made for the Erlang VM that I'm not familiar with yet that makes the container move more dubitable.
There is surprisingly little online discussion/documentation on the intersect of Resource Management and Container orchestration. Not sure if it is too early in the curve, a dark art, not actually done, or what...
OO comes with a bunch of really nice quality of life improvements that are missing (but in a lot of cases can be added via 3rd party TPR/CRDs/etc) in k8s but you aren't deviating so far from k8s that you won't be able to go work on vanilla k8s in the future. Not at all. Most of the additional stuff are annotations that simply wouldn't do anything and you'd remove them if you moved from OO to k8s.
I think that if you're brand new to the environment OO can really help you get running quickly. You just have to make sure that you do in fact dive in to the actual k8s yaml and deal with ingresses, prometheus, grafana, RBAC, etc at some point. I haven't used OO in awhile but I believe you could successfully do most of what I do day to day via yaml/json through the OO UI.
On the flip side a lot of people will probably tell you to start from k8s, whether that's GKE or AWS or minikube or wherever and go through the k8s the hard way. Personally, and I help people quite frequently on the k8s slack, I feel like that leads a lot of people down a path of frustration. It may be perfect for your style of learning or it may just scare you off.
Now when it comes to OO you are at the behest of their releases. Their most recent release was Nov 2017 and k8s 1.10 was just cut. I'm not sure what version OO is on now, evidently they changed their versioning numbers to not correspond to the k8s version.
Join #kubernetes-users and #kubernetes-novice on slack.k8s.io if you need any help. It's a vibrant community. You can message me directly @ mikej if you'd like.
edit: Ok OO 3.7 is k8s 1.9, that's perfect. I wait a few months before jumping into new major.
> An OpenShift Origin release corresponds to the Kubernetes distribution - for example, OpenShift 1.7 includes Kubernetes 1.7.
I had to dig around to figure out the version, might want to update that. :)
And great work, btw, OO is fantastic.
Kubernetes base is more flexible of course, but with that flexibility comes more stuff you have to do yourself...
Either way they're both based on the same tech., so experience with one largely translates across to the other
Personal anecdote: We've had OpenShift Origin running in production for a year and a half, I'm the DevOps guy and know how to poke around in the internals but the developers on my team are just that, developers. They just want to say "here's my code, go make it work", S2I, templates and Jenkins pipelines let them do just that - they go paste the URL of a git repository into the web interface and watch the project build and automatically deploy. It's a pretty magic experience watching a Junior developer deploy 13 applications over the course of a year with no supervision or hand-holding from a more-senior developer.
For example, Helm charts work on OpenShift: https://blog.openshift.com/getting-started-helm-openshift/
(Disclosure: I'm the executive director of CNCF and help run the Certified Kubernetes program.)
Are you using Kubernetes? And if not, what is your reason not to?
We are all-in-all very satisfied with Kubernetes, except that it's quite messy to do right on bare metal.
[1] Also available as open-source: https://github.com/sapcc/kubernikus
Disclosure: I'm executive director of CNCF, which funds the case study write-ups.
Kubectl auth login, is it available yet ?