Bare-Metal Kubernetes with K3s
blog.alexellis.io
blog.alexellis.io
I'm not really sure what people expect gain from these kinds of article, they're great as notes, but it's not something I'd use as a starting point for installing a production Kubernetes cluster.
The initial setup of a Kubernetes cluster is something most HN readers could do in half a day or so. Learning to manage a cluster, that's tricky. Even if you resort to tools like Rancher or similar, you're still in deep waters.
Also why would people assume that there's any difference in installing Kubernetes on an operating system running on physical hardware vs. on virtual machines?
This times 100. Deploying basic clusters is easy. Keeping a test/dev-cluster running for a while? Sure. Keeping production clusters running (TLS cert TTLs expiring, anyone?), upgrading to new K8s versions, proper monitoring (the whole stack, not just your app or the control-plane), provisioning (local) storage,... is where difficulties lie.
Our cluster configuration is public[1] and I’m almost done with a blog post going over all the different choices you can make wrt the surrounding monitoring/etc infrastructure on a Kubernetes cluster.
Another question: how about things like autoscaler which automatically edits numbers in deployments, how to git track that?
second you should always be ready to start from scratch, which is also pretty simple, because of terraform.
a lot of people are scared of k8s but they did not even try. they prefer to maintain their scary ansible/puppet whatever script that works only half as good as k8s.
Fair enough. I'll admit I have no direct experience with K3s. There are, however, many K8s deployment systems out there which I would not consider 'production-ready' at all even though they're marketed that way.
> second you should always be ready to start from scratch, which is also pretty simple, because of terraform.
That may all be possible if your environment can be spawned using Terraform (e.g., cloud/VMWare environments and similar). If your deployment targets physical servers in enterprise datacenters where you don't even fully own the OS layer, Terraform won't bring much.
> a lot of people are scared of k8s but they did not even try. they prefer to maintain their scary ansible/puppet whatever script that works only half as good as k8s.
We've been deploying and running K8s as part of our on-premises storage product offering since 2018, so 'scared' and 'didn't try' seems not applicable to my experience. Yes, our solution (MetalK8s, it's open source, PTAL) uses a tech 'half as good' as K8s (SaltStack, not Ansible or Puppet) because you need something to deploy/lifecycle said cluster. Once the basic K8s cluster is up, we run as much as possible 'inside' K8s. But IMO K8s is only a partial replacement for technologies like SaltStack and Ansible, i.e., in environments where you can somehow 'get' a (managed) K8s cluster out of thin air.
My start was the official kubernetes docs, step by step. Try everything, write it down, make ansible playbooks. Even used some of their interactive training modules, but I quickly had my own cluster up in vagrant so didn't really need the online shell.
Now we have two clusters at work, I have ansible playbooks I'm really happy with that help me manage both our on-prem clusters and my managed LKE with the same playbooks.
I'm completely sold on this container thing. :)
Running your two tightly coupled distributed monoliths of your front and back office systems no doubt!
They are focused only on CentOS 7, you might be able to find them yourself. As far as I can tell they include everything except persistent storage and HA control plane.
So it made a lot of sense to start at the official docs. And while I'm reading the docs, why not build the cluster at the same time. And while I'm building the cluster, why not write each step down in Ansible so I won't have to repeat myself.
So end result is my own ansible setup that I'm happy with and I know inside and out.
It always bothered me to run other people's ansible playbooks. I'm too much of a control freak.
The main difference between the different orch. tools is how they pipe the data between Outside and the service cluster. Mesos didn’t help you, you had to build your own service discovery. ECS uses Amazon load balancers and some custom black-box daemons. Kubernetes famously uses a complex mesh of “ingress controllers” and kernel iptables/ipfw spoofing. All take their configurations in some form of JSON or YAML.
Is there some sort of a PID-feedback-loop control that monitors the CPU/memory load and helps spin up more instances if it sees more traffic? If Kubernetes doesn't do that, what piece of software can help automatically scale if there is a huge traffic spike?
Load balance AFAIK doesn't do that. It just helps distribute the load.
[1]: https://kubernetes.io/docs/tasks/run-application/horizontal-...
The term was created for virtual machines.
It's neat. It is a lot of moving parts tho. I am just now trying it in a big infrastructure because we have so many bespoke parts that we have to glue together that we might as well try to use what is standard now...
You now have to do the configuration in YAML, which is MUCH MUCH worse.
Care to share an "easy" recipe, because I haven't found one that actually works. It always falls apart for me with the networking.
Let's say I have a cluster of 8 physical nodes and a management node available in my data center and I want to set up k8s for use by my internal users. I'm a solo admin with responsibilty for ~250 physical servers so any ongoing management necessary will be very much a task among many.
Is this blog post a good guide? Is there a better one?
Given you're running on physical infrastructure, MetalK8s [1] could be of interest (full disclosure: I'm one of the leads of said project, which is fully open-source and used as part of our commercial enterprise storage products)
Using the 2017 instructions went about like you'd expect in this space, with everything moving as fast as it does.
Using the instructions here has it complaining "Terraform initialized in an empty directory!"
mucking about and thinking the main.tfvars file was a typo and it wanted main.tf got things a little further along then generated another error because it WAS meant to be main.tfvars...
And at this point I'm frustrated, confused, and no closer to understanding the concepts...
This is as bare-metal as it gets since it's not running on a VM.
How much of linux isn't about cgroups and namespaces? Docker I believe needs about a 100 system calls to get containers to work. How much of the tree could you shake and still have containers work? Would you still call the host Linux, or something else? And what would that system do for Windows and OS X users? Anything?
Could you maintain this 'something else' as a permanent fork, a la Red Hat?
"A bare-metal server is a computer server that hosts one tenant, or consumer, only.[1] The term is used for distinguishing between servers that can host multiple tenants and which utilize virtualisation and cloud hosting.[2] Such servers are used by a single consumer and are not shared between consumers. Each server may run any amount of work for a user, or have multiple simultaneous users, but they are dedicated entirely to the entity who is renting them. Unlike servers in a data centre, they are not being shared between multiple customers.
Bare-metal servers are physical servers. Each server offered for rental is a distinct physical piece of hardware that is a functional server on its own. They are not virtual servers running in multiple pieces of shared hardware."
> In computer science, bare machine (or bare metal) refers to a computer executing instructions directly on logic hardware without an intervening operating system.
Bare-metal server has indeed been used to refer to non-virtualized servers, but it's a misnomer. The current terminology for this is "dedicated server".
Today when most tech people talk about bare metal they refer to a server that is not virtual.
If you disagree with what's on the page, please add or edit the information there. With proper references, it will be appreciated by everyone that uses Wikipedia.
But you still can't stop language from changing no matter how "right" you are you.
Not many things are written to do that, of course. Oracle used to offer an installation mode like this. It was generally a gimmick - you pay for a tiny bit of performance with a ton of flexibility. There are probably use cases where it makes sense, but not that many.
This usage has been, in my experience, a lot more widespread.
Oracle, and BEA before them, used to offer a JVM which ran on top of a thin custom OS designed only to host the JVM, you could call it a "unikernel". Product was called JRockit Virtual Edition (JRVE), WebLogic Server Virtual Edition (WLS-VE, when used to run WebLogic), earlier BEA called it LiquidVM. The internal name for that thin custom OS was in fact "Bare Metal". Similar in concept to https://github.com/cloudius-systems/osv but completely different implementation
I think one thing which caused a problem for it, is a lot of customers want to deploy various management tools to their VMs (security auditing software, performance monitoring software, etc) and when your VM runs a custom OS that becomes very difficult or impossible. So adopting this product could lead to the pain of having to ask for exceptions to policies requiring those tools and then defending the decision to adopt it against those who use those policies to argue against it. I think this is part of why the product was discontinued.
Nowadays, Oracle offers "bare metal servers" [1] – which are just hypervisor-less servers, same as other cloud vendors do. Or similarly, "Oracle Database Appliance Bare Metal System" [2] – which just means not installing a hypervisor on your Oracle Database Appliance.
So Oracle seems to have a history of using the phrase "bare metal" in both the senses being discussed here.
[1] https://www.oracle.com/cloud/compute/bare-metal.html
[2] https://docs.oracle.com/en/engineered-systems/oracle-databas...
I will say, this comment section is the first time I'm hearing about "bare-metal" meaning "without an OS", but the above question is genuine curiosity.
I've never before encountered, "Includes a full feature-rich OS" as "bare-metal" before. Reading the title I assumed someone managed to get some flavor of Kubernetes running right on the hardware as the lowest-level software layer of the system. That would have meant bare-metal to me. What's described here is running Kubernetes on a physical host rather than a virtual host from what I can tell, but it's not running Kubernetes "bare-metal" because between Kubernetes and the "metal" is Linux.
Or at least that's what it would mean in my world, but the interpretation appears to be different for others. Outside of confusion, I was also just disappointed. This article is just basically setting up Kubernetes. That it's on a physical host is a lot less interesting and novel to me than if they'd managed to implement some shape of Kubernetes as the OS itself, which is what I'd originally interpreted the title to mean.
> In computer science, bare machine (or bare metal) refers to a computer executing instructions directly on logic hardware without an intervening operating system.
Bare-metal = no VMs or other virtualization involved.
I imagine it also has some meaning for the music community.
https://apps.dtic.mil/dtic/tr/fulltext/u2/a219356.pdf (search for bare-metal)
I suspect it comes from the automotive paint industry. Sanding down to the "bare metal" for the best finish...where the primer and paint are as close to the substrate as they can be.
One of these guests is alpine with k3s (tried k3os, very limiting in a good way) that allows me to pass a host directory directly into the vm using p9 (tried nfs but lil heavy)
So any storage needs get the benifit of regular snapshots, compression and sync to nas.
Really wish it would be more know that k8 single node is all most people need to get started!
Actual kubernetes deployments and services etc deployed with ansible managed helm
I used to have 3 low-power Nuc-style computers, and after playing around for a few months I did the same thing as you have - replaced them with a single beefier machine, and its a lot more practical.
There was a user who was paying 7 USD / mo per site to host 20 side-project sides.. expensive. They switched out to a computer under their desk and saved a lot of money that way.
You mount this in your guest vms fstab and specify 9p in the options
Using bare metal servers without VM layer is actually a simplification. Cutting out a layer that's not strictly necessary.
Test environments were in AWS. There is a load balancer outside of the cluster (highly available HAProxy as a service). I wouldn't say it's particularly difficult or easy. It's pretty cost effective. After the initial setup, scripting and testing, is done, you spend at most a few hours per month with maintenance and the difference in cost of severs is huge. Also, unmetered bandwidth.
The pain points are mostly storage (nothing beats redundant network storage ala EBS) and having to plan at least a few months in advance because you're renting larger chunks of HW.
I've installed k8s with ansible on baremetal (kubespray), more or less just followed the steps here: https://kubernetes.io/docs/setup/production-environment/tool...
No network virtualisation, just Calico. Announce the service ips via BGP from each node running the service and ECMP gives you a (poor mans) load-balancing. Ingress gets such a service-ip. I used simply nginx.
Important here though is, that the router needs to be able to do resilient hashing: Removing a node or adding a node otherwise causes a rehash of all connections leading to breaking connections.
Calico? Network virtualization? BGP? ECMP? Resilient hashing?
No big surprise all this stuff is easy for you.
Networking is handled in kubernetes with CNI plugins, Calico is one of them. They define how one pod can talk to another.
Probably best described in how it does it is by the project itself: https://docs.projectcalico.org/about/about-networking
My simplyfied version: Calico uses the IP routing facilities to route IP packets to pods over hosts. Either from another pod or from a gateway router.
BGP is a protocol to exchange routing information, so it can be used to inform the router or kubernetes nodes (in this case physical hosts) about where to send the IP packets.
If a pod is running on a node, the node announces with BGP that the pod IP can be routed over the IP of the node. If the pod provides a service (in the kubernetes sense), the node can also announce that the service IP can be routed over the same host. Now, if two pods on different nodes are providing the same service, then both are announcing the same service IP. So, there are multiple routes or multiple paths for the same IP. That are the last to letters of the acronym ECMP (Equal Cost Multiple Path). Equal cost, because we do not express a preference over one or the other.
The router then can make a decision where to send the packets to. Usually that is done by hashing some part of the IP packet (IP and port of source and target for example).
Now the question is how is that hash deciding to which host it goes? In most cases it is very simply that you have an array of hosts, and the hash modulo the length gives you the host. Problem is, if you add or remove one item from that, practically all future packets will end up at a different host than before you did so. And they don't know what to do with it, breaking the connection (in case of TCP). Resilient hashing describes a feature that the mapping won't change under changes.
Not sure what you mean re: "network virtualisation" though?
k3s is actually pretty simple to use now. the tricky part was to integrate with https://github.com/kubernetes/cloud-provider-aws and https://github.com/DirectXMan12/k8s-prometheus-adapter
The hardest part is to get it to work with spot instances. we use https://github.com/AutoSpotting/AutoSpotting to integrate with it.
i have 3 cheapish VPS with public IPs and no VPC . Am but struggling to find a way to get k3s HA working with MetalLB, I dont have EIP /NodeBalancer resource, and dont want to resort to mere 1master+2agents cluster. Any tips or links are appreciated
I would probably suggest you go for 3x servers in HA, it uses marginally more memory but can tolerate a failure.
You can always go and use DigitalOcean managed K8s, or self-host at home using inlets that I mention in the post.
https://kube-vip.io does two things: - Control Plane HA w/BGP or ARP - Service Type: LoadBalancer w/BGP or ARP
It is similar to metallb but has a number of differences under the covers in how it works inside Kubernetes, it is cloud controller agnostic so as long as something attaches an IP to spec.IngressIP then kube-vip will advertise it. For edge deployments loadBalancers addresses can use the local DHCP for addresses etc..
This was atually the part I was curious about. Could you elaborate? Or is there a design doc somewhere?
Think various boards bootloaders, not even BIOS to rely on, that's running on bare metal. Calling something running on a normal x64 OS launched with UEFI as such is silly and inaccurate.
This usage has been commonplace for a long time, probably two decades.
https://en.wikipedia.org/wiki/Close_to_Metal
https://en.wikipedia.org/wiki/Bare_machine
https://wiki.c2.com/?CloseToTheMetal
This term goes back to early 2000 at least.
Can we have nice things without it being a hook to buy something?
If you want to go cheap, you can host at home on old servers or RPis.. there's even links included to help you with that.
Or are you making a metacomment about Terraform being painful at times?