MicroK8s – Low-ops, minimal Kubernetes, for cloud, clusters, Edge and IoT
microk8s.io
microk8s.io
- https://docs.k0sproject.io (https://github.com/k0sproject/k0s)
- https://k3s.io (https://github.com/k3s-io/k3s/)
k0s is my personal favorite and what I run, the decisions they have made align very well with how I want to run my clusters versus k3s which is similar but slightly different. Of course you also can't go wrong with kubeadm[0][1] -- it was good enough to use minimally (as in you could imagine sprinkling a tiny bit of ansible and maintaining a cluster easily) years ago, and has only gotten better.
[0]: https://kubernetes.io/docs/reference/setup-tools/kubeadm/
We switched to rke, it’s much better.
The problem with k3s is that the architecture level libraries are a bit outdated. Early on, it was for a particular reason - ARM64 (raspberry pi) support. But today, like everyone is on ARM - even AWS.
For example the network library is Flannel. Almost everyone switches to Calico for any real work stuff on k3s. it is not even a packaged alternative. Go-do-it-urself.
The biggest reason for this is a core tenet of k3s - small size. k0s has taken the opposite approach here. 50mb vs 150mb is not really significant. But it opens up alternative paths which k3s is not willing to take.
In the long run, while I love k3s to bits....I feel that k0s with its size-is-not-the-only thing approach is far more pragmatic and open for adoption.
Most of the other choices that k0s makes are also right up my alley as well. I personally like that they're not trying to ride this IoT/Edge wave. Nothing wrong with those use cases but I want to run k8s on powerful servers in the cloud, and I just want something that does it's best to get out of my way (and of course, k0s goes above and beyond on that front).
> The biggest reason for this is a core tenet of k3s - small size. k0s has taken the opposite approach here. 50mb vs 150mb is not really significant. But it opens up alternative paths which k3s is not willing to take.
Yup! 150MB is nothing to me -- I waste more space in wasted docker container layers, and since they don't particularly aim for IoT or edge so it's perfect for me.
k3s is great (alexellis is awesome), k0s is great (the team at mirantis is awesome) -- we're spoiled for choice these days.
Almost criminal how easy it is to get started with k8s (and with a relatively decent standards compliant setup at that!), almost makes me feel like all the time I spent standing up, blowing up, and recreating clusters was wasted! Though I do wonder if newcomers these days get enough exposure to things going wrong at the lower layers as I did though.
[0]: https://itnext.io/benchmark-results-of-kubernetes-network-pl...
Things like proxy protocol support (which is pretty critical behind cloud loadbalancers), network plugin choice, etc is going to be very critical.
What's the tradeoff? Why not flannel for Real Work™?
- network policy enforcement
- intra-node traffic encryption with wireguard
- calico does not use VXLAN (sends routes via BGP and does some gateway trickery[0]), so it has slightly less overhead
[0]: https://stardomsolutions.blogspot.com/2019/06/flannel-vs-cal...
If developing locally with k8s would likely be a better workflow, are any of these options better than the others for that?
The biggest benefit there is no need to have a docker compose or have other resources running locally you just can run the test cases if you have docker installed. [0] https://github.com/ory/dockertest
I usually only reach for it when I am building out a helm charm for a project and want to test it. Otherwise docker-compose is usually enough and is less boilerplate to just get an app and a few supporting resources up and running.
One thing I have been wanting to experiment with more is using something like Tilt [1] for local development. I just have not had an app that required it yet.
[0] https://k3d.io/ [1] https://tilt.dev/
If folks are interested in this kind of K8s deployment, they might also be interested at what we're doing at Talos (https://talos.dev). We have full support for all of these same environments (we have a great community of k8s-at-home folks running with Raspberry Pis) and a bunch of tooling to make bare metal easier with Cluster API. You can also do the minikube type of thing by running Talos directly in Docker or QEMU with `talosctl`.
Talos works with an API instead of SSH/Bash, so there's some interesting things about ease of use when operating K8s that are baked in like built-in etcd backup/restore, k8s upgrades, etc.
We're also right in the middle of building out our next release that will have native Wireguard functionality and enable truly hybrid K8s clusters. This should be a big deal for edge deployments and we're super excited about it.
- How does Talos handle first getting on to the network? For example, some environments might require a static IP/gateway for example to first reach the Internet. Others might require DHCP.
- How does Talos handle upgrades? Can it self upgrade once deployed?
- What hardware can Talos run on? Does it work well with virtualisation?
- To what degree can Talos dynamically configure itself? What I mean by this is if that a new disk is attached, can it partition it and start storing things on it?
- How resilient is Talos to things like filesystem corruption?
- What are the minimum hardware requirements?
Please forgive my laziness but maybe other HNers will have the same questions.- How does Talos handle first getting on to the network? For example, some environments might require a static IP/gateway for example to first reach the Internet. Others might require DHCP.
For networking in particular, you can configure interfaces directly at boot by using kernel args.
But that being said, Talos is entirely driven by a machine config file and there are several different ways of getting Talos off the ground, be it with ISO or any of our cloud images. Generally you can bring your own pre-defined machine configs to get everything configured from the start or you can boot the ISO and configure it via our interactive installer once the machine is online.
We also have folks that make heavy use of Cluster API and thus the config generation is all handled automatically based on the providers being used.
- How does Talos handle upgrades? Can it self upgrade once deployed?
Upgrades can be kicked off manually with `talosctl` or can be done automatically with our upgrade operator. We're currently in the process of revamping the upgrade operator to be smarter, however so it's in flux a bit. As with everything in Talos, upgrades are controllable by the API.
Kubernetes upgrades can also be performed across the cluster directly with `talosctl`. We’ve tried to bake in a lot of these common operations tasks directly into the system to make it easier for everyone.
- What hardware can Talos run on? Does it work well with virtualisation?
Pretty much anything ARM64 or AMD64 will work. We have folks that run in cloud, bare metal servers, Raspberry Pis, you name it. We publish images for all of these with each release.
Talos works very well with virtualization, whether that's in the cloud or with QEMU or VMWare. We've got folks running it everywhere.
- To what degree can Talos dynamically configure itself? What I mean by this is if a new disk is attached, can it partition it and start storing things on it?
Presently, the machine configuration allows you to specify additional disks to be used for non-Talos functions, including formatting and mounting them. However, this is currently an install-time function. We will be extending this in the future to allow for dynamic provisioning utilizing the new Common Operating System Interface (COSI) spec. This is a general specification which we are actively developing both internally and in collaboration with interested parties across the Kubernetes community. You can check that out here if you have interest: https://github.com/cosi-project/community
- How resilient is Talos to things like filesystem corruption?
Like any OS, filesystem corruption can indeed occur. We use standard Linux filesystems which have internal consistency checks, but ultimately, things can go wrong. An important design goal of Talos, however, is that it is designed for distributed systems and, as such, is designed to be thrown away and replaced easily when something goes awry. We also try to make it very easy to backup the things that matter from a Kubernetes perspective like etcd.
- What are the minimum hardware requirements?
Tiny. We run completely in RAM and Talos is less than 100MB. But keep in mind that you still have to run Kubernetes, so there's some overhead there as well. You’ll have container images which need to be downloaded, both for the internal Kubernetes components and for your own applications. We're roughly the same as whatever is required for something like K3s, but probably even a bit less since we don’t require a full Linux distro to get going.
A brief description of what CSI is - https://kubernetes.io/blog/2019/01/15/container-storage-inte...
There are also SeaweedFS CSI Driver: https://github.com/seaweedfs/seaweedfs-csi-driver
[1] https://docs.min.io/docs/distributed-minio-quickstart-guide....
A few reasons:
- Rook is "just" managed Ceph[2], and Ceph is good enough for CERN[3]. But it does need raw disks (nothing saying these can't be loopback drives but there is a performance cost)
- OpenEBS has a lot of choices (Jiva is the simplest and is Longhorn[4] underneath, cStor is based on uZFS, Mayastor is their new thing with lots of interesting features like NVMe-oF, there's localpv-zfs which might be nice for your projects that want ZFS, regular host provisioning as well.
Another option which I rate slightly less is LINSTOR (via piraeus-operator or kube-linstor[6]). In my production environment I run Ceph -- it's almost certainly the best off the shelf option due to the features, support, and ecosystem around Ceph.
I've done some experiments with a reproducible repo (Hetzner dedicated hardware) attached as well[7]. I think the results might be somewhat scuffed but worth a look maybe anyways. I also have some older experiments comparing OpenEBS Jiva (AKA Longhorn) and HostPath [8].
[0]: https://github.com/rook/rook
[1]: https://openebs.io/
[3]: https://www.youtube.com/watch?v=OopRMUYiY5E
[5]: https://github.com/piraeusdatastore/piraeus-operator
[6]: https://github.com/kvaps/kube-linstor
[7]: https://vadosware.io/post/k8s-storage-provider-benchmarks-ro...
[8]: https://vadosware.io/post/comparing-openebs-and-hostpath/
There are even now higher-level tools such as k3os and k3sup to further reduce the initial deployment pains.
MicroK8s prides with 'No APIs added or removed'. That's not that positive in my book. K3s on the other hand actively removes the alpha APIs to reduce the binary size and memory usage. Works great if you only use stable Kubernetes primitives.
[0] https://k3s.io/
I used Microk8s on a client project late last year and it was really painful, but I am sure it serves a particular set of users who are very much into the Snap/Canonical ecosystem.
In contrast, K3s is very light-weight and can be run in a container via the K3d project.
If folks want to work with K8s upstream or development patches against Kubernetes, they will probably find that KinD is much quicker and easier.
Minikube has also got a lot of love recently, and can run without having a dependency on Virtual Box too.
[0] https://k3sup.dev/ [1] https://kind.sigs.k8s.io/docs/user/quick-start/
I run a couple of small clusters and my Ansible script for installing them is pretty much:
* Set up the base system. Set up firewall. Add k8s repo. Keep back kubelet & kubeadm.
* Install and configure docker.
* On one node, run kubeadm init. Capture the output.
* Install flannel networking.
* On the other nodes, run the join command that is printed out by kubeadm init.
Running in a multi-master setup requires an extra argument to kubeadm init. There are a couple of other bits of faffing about to get metrics working, but the documentation covers that pretty clearly.I'm definitely not knocking k3s/microk8s, they're a great and quick way to experiment with Kubernetes (and so is GKE).
I haven't done a manual deployment since. I hope it got significantly better and I may be an idiot but the reputation isn't fully undeserved.
The problem back then was also that this was usually the first thing you had to do to try it out. Doing a complicated deployment without knowing much about it doesn't make it any easier.
The thing that concerns me the most is managing the internal certificates and debugging networking issues.
I haven't yet set it up, but https://github.com/kontena/kubelet-rubber-stamp is on my list to look at.
> debugging networking issues
In this regard, I have had much more success with flannel than with calico. The BGP part of calico was relatively easy to get working, but the iptables part had issues in my set-up and I couldn't understand how to begin debugging them.
The server itself, is, probably, overprovision, but I still struggle with responsiveness, logging, and ingress/service management. What's also funny, that using the Ubuntu's ufw service is not that seamless together with microk8s.
I am think of moving now to k3s. The only thing that's holding me back is that k3s doesn't use nginx for ingress so I'll need to change some configs.
Also, the local storage options are not that clear.
[0]: https://vadosware.io/post/ingress-controller-considerations-...
[1]: https://vadosware.io/post/stuffing-both-ssh-and-https-on-por...
At previous employment with large scale clusters (thousands of nodes) Traefik seemed to be heavily preferred by the SRE’s in my org.
You could get Traefik working inside your existing cluster actually by just letting NGINX route to it and seeing how easy it is to use that way -- though that may be more difficult than just spinning up a cluster on a brand new machine (or locally) and feeling your way around.
I can say that Traefik's dashboard makes it much easier to debug while it's running as it gives you a fantastic amount of feedback, Prometheus built in, etc.
Last year I checked and every physical machine in a K8S cluster was burning CPU at 20-30% - with zero payload, just to keep itself up!
Don´t you feel that this is totally inacceptable in a world with well understood climate challenges?
Every time I see it mentioned, I check to see if they ditched snapd yet, alas, today is not the that day.
I am the creator of the k3sup (k3s installer) tool that was mentioned and have a fair amount of experience with K3s on Raspberry Pi too.
You might also like my video "Exploring K3s with K3sup" - https://www.youtube.com/watch?v=_1kEF-Jd9pw
Kube is designed for nodes to have continuous connectivity to the control plane. If connectivity is disrupted and the machine restarts, none of the workloads will be restarted until connectivity is restored.
I.e. if you can have up to 10m of network disruption then at worst a restart / reboot will take 10m to restore the apps on that node.
Many other components (networking, storage, per node workers) will likely also have issues if they aren’t tested in those scenarios (i’ve seen some networking plugins hang or otherwise fail).
That said, there are lots of people successfully running clusters like this as long as worst case network disruption is bounded, and it’s a solvable problem for many of them.
I doubt we’ll see a significant investment in local resilience in Kubelet from the core project (because it’s a lot of work), but I do think eventually it might get addressed in the community (lots of OpenShift customers have asked for that behavior). The easier way today is run edge single node clusters, but then you have to invent new distribution and rollout models on top (ie gitops / custom controllers) instead of being able to reuse daemonsets.
We are experimenting in various ecosystem projects with patterns would let you map a daemonset on one cluster to smaller / distributed daemonsets on other clusters (which gets you local resilience).
Where have you been suffering from this?
I don't want to have to restart the whole thing on each site every time it happens. I'd like a deployment/orchestration system that can work in such scenarios, showing a node as unreachable but then back online when it gets network back.
Isn't that exactly what happens with K8s worker nodes? They will show as "not ready" but will be back once connectivity is restored.
EDIT: Just saw that the intention is to have some nodes in a DC and some nodes in the edge and the intention is to have a single K8s cluster spanning both locations with unreliable network in between. No idea how badly the cluster would react to this.
Don't do this. Have two K8s clusters. Even if the network were reliable you might still have issues spanning the overlay network geographically.
If you _really_ need to manage them as a unit for whatever reason, federate them(keeping in mind that federation is still not GA). But keep each control plane local.
Then setup the data flows as if K8s wasn't in the picture at all.
No, it's ok. What you don't want to have is:
* Poor connectivity between K8s masters and ETCD. The backing store needs to be reliable or things don't work right. If it's an IOT scenario, it's possible you won't have multiple k8s master nodes anyway. If you can place etcd and k8s master in the same machine, you are fine.
You need to have a not horrible connection between masters and workers. If connectivity gets disrupted for a long time and nodes start going NotReady then, depending on how your cluster and workloads are configured, K8s may start shuffling things around to work around the (perceived) node failure(which is normally a very good thing). If this happens too often and for too long time it can be disruptive to your workloads. If it's sporadic, it can be a good thing to have K8s route around the failure.
So, if that is your scenario, then you will need to adjust. But keep in mind that no matter what you do, if network is really bad, you would have to mitigate the effects regardless, Kubernetes or not. I can only really see a problem if a) network is terrible and b) your workloads are mostly computing in nature and don't rely on the network (or they communicate in bursts). Otherwise, a network failure means you can't reach your applications anyway...
* I don't like Snap. Like, a lot. But unfortunately there aren't any other options at the moment.
** I have HAProxy load-balancing both the k8s API and the ingresses. Both on L4 so I can terminate TLS on the ingress controller and automatically provision Let's Encrypt certs using cert-manager[1].
*** Keepalived[2] juggles a single floating IP between all the nodes so you can just run HAProxy on the microk8s nodes instead of having dedicated external loadbalancers.
https://github.com/ubuntu/microk8s/issues/988#issuecomment-6...
What kind of workloads are you running on it? Whoever has OpenFaas installed on your rpi, what type of functions are you running?
The real value I get is Infra as Code, and having access to the same powerful tools that are standard everywhere else. I can drop my current (dedicated) hardware and go publish all my services on a new box in under an hour, which I did recently.
From my point of view, I already pay the complexity by virtue of k8s being The Standard. The two costs of complexity are 1) Learning the Damn Thing 2) Picking up the pieces when it breaks under its own weight. 1) I already pay regardless, and 2) I'm happy to pay for the flexibility I get in return.
(Disclaimer: I work for Canonical)
Are these essentially different choices available for people who wish to self host their k8s clusters?
Are there any good resources that summarise the different advantages of each choice? What is the same between all the choices, what are the differences?
Most distributions will use the stock kubelet process (what runs containers on nodes), but you don't have to--there are kubelet compatible processes to run VMs instead of containers, run webassembly code (krustlet), etc.
Everything in k8s can be swapped out and changed, the API spec is the only constant.
Microk8s pros:
- bind-mount support => it is possible to mount a project including its node_modules and work on it from the host while it hot-reloads in the pod.
- The addons for dns, registry, istio, and metallb just work.
- Feels more snappy than minikube.
Microk8s cons:
- Distributed exclusively via snap => can't be easily installed on nix/nixos.
Minikube is very effective on my huble opinion to STUDY K8s becauese it is always strong aligned with K8s releases.
I do not think minikube is a good choice for production but hey, I could be wrong... someone want to share any experience?
After shipping a number of MQTT backends for deployments of few million devices each, I'd say the most troublesome part was getting the anycast network setup, uplinks, routing, and getting the "messaging CDN" second to it.
It's very hard to do without having an own ASN in which you have complete freedom.
Even in the case of 8 million shipped units fire/burglary alarm client, with ~4M constantly on units, which all have to send frequent pings to signal that they are online, I haven't seen any need in any kind of sophisticated clustering.
Sure if you run an entire control plane on the edge you're adding more complexity... but you don't have to do that, and control planes are complex beasts by their nature.
Kubernetes has made efforts for nodes to work with poor network connectivity.
The node requirements aren't that big.
A lot of use cases are relatively easy to containerise.
And edge / IoT devices are getting more powerful as well.
They aren't talking about consumer IoT here as well.
It has become a meme that Kubernetes is complicated. But it solves a lot of orchestration problems that would need to be implemented in other ways.
On the backend where you have services being fed, processing, and presenting all that data sure Kubernetes that part up, but that doesn't seem to need a special distro of Kubernetes.
when I think IoT, I think small single purpose devices: a ring door bell, a "smart" thermostat or fire alarm, a security camera. There at most a couple of processes running, what is there to orchestrate on an IoT doorbell?
Exactly!I suppose it might depend on what you count as "edge", but we're using kubernetes to distribute a complex product to customers onprem. The product has multiple databases, services, transient processes, scheduled jobs, and machine learning. It needs to be able to run on a single machine or a cluster depending on customer requirements. It needs to support whatever Linux variant the customer allows. Using Kubernetes solves a lot of problems for us.
For example, SQLite advertise itself as "database on edge".
Just as an example (not saying it's authoritative):
> "Edge computing is often referred to as 'on-premise.'"
-- https://dzone.com/articles/demystifying-the-edge-vs-cloud-computing
But these days, people even refer to systems hosted on a company's own cloud account as "on-premise", so these terms get increasingly fuzzy over time.Btw since you mention SQLite, the k3s system uses SQLite instead of the default etcd used by Kubernetes, for the reason you mention. These systems really are intended to support true edge scenarios. K3s is distributed as a single 40 MB binary, and you can run it on non-PC edge hardware.
For anyone who runs a system that involves multiple containers on a single machine, it can be worth looking at systems like k3s as an alternative to e.g. Docker Compose. There's not much downside other than some learning curve, and it gives you a wealth of capabilities that you otherwise tend to end up hacking together with scripts or whatever.
Technically they’re all edge, but nobody thinks K8s can run on the latter, but it might work on the prior two.
Pity Istio doesn’t offer an ARM build
- istio SC member
The docs say recommended 4gb memory, but could I run a huge swap partition for that?
[1] https://ubuntu.com/tutorials/how-to-kubernetes-cluster-on-ra...
If you mean "docker" the binary used to build container images, you don't even need that right now -- there are multiple projects that will build container images without involving docker or dockerd.
If you mean "dockerd" the container management engine that one controls via "docker" to start and stop containers, then yes microk8s will help as they appear to use containerd inside the snap (just like kind and likely all such "single binary kubernetes" setups do): https://github.com/ubuntu/microk8s/blob/master/docs/build.md...
I would love for something with more sanity to catch on, because Dockerfile as a format has terrible DX, but "wishes horses etc etc."