Canonical introduces high-availability Micro-Kubernetes
zdnet.com
zdnet.com
I think Kubernetes would do well to make swapping out the storage backend possible without having all of these forks. Kubernetes is too tightly coupled to etcd, and for little benefit. I would wager a lot of customers would trade the guarantees etcd provides for a simpler deployment.
In general, I'm not blown away with Canonical's track record. I used microk8s in its early days and it kind of blew up my coworker's computer. Networking stopped working. Making a request to localhost:5000 would return an nginx error page, even though nginx wasn't running anywhere on that machine or the network. It even persisted itself through reboots! We just reinstalled the machine eventually. It was weird stuff and I haven't touched it again. (I prefer kind locally and k3s for small "real" clusters.) Then there's Snap, major breaking changes to Ubuntu for no reason, etc. I just don't trust Canonical much, and I'm not quite ready to take a leap of faith on their new distributed database. But maybe it's great, and we'll all be using this in a year. A stopped clock is right twice a day.
They lost me when I've tried their ubuntu server when I was lazy one day and greeted with their Landscape advertisement in the MOTD display.
Then I installed armbian's ubuntu version because Debian version was not ready and found out that MOTD was downloaded from web every time I log in.
Add analytics (now opt-in), forcing snaps and their silent-ish efforts to monopolize the landscape, I avoid them altogether.
We use it everywhere in my current place of employment and are very satisfied with it.
(maybe it's because we deploy servers using pxe — it makes good ubuntu installs, with no ads in motd & no snapd)
Deploying servers with PXE is fun though. We use PXE and XCAT to manage our fleet. Commissioning ~200 servers in 15 minutes with three commands while sipping coffee is really satisfying.
It doesn't put a "customized image of some sort" on servers, it literally iterates through installer steps with the autoanswers provided by preseed file.
You might say that it is customized since we wrote a preseed file. Well maybe, but it's not a "customized image". Also we didn't write any specific code to remove unwanted packages. It just doesn't install them. No weird packages like byobu :)
I say all this above, but I'm sure there are still shops out there burning ISOs to CDs (maybe USB now!) and hand installing everything - if I've learned anything, it's that nobody seems to agree on how to do stuff the same way. :) Each company is it's own beautiful unique snowflake full of their own design and deployment patterns.
The comment I am replying to is complaining about Ubuntu calling home, IBM OpenShift is the same story.
Not complaining - If you want me to complain about IBM cluster technology that would be a whole different story.
And no I'm not gonna ask about disconnected clusters :).
https://docs.openshift.com/container-platform/4.5/support/re...
I will follow up on opt-out being the default for the evaluation version of OpenShift. It is already opt-in for OKD.
It's funny when a corporation uses tactics of another corporation it wants to beat.
Debian has 24 hour security fixes for 15+ years. It also supports "OldStable" in terms of security & backports. For some time it also has "Long Term Support" teams which supports older releases.
Debian is "the original" install & forget distro. Ubuntu has some commercial sauce over it but, unless you have a special need, Debian can handle everything you throw at it.
We have so much servers so that we sometimes forget some of our background service servers' and they hum all-along with all security updates applied.
Edit: they deprecated the option https://rancher.com/docs/k3s/latest/en/installation/ha-embed...
No modern clustering system is much different. Services aren't killed off if a leader can't be elected. It's just that having all your servers alive but uncontrollable is almost as scary as them being down.
In fact, for something like what etcd is used for in k8s, as the root and lynchpin of all state, but not direct application I/O (pods use their own data stores, and even for cluster management etcd usually just contains pointers to external data), a leaderless architecture makes the most sense. Protocols to choose a leader are latency and throughput optimizations, but the core state for a cluster (members, pod metadata, etc) shouldn't require very much state and therefore should be relatively low traffic as compared to most database applications. And as shown in the above paper, for certain scenarios (including, arguably, the k8s scenario) a leaderless architecture can have better latency and throughput, not to mention availability.
Note that Kubernetes requires a total ordering of writes (which simplifies how hard it is for us puny humans to reason about) AND requires strong consistency in order to provide guarantees like “this pod only runs on one machine at a time” and “PVs aren’t released until the pods are really stopped”. Leaderless is a simple tradeoff - singleton or three instances. That’s the best possible choice in the world and it’s etcd and it’s relative simplicity that make it possible.
I’ve never seen a production Kube system with HA etcd go down due to non-human error, so I don’t believe single instance is going to give you better availability when single machine faults happen (they happen frequently; about 0.5-1% of the machines in the OpenShift fleet - cloud and on-premise - are down at any one time due to power, software, or hardware issues). Almost all of those clusters tick along fine when they lose that machine.
I have, with OpenShift, several times. The worst one took hours to get it partially back online, days to fully recover. At least I got to have one free drink before I had to leave our holiday party to start fixing it. (Ok, technically it was non-prod; luckily prod was ECS... long story)
> I don’t believe single instance is going to give you better availability when single machine faults happen
Depends on your definition and context of availability. It is possibly the simplest thing to destroy any existing nodes (and in the event that a node can't be destroyed, null-route it and/or disable its network port) and bring up a new node. So single instance is fairly easy to recover. But with multiple nodes, for each function of each instance that is required to reach consensus, the likelihood of consensus failure increases; you actually need more nodes and variety in the cluster to resist consensus failure. Yet ironically, the more nodes you have, the higher the probability of failure. In the end, the real-world reliability of the system is based on additional factors besides the network model.
Yeah, they really wanna be like Apple, but end up being more like Google, starting a lot of projects and abandoning them in couple of years.
The only place etcd might not pay off is in disposable dev environments or something, but do you really want your prod setup to page only to discover a complete cluster rebuild is necessary to resolve the problem?
Also, there was a bug with kube api where it wouldn't failover to other etcd members during an outage. So I would say most customers have ran kubernetes with a "single backend" at some point.
It’s more that the core team doesn’t have the time to support these, and we already have a fairly tight contract with storage, and we still find edge cases that need to be fixed, so the extra cost for the committers to support this is high.
though with kine (kine is not etcd) you can easily use mysql/sqlite/dqlite/postgresql.
Surprisingly, Etcd is neither scalable or high performance in our small cluster with <10 nodes on GKE. It had been the single most frequent source of prod outages.
But it appears that k8s community do not section a list of compatible storage engine that can plug and play with k8s. Or I might have missed some recent development.
The only remaining headache would be database. Setting up reliable mission-critical HA Postgres is kind of a pain, and it's nice to not have to worry about it. If that can also be canned, then this kind of rig would be a really compelling alternative to the managed clouds... unless you need other things like S3, etc.
ETCD is a nice piece of nick, but I am still disappointed that other options haven't displaced it, such as consul. That said, I am not sure I am ready to trust a new distribute store for this. Its hard to get that right.
How many folks are running self-managed k8s in production though, out of curiosity?
It seems so economical to deploy k3s or use something like Rancher on dirt cheap VPS's from some place like Hetzner -- but what's the ops burden and failure risk like?
Never tried it myself because it seems intimidating. Use managed k8s services.
i tried both in my homelab and this is my experience as well. microk8s seems to has less bug for me.
Oh wow.. Why so many clusters, and what kind of resources per cluster? Are you running 100 000s of physical servers?
Or are you just using k3s over three physical nodes as a way to achieve hardware redundancy and rolling hardware replacement?
Each restaurant has three Intel NUCs, so a little over 6K-7K devices. We're using k3s partly for the hardware redundancy, and partly for the ease of managing the services running at the edge - there's a local OAuth provider, MQTT server, along with some other applications that need to be up a majority of the time.
There's a cloud component to all of the software running at the edge as well, and since that's run on k8s, we wanted something similar but more lightweight at the edge.
Do you frequently see a nuc dying, but the cluster staying up?
Yes, in fact - we've seen clusters still running when two of the NUCs have disappeared from the cluster.
You mean when you want to run a small cluster of let's say less than 10 nodes (anything in single digits)?
Why doesn't normal k8 work this way? like same tech but just less scale?
(I should probably read more on the side myself as well, new to this K8 world).
I am guessing there were valid reasons for the offshoots (k3s, microk8s etc.).
As a k3s user (for local dev & my very small personal prod envs), I end up having to assemble a lot of these solutions myself. I prefer my own picks, my own solutions, and learning, but microk8s having these instantly available would be really good for a lot of folks. Until now though it has never felt like those advantages would be useful in a real prod environment, that microk8s was not interested in being in prod, but HA signals to me that they are interested in broader adoption.
For me, using k3s for development (not prod), the killer feature is running it in --docker mode where the node uses the local dockerd to run containers (vs. managing its own containerd instance).
This allows building images locally with `docker build` and immediately using them in kubernetes pods _without_ first pushing to a (possibly local) image repository and having k3s pull the images.
Last time I investigated, none of kind, microk8s, or minikube supported this mode. For large images (gigabyte or more, and I've got a handful of these), it's very space-inefficient to have a copy in my docker and in a local registry _and_ in k3s's containerd at the same time. How is this problem typically solved?
(I note the kubernetes included in Docker Desktop on macOS works in the same way: images built with `docker build` are available to kubernetes without going through a registry.)
I was surprised to be unable to otherwise find a good local / remote container development workflow, but this was built to replace what we were previously doing, which is setting DOCKER_HOST to point to the remote (single-node) cluster's Docker daemon (over SSH), so that docker CLI commands would execute on that remote box.
In both cases, you'll still want to take steps to minimize the size of your container image build context, but the size of the images doesn't matter. I'm not sure if it'd fit your needs or not, though.
https://minikube.sigs.k8s.io/docs/handbook/pushing/#1-pushin...
If you want to turn your host into a node you can do this with --vm-driver=none in minikube or better yet just use `kubeadm init` directly. In KIND we point people to the latter -- the main thing we're doing is running inside a disposable container "node" of which you can have many.
Assuming you don't actually want to turn your host machine into a node managed by Kubernetes, you'll want to stick Kubernetes in a VM or container. If the rest of Kubernetes is in this container or VM, it doesn't make sense to be running containers out on the host, things like mounting volumes won't work, you need a consistent filesystem between kubelet and the container runtime.
With kind it's also important that we simulate multi-node and multi-cluster, which is not viable with a single container runtime instance.
Without actually running Kubernetes against the hosts's runtime you can't share storage. The way docker desktop does this is to run Kubernetes with docker as the node's container runtime while only supporting a single node/cluster in a VM and expose the same runtime for building.
For our test workloads it's important to have different clean clusters constantly for different tests / projects.
KIND and microk8s have made a bet on containerd, as kubernetes is actively moving away from dockershim towards CRI, so even if we exposed a node runtime you can't build with it.
It's indeed space-inefficient, but it's a tradeoff in isolation between projects etc. For multi-node you're going to wind up with multiple copies anyhow, and a lot of projects we work with wind up needing some multi-node testing.
It looks like Canonical is just using dqlite directly.
I don't see all of these in microk8s. I'm not sure if this is on the roadmap.
I'm also not sure how customisable microk8s is. We run k3s with haproxy ingress (which is not the default) and calico for network (again nog the default)
I'm using it and it's been great so far.
I run Docker Swarm at home because my cluster is older, decommissioned machines. The Swarm daemon uses about 50MB of resident memory, which leaves a lot of room to run containers with little overhead.
It doesn't use a ton, but in order of least-to-most lightweight IME it's Minikube -> Microk8s -> k3s
But again, the overhead is marginal so it's not a world of difference.
If you can run Swarm you can definitely run one of the lightweight k8s distributions.
https://devblogs.microsoft.com/aspnet/introducing-project-ty...
Of course it does shift the complexity into the underlying K8S so installing and operating it can be difficult but there are manys to avoid that as a user. Overall you gain much more in productivity and usability, which is why it's taken off so much.
Is it? I can take a program written for Windows 95 and run it on Windows 7 (maybe even newer) just fine and it will run reliable and efficiently and integrate better than containers.
It is problem only on Linux because user space ABI keeps breaking.
Notice that the containers run on the same linux kernel and not in VMs, why? Because "we do not break userspace!".
To answer your question, I'm talking about Linux userspace ABI - you can't rely on ABI of essential libraries, openssl for example. That's why docker was born back then.
What is your replacement for all that?
I thought we were talking about running applications reliably. Why is complete linux userspace bundled separately in each container?
> Environments and languages change. Yes.
> That's entirely different to running programs with zero-downtime deployments, load-balancing and traffic management, health monitoring, logging and observability, secret and config management, storage volumes, security roles, and much more.
That's a lot of new requirements in addition to "running applications reliably". Most applications simply do not need that. And I believe this is the point of the original comment you replied to.
- zero-downtime deployments -> not needed for most applications (for example, twitter outages are not a big deal either). Btw how do you do zero-downtime of (websocket) streams with kubernetes? ;)
- load-balancing and traffic management -> in standard k8s you are pushing all traffic through one active LB (nginx) anyway => strip most of the extra layers and you dont even need that LB
- health monitoring, logging etc. -> you can use an existing solution that provides only the functionality you need, most of the work will be in your app anyway (every application needs different metrics ..)
The argument is that most applications do not need to scale at this level (until you need anycast DNS returning per-node IPs or at least geodns, you are not scaling that much anyway) and can be implemented in simpler manner hence easier and cheaper to maintain, audit and secure.
I do not want security roles, storage volumes and config management or observability I want my application to reliably and quickly serve my customers and be easy to maintain and debug.
If you want to discuss how to design architecture for scalable applications which keep all state in distributed databases but based on a platform with stable ABI that would surely be an interesting debate as well.
If you don't need K8S then don't use it. What's the problem? Run your app on your server and ignore everything else.
But most of these features have nothing to do with scale and are more about usability, reliability and consistency. Sure you can do it yourself but that's less efficient than just letting K8S do it all in one standardized way and interface.
> "I want my application to reliably and quickly serve my customers and be easy to maintain and debug."
That's what K8S helps with. I've spent 10 years running large distributed applications handling billions of requests per day in multiple regions. I don't care about the ABI and don't see why that's relevant, but I do know that K8S has made many things easier in actually running these apps.
See the original comment you replied to, he is clearly complaining about the whole infrastructure and solutions getting too complex for very little benefit, I just elaborated on that point because there is some truth to it. It is not the case for your scenario handling billions of requests per day in multiple regions - that's where it makes a lot of sense to use k8s! But very few applications need that.
> It's a container and can be as thin or fat as you want with it's contents.
But you can't rely on the platform, except for the kernel because linux kernel ABI is stable hence why the containers are done in this way. I am not complaining about it, I am exaplning the reasoning. Now imagine if you could rely on and share more services provided by the platform that just the kernel ;).
> I don't care about the ABI and don't see why that's relevant
Fair enough but then I don't understand why you replied to my comment saying the containers are designed in this way because of unstable userpace ABI if you don't care about this.
> "I want my application to reliably and quickly serve my customers and be easy to maintain and debug." >> That's what K8S helps with
For certain solutions, absolutely! For other solutions a simple stateful applications is simpler and easier to maintain and debug (again, that's how I read the first comment in this thread).
* It doesn't work as well
* Nobody outside your company has heard of it
* Most of the code is once-off hacky scripts that nobody really maintains
* Breakages will occur when underlying dependencies get upgraded when you patch your OS.
I'll even wager that nobody inside your company completely understand how it works. And when it breaks, you are 100% on the hook for it. There's no chance of official documentation or training, and no point asking for help on stackoverflow.
Gotta say, the docker ps output on those machines looks like line noise to me.
Not all movement is progress. Young developers are incentivized to embrace new technology because it levels a playing field where time with a tool is your most important asset. That playing field is not real, but a lot of managers seem to think it is, so the strategy works. I don't know how we sell a different version of reality there, but we need to figure it out.
At my first big job, the oldest developer told me shortly before he left that essentially we keep facing the same set of problems in a loop, and that if I watch for it I'll see it happening. That was in many ways a very big shoulder to stand on.
That challenge, added to historical information I had learned in a class on distributed computing, changed my perception of my first loop, like I'd found a shortcut. In my second loop, I found myself having productive conversations with people on their third loop, while my coworkers were still chirping on about how it's going to be different this time.
We just keep playing a game where the rules (like cost inequalities between resource classes) get tweaked every game, but quite often they revert back to the previous rules in the following game (because the people who make hardware for that resource finally figure out how to fix their bottleneck). But instead of recycling or democratizing the tech that worked last time the rules looked that way, we reinvent it badly with new names.
Elixir is a rare exception in this case, which is part of what attracts me to it. It's essentially recycling 25% of Erlang and 45% of Rails and being transparent about it, creating a new recipe out of old ingredients that have worked well in the previous 4 tech cycles.
That is why we are stuck in fashion industry nowadays.