Why Fix Kubernetes and Systemd?
medium.com
medium.com
systemd is a real power tool and the more I learn about it the better I like it.
When you need to get something done, systemd often as not has your back.
Simple example - recently I needed to ensure a given service could not send more data than a given limit. Easy - systemd includes traffic accounting on a per service basis - all I needed to do was switch it on with "IPAccounting=yes" and read the values and kill the service if its data egress was beyond my maximum.
This is just one simple example of how systemd makes life easy - there's many, many more examples.
For example did you know that systemd/nspawnd lets you do virtualisation with all the benefits of containerisation such as resource/memory optimisation? The more you dig, the more you'll find to like about systemd.
systemd haters gonna hate but show me something better.....
If anything, systemd influences what modern Linux is so heavily that modern Linux should be referred to as linux/systemd instead of gnu/linux.
It's like I'm in a hall of mirrors......
It's the integration of systemd, the common approach that yields the giant payoff. I understand the "lots of independent utilities" Unix philosophy but systemd's consistent broad scope yields increasing benefits.
See also vim vs Emacs.
I think there are other reasons why people dislike systemd (and recalling a few articles with people ranting about it; I think there are many around), and usefulness of standards and alternatives/choice wouldn't necessarily make sense either, but here is at least one of the ways to view it.
As for init scripts, I think it can be even a bit nicer than that, since systemd provides some backward compatibility with sysV-style init scripts [1], but in that systemd acts almost like a swappable implementation.
[1] https://www.freedesktop.org/software/systemd/man/systemd-sys...
Well, but then you must be aware that you are leaving lot of effectiveness, unification and deduplication on the table.
See, many people think of systemd as if it was init replacement that grew timer support too. They are looking at the wrong level of abstraction: it is an event machine, that manipulates units when events happen. That event might be boot to certain runlevel, timer, hotlplug, incoming request on socket, dbus activation, whatever.
By cutting and dragging out one of the event handlers outside into special-cased, custom flow component that duplicates existing functionality anyway, you are ending up with a worse system.
But this still makes no sense. You can just develop another thing that converts from the systemd format to something else, or make the other daemon capable of loading the systemd format directly. It isn't even hard to do this, I've seen various things do it. And depending on where you are the systemd formats are arguably more standardized than old crons and sysloggers in the current year.
Plus, it creates fragmentation instead of concentrating efforts on one project.
It's easy to get situations like "I need feature F but that's only in B, we use A because C needs it".
Use of any features outside the lowest common denominator quickly makes things nonswappable.
I saw he was commenting in the thread on this page and I said "Hey we're in the same thread!"
Then some other HN joker created an account called andrewstuart3 and joined in the thread.
Now you know!
I do not know what customkitchen is.
No, Systemd is anti-Unix and unnecessary. Its demand never comes from the community and many non-systemd distributions work much better than redhat, like alpine/gentoo/Slackware.
[0]: dbus normally is used over a unix socket but afaik you can use dbus over a regular tcp socket
If there is a network "gRPC server" as well, I suspect it would be somewhere in the 20+ department.
I don't anticipate exposes the actual pid 1 over a network. I'm not a monster. I suspect there will be be an init/jailer mechanism that manages bringing some of the basics online such a system logger and any kernel services (EG: ZFS) right away. One of the first "services" would be a d-bus alternative that is written in Rust and leverages gRPC.
The main motivation behind gRPC is cloud, mTLS, and the support in rust. It comes with the ability to implement load balancing and connection error management/retry capabilities. I have week opinions on the technical detail as I don't suspect the network traffic will be very large. gRPC is more familiar for folks in cloud, as well as supports a large number of client languages for generating clients.
https://en.wikipedia.org/wiki/Beowulf_cluster
https://en.wikipedia.org/wiki/MOSIX
That is the thing with repeating history.
...and they also create lock-in wrt k8s.
Actually said company makes packaging tools and Linux packages are fairly simple so we started with the ability to make server packages that use systemd.
That support is still there and makes it fairly easy to build a signed apt repository from build system outputs (e.g. jars, native binaries), or even by re-packaging an existing server download [1]. What you get out is debs that start up the service on install, ensure it starts automatically at boot, restarts it across upgrades and shuts it down cleanly if the machine shuts down. Such packages can be built from any machine including Windows and macOS. The nice thing about doing packaging this way is you can specify dependencies (both install and service startup) on things like postgresql, and you get a template nginx/apache2 reverse proxy config as well. Unfortunately there doesn't seem to be a way to make those templates 100% usable out of the box due to limits in their config languages, but it's still nice to have.
There's also pre-canned config snippets for sandboxing, setting up cron jobs and other useful systemd features [2]. We package and run all our servers this way. One big rented colo machine is sufficient right now. If there was a need for clustering we'd just point the machines at the same apt repo and use some shell scripting or Ansible, I think.
There's lots of tools for making Docker containers out there already, but I never really clicked with Docker for various reasons. Conveyor is mostly focused on desktop apps but perhaps there's more demand for easily making systemd using packages than I thought. The server side support is there and we dogfood it, but isn't really advertised and it lacks a tutorial.
If someone is building say an embedded system that has to run different services (say a NAS, or a router), it's pretty compelling alternative to a bit more heavy weight containers like docker.
Many sandboxing features require a relatively modern systemd and will do nothing on older distros (we run RHEL9).
Also like the "negative" commented out config, that's something I often do myself with footguns that feel "obvious".
That was still early days for systemd-nspawn, folks would be using docker containers under systemd services. And that caused heaps of other problems because docker was aspiring to do PID 1's jobs as well.
But does systemd do these things in a user-friendly way? Or at least as friendly as Kubernetes?
Unfortunately, no.
Don't underestimate the importance of UX.
No one should ever confuse kubernetes for user-friendly.
Kubernetes is user-friendly, specially when compared to each and any of its alternatives.
And moreso when compared with systemd.
We live in a day and age where it's possible to get a whole web app up and running in a Kubernetes cluster from a fresh Ubuntu install with a single snap installation and a kubectl apply -k <kustomize dir>. How long would it take to get systemd to containerize a single app?
Anyway, you're right in the general case, but specific cases really depend on how complicated your app actually is. If we're allowed to introduce extra tools then gosh darn it I'm going to promote Conveyor again because in that case it gets easier to use systemd too. Here's what a server config looks like:
https://gist.github.com/mikehearn/5485a7343d9fe838d33d0b0281...
All of ~25 lines, some of which is just optional demo stuff. To use it you'd compile your app (a JVM app in this case), run "conveyor make debian-package", upload the resulting package and install it with "apt install ./whatever.deb". Or alternatively upload it to a static file server (s3 bucket or whatever) and then run "apt-get update && apt-get upgrade" on each host. The server will start/restart automatically. It's not containerized in the Docker sense but it does run in a lightweight sandbox using the DynamicUser feature, and you can lock it down further if you want by setting the right systemd keys.
Now, you're going to say that Kubernetes does a lot more for you, that it can deploy many kinds of apps simultaneously, configure networking, let replicas find each other etc and so it's easier to use for 'real' apps. Granted, all true. The above workflow isn't optimized, does less and would need more work to be competitive. Also the resulting packages assume there's an apt repository somewhere, which takes a bit of work to set up (this isn't strictly needed in the server case and we'll fix it at some point).
Still, whilst maybe it's better for everyone to just learn Kubernetes at some point, but a lot of us have learned Linux/UNIX in the past and have needs that can be met cheaply with that toolset.
But I was then quite surprised a few years later when it was just the normal thing for me when people in Debian were basically fighting a war over it.
I dislike its "let's just reinvent features that worked entirely fine" (like journalctl binary logs WITH NO FUCKING INDEXING, so they are slower than text files for usual operation somehow), but at its core competence it's extremely useful.
Seems to be quite serviceable on the happy path, and I quite like it for desktop linux honestly.
But I find it's usually a bit too opaque and "magic" for servers; even if it's causing less and less issues.
I liken the issue to the same one pulseaudio had: a software ecosystem that is poorly designed but foisted together into a perfectly serviceable product by incredible amounts of effort for all involved in its development and release to public.
I feel the same way about systemd, its safe, reliable, and always a good choice for "dinner".
Basically I am saying that systemd has withstood the test of time and has never disappointed me.
> Basically I am saying that systemd has withstood the test of time and has never disappointed me.
IMO, when you make a reference specifically to chicken tenders, there is an implication of childishness - implying that systemd is a bit of a toy implementation. It doesn't sound like you mean that particular interpretation.
It would be the same if you mentioned red jello, milk boxes, capris suns, fruit snacks, or any other stereotypically children's food.
Lack of Unix philosophy
[...]
Missing rest/gRPC/json API
OK ...Edit after further reflection: Alternatively, perhaps both of those complaints could be summarized as poor ability to interoperate with other tooling, in which case they really do fit together.
[1] https://www.freedesktop.org/software/systemd/man/org.freedes...
systemctl --host=whatever.abc status apache2
It uses ssh and UNIX domain socket forwarding behind the scenes. There's a writeup of how to do this with minimal privs here:https://sleeplessbeastie.eu/2021/03/03/how-to-manage-systemd...
If you want to speak the protocol without the systemctl program then you can do the same trick. Use an SSH library, connect to the dbus socket and connect to systemd. You do need a dbus library though. They are less common than HTTP stacks.
libsystemd includes a d-bus library for exactly this reason.
I don't think "Unix philosophy" means "communicate via streams" in this context. I think it means things like the single responsibility principle. Once we introduce APIs, we can't just chain services together with pipes.
But I agree that this use of "Unix philosophy" is confusing.
I think the original sound-byte I was trying to capture was "do one thing" which in my opinion neither Kubernetes nor Systemd do. To be fair -- neither would Aurae. So I just scrapped the entire comment.
I wasn't trying to nitpick systemd as much as I was trying to draw attention to the fact that it does in fact -- get nit picked -- and often unnecessarily.
Yeah, right.
so, I say yes
Complex systems are hard to replace, thus stay in place.
This is the way of all things unless there is a design constraint placed on simplicity. In mechanical engineering they design things to be cheap to manufacture: this is the design constraint.
Computers do not have constraints on them, you can burn resources or be as complex as you want: more powerful hardware is around the corner, and after all: why shouldn't abstractions be abstracted ad infinitum; as long as it's easier for people?
Damn, this is very good and clarifying. Like Gresham's law, but for software.
> why shouldn't abstractions be abstracted ad infinitum; as long as it's easier for people?
Not pessimistic enough. Should be: Even if it's no easier on anyone
VM is essentially just the last part and usually simpler too ("just connect it to switch like it is separate machine" is common solution"). Sure you still have some complexity around storage but that's just "here is a bunch of blocks", no need to make overlay filesystem or mount image layers.
Then we are at level of generic OS. The features docker/systemd/k8s uses were not "designed to make containers", they were designed to be used for anything you wanted, from as complex as containers to as simple as "let's put this processes in a cgroup so we can limit memory usage together. And with flexibility comes complexity.
Then we have a problem of both systemd and k8s/docker managing same interfaces so both have to play nice with eachother. And we get to even more complexity when kubelet-controlled containers also need to manage system stuff, like say networking via kube-router.
It would be simpler if say k8s' kubelet directly managed everything all at once but that means it would also have to do what systemd does (boot system, setup devices and partitions etc.) so while total complexity would be lower, the kubelet itself would be more complex.
Assume you have a $25k budget for a Dell server and $20k for attached storage . Go see what you can build on Dell.com's online configuration tool.
You could colocate everything you can buy for $45k in a half rack at a reputable datacenter for $750 per month including a burstable to 1gbit internet connection.
I estimate that at least 75% of the people who are commenting on these threads, will not have a product or service that could cause that server to choke on the load, nor saturate the 1gbit feed with legitimate traffic from their applications.
But they will instead recommend that for 600k/year in spend you over-engineer all aspects... I call it CVops instead of DevOps. CV being another word for resume...
Then again we got a bunch of racks and our ops dept is 3 people.
We have few dozen different apps running on it, anything from "just a wordpress" to k8s cluster (mostly because our customers want us to deploy app they are paying us to develop on k8s).
So far the only actual value k8s provides is self-service aspect of it. Nothing that is running on it needs anything special to k8s, all of it is just "few app servers + one or few DB services".
Sure, there is value in not bothering ops to install yet another VM and plumb the network for it, but you need pretty big team for that kind of savings to be worth it.
There is also value in having complex app deployment centralized in one manifest that can be run on prod or on dev machine, but you can you know... not overcomplicate your apps massively by making a bunch of microservices for a thing that really should be just well structured single app. Again, not really a benefit for small (let's say below 50) teams, for bigger orgs, sure
> But they will instead recommend that for 600k/year in spend you over-engineer all aspects... I call it CVops instead of DevOps. CV being another word for resume...
It's a mix of bad goals and devs wanting to work on cool stuff. I did mildly overengineer a lot of things but about 8/10 out of them eventually became useful, but that's because I'm in ops so everything we do is long term by default.
AWS skills are as arcane or more than traditional hardware management skills; I’m not sure I could make the argument in good faith that you need less people overall with cloud.
An (anecdotal) example: I worked on an e-commerce SaaS platform responsible for 1% of internet traffic in 2011 with a team of 6 sysadmins, with 60~ developers working purely on the product.
My last company had 23~ “SREs” (that did not code aside from DSLs) operations staff handling AWS deployments.
And it always ends with a bunch of people doing essentially same thing but with different job title.
> AWS skills are as arcane or more than traditional hardware management skills; I’m not sure I could make the argument in good faith that you need less people overall with cloud.
The main difference is that it is impossible to dig deep. You can do the deep dive, almost to the bare hardware level when you own the hardware. If you just call a bunch of APIs you're essentially talking with black box.
Like, I do not exactly miss the time spent finding out why our Ceph cluster misbehaves and getting to some driver bug causing ~0.5-1% packet drop after few weeks running with irqbalance daemon on, but it is possible and it can be mitigated, meanwhile in cloud, black box, tough shit, live with it.
I might borrow this, although I'm leaning towards Res[ume]Ops
Here's Tom Preston-Werner's "Readme Driven Development" in 2010:
https://news.ycombinator.com/item?id=17427593
https://tom.preston-werner.com/2010/08/23/readme-driven-deve...
// Same energy as Amazon 'Working Backwards':
https://www.allthingsdistributed.com/2006/11/working_backwar...
This article was scratching different itches than the ones that bother me. I'm not interested in Kubernetes, or any other enterprise cloud stuff for that matter. I don't need orchestration, my hardware never changes (so I don't need anything like UDEV), and if there ever is an ad-hoc change I can ad-hoc configure it.
You don't own any USB peripherals?
If we do this "right" it kind of needs to go at the pace of the Kernel, and hold very tightly onto the API scope to protect the project from scope creep.
It looks like the Aurae language leverages Rhai? Or is that temporary?
In "Aurae Scope", would logging be in scope as well?
Do you think Aurae would be a nice successor to docker swarm, or is that not a product goal?
Here's how to fix Kubernetes and systemd: make a non-sucky build system and then use it to build systems.
Systemd builds a running system. Kubernetes builds a distributed system. Make a build system that can build both. Then you're golden.
And it will also not be monolithic (i.e., it will follow the Unix philosophy) because a build system just hooks together smaller tools.
Yes, the build system needs to be able to respond to events, but that's a simple extension.
Disclaimer: I'm currently building such a build system. I wish the author luck, but I am a competitor.
Edit: I might have sounded snarky. I'm not trying to be; it's just late at night. The author has a lot of good ideas; I just think that the ideas could be followed just a little further to something great.
how to fix this "most-open unit of computing" will depend on "what" do you think this 'unit of compute' even means, which of course depends on who you are and what do you do with 'units of compute' (buy? sell? resell? use up? build? oversee??)
That is sort of what we're talking about when we talk about unix philosophy. If your daemons for handling all those aren't tightly coupled then it's easier to use the same tools for different tasks.
I'm a bit confused by the early complaints though.
> It assumes there is a user in userspace. (EG: Exposing D-Bus SSH/TCP)
What does this mean?
> Binary logs can be difficult to manage in the event the system is broken and systemd can no longer start.
What kind of managing? To read them you can use --directory to read a specific set of logs from a recovery system and systemd doesn't need to run. What else am I missing?
I think there’s a few things in here that I completely unashamedly love.
Ipv6 by default too was something I assumed would make kubernetes easier (since no double NAT traversal).
Since you can’t make a comment on hacker news without having something critical to say: I wish they hadn’t chosen discord for their chat platform; when things like zulip exist for free and are used already by the rust community.
Like, it appears to target doing same thing systemd does but with different APIs which is like... okay ? Sure, some consistency on one side, on other it is now entirely dissimilar from anything that looks like normal Linux userspace.
I could introduce my own load balancer proxy like Envoy but I want to reduce complexity as much as possible for systemd to be a viable alternative to k8s.
Also iptables would be a simple option for a round-robin or random load balancer. No health checks or anything though. See: https://scalingo.com/blog/iptables for an example.
http://0pointer.de/blog/projects/socket-activation.html
I've done it with Python it's very easy.
systemd listens on a port and when a connection request comes in it starts your service. You could spawn a new process and let systemd keep waiting for new inbound connections.
My dream is VMs as easy to manage (build, deploy, etc.) as containers, hosting containers, and managed across a cluster of hardware by an orchestration system. Trying to make container orchestration do the things VMs do with ease produces nothing but pain and waste.
kris-nova: yes, fully integrate VM management and fill this yawning chasm. There is enormous demand for this, naysayers notwithstanding.
It's similar to how --machine works for containers under systemd, but using ssh to communicate.
On one hand I over-engineer a systemd hypervisor that is only meaningful to me. On the hand I create another ambiguous junk drawer that is meaningless without a team of experts to tell you how to configure everything.
I think having what kubernetes calls "namespaces" as an isolation boundary on each node running as a VM is the move here. It SHOULD run like this as a default. Pods are another story. Namespaces however -- should always have a VM boundary.
Getting the network device integration is going to be a big thing here. I suspect this means each namespace now has 1 or more NICs it will be able to leverage.
Firecracker went with the bridge mentality which I kind of disagree with: https://github.com/firecracker-microvm/firecracker/blob/main...
I want to see tools like Tailscale that leverage network devices as the "true network interface" find value in the guest namespace paradigm.
Hope this helps!
In an ideal world, where virtualization has no performance penalty, it might make sense to wrap everything in VMs but in the real world I think having the option to switch isolation mechanisms might be the best idea.
Some may need "better" (subjective) security and opt for VMs which could be the default platform. Others may be fine with something more lax like gvisor or even just having different users for each namespace.
I just think of it as a node API more than anything. Having a comprehensive set of features/library/API for the node seems like it would unlock a lot of features we are seeing in large service mesh and large platform shops are turning to sidecars to solve.
Neat project.
"Problems with Systemd"
Seems pretty unserious. Vague reference to Erik Raymond the Unix Philosophy (actually very debatable and not a useful way to think about Linux or Unix).
Links to pretty useless and unserious anti systemd posts (skimming them was a waste of time).
"""I personally don’t have a deep seeded loathing for systemd like many of my peers."""
Author should definitely get out of their local slack channel more, it not the result of a plot by Redhat and Sekrit powers that systemd is everywhere.
If the author is going to replace systemd they'd be wise to understand why it works the way it does and the epic amount of problems it solved first.
The cognitive dissonance of systemd concerns expressed and then somehow merging with some odd mutant kubernetes speaks for itself....