Systemd as tragedy
lwn.net
lwn.net
Two years later, you can find hundreds of support requests across the internet, from frustrated users who are having their sessions killed by systemd.
Bugs are annoying, but that's life. On the other hand, when you're an impacted user who's lost work, and researching the bug leads you to a years-old discussion in which someone is actively denying that the bug exists and refusing to fix it, that's infuriating. I don't think systemd's developers deserve the trust that maintaining a core piece of infrastructure requires; they don't seem to care enough about whether they've broken things.
Now I am used to taking blame for apparently everything that every went wrong on Linux, but you might as well blame your downstream distros for this as much you want to blame us upstream about this, as it's up to them to pick the right compile-time options matching their userbase and requirements in compatibility, and if they didn't do that to your liking, then maybe you should complain to them first.
(And yes, I still consider it a weakness of UNIX that "logout" doesn't really mean "logout", but just "maybe, please, if you'd be so kind, i'd like to exit, but not quite". I mean, that's not how you build a secure system. We fixed that really, fully knowing it would depart from UNIX tradition, but that's why we made it both compile-time and runtime configurable)
(Also, nobody has to "incorporate" systemd's library to avoid the automatic clean-up. In fact, there's no library we provide that could do that. What was requested though is to either run things as child of systemd --user or just register a separate PAM session, neither of which requires any systemd-specific library.)
Lennart
It's up to you as a systemd developer to pick sane defaults. Claiming that it's okay to introduce opt-out breaking changes upstream and then abdicate responsibility is a quite bit like walking around while waving your hands and arms around and then blaming whoever you hit for walking into you.
IOW the distros maintainers made a mistake by picking systemd? Agreed.
IMO, it is part and parcel of designing great software that you pick as universally agreeable defaults as possible.
That you can configure systemd to behave in a less obnoxious manner is well beside the point. Systemd should be unobtrusive and predictable without any extra action on the part of the distribution folks or end users.
That the suggestion is to simply read the code or documentation is the height of arrogance considering how sloppy and insecure the systemd code is (parse error equals root privileges? come on…).
And sometimes security requires breaking compatibility.
The options for fixing the bug are:
* nohup, tmux, emacs, etc all take dependencies on systemd and use the new systemd daemonization procedure. This is not a viable path because the maintainers of those utilities have refused (see https://github.com/tmux/tmux/issues/428), and because there are too many of them.
* Each distro separately works around the problem by maintaining forks of nohup, tmux, etc. This is not a viable solution because it's way too many forks; people will be finding broken distro+utility pairs forever.
* Each distro separately works around the problem by putting loginctl enable-linger in /etc/profile and KillUserProcesses=no. This would effectively be overruling a systemd's decision. Some distros won't know they need to do this, and the github systemd repo becomes a trap.
* Or: systemd backs down and changes the defaults so that the old daemonization APIs work again.
If you have a fifth option, we'd all love to hear it. But the status quo is that there's a user-facing bug, and the bug is still there. Rather than make the case for it not being a bug, you're currently making the case for it being someone else's bug, but the "someone else" doesn't actually have the power to fix it. You are the only one with the power to fix this bug.
Replace systemd with something else.
As an aside this is the height of arrogance to suggest that the systemd is somehow a more secure alternative. Lest this be considered an empty ad hominem attack, let me quote the pwnie you won in 2017[1]:
> Where you are dereferencing null pointers, or writing out
> of bounds, or not supporting fully qualified domain names,
> or giving root privileges to any user whose name begins with
> a number, there's no chance that the CVE number will
> referenced in either the change log or the commit message.
> But CVEs aren't really our currency any more, and only the
> lamest of vendors gets a Pwnie!
https://github.com/systemd/systemd/issues/6237
oh my god, what a spectacular issue. And, seriously, the Poetterings' response is basically "not my job" and "not a bug". And this person develops something that sits at the core of a modern linux system...
All the while Lennart claims that he's making Linux more secure. FFS.
Edit: I forgot about this
https://igurublog.wordpress.com/2014/04/03/tso-and-linus-and...
> He (Theodore Ts’o) goes on to describe how he previously had to neuter policykit’s security (rendering his system very vulnerable) just to get his system working, and how he has found systemd "very difficult sometimes to figure out".
And:
> As for Kay Sievers, maybe he should rename himself to Kay Sewers, because that’s exactly what he smells of. He told to IETF internet area director and previously DHCP working group co-chair “Tod Lemon” to lmgtfy when he asked about a systemd related git repository.
This gem sums it up perfectly though:
> Yet just two days ago, we see Linus Torvalds (the creator of Linux and maintainer of the Linux kernel), launching into a tirade against – yes, you guessed it – systemd developers because of their atrocious response to a bug in systemd that is crashing the kernel and preventing it from being debugged. Linus is so upset with systemd developer Kay Sievers (gee, where I have heard that name before – oh, that’s right, he’s the moron who refused to fix udev problems) that Linus is threatening to refuse any further contributions from this Red Hat developer, not just because of this bug, but because of a pattern of this behavior – a problem for Kay because Red Hat is also foaming at the mouth to have their kernel-based, no doubt bug- and security-flaw-ridden D-Bus implementation included in our kernels. Other developers were so peeved that they suggested simply triggering a kernel panic and halting the system when systemd is so much as detected in use.
As to the stuff mentioned in the pwnie. Those sound like great contributions that would be appreciated.
You could also take your concerns to the distro development group. If that doesn't work you could also customize your distro with a custom build of systemd.
If you still don't get satisfaction you can stop using it.
If you dislike how they do thing you have options. Or, you could just be mean on a forum...
When I switch distro, it's almost always systemd, and not the system du jour, so I know how it works. Creating service files is a google query away, and makes common use cases a breathe, while advanced features that were hard to bash script yourself into, are now just a few options to type.
I understand that many people may have problems with systemd for their particular situation, but that's not my experience.
As a dumb user with a few laptops and servers that needs an occassional daemon, I'm glad systemd won. I know you get a lot of heat since it came out, so thank you for working on it.
What is not as good: (1) systemd takes over or duplicates functionality not related directly to its primary purpose, and (2) is not solid enough to trust it in a number of cases, while (3) the developers' attitude does not give a lot of hope that the situation will materially improve.
(Of course, I run a distro without systemd.)
Ok, but UNIX and it's behaviour has evolved over forty years, and users have a certain set of expectations about it.
Also, it should be noted, systems like UNIX are cultural artifacts. The way they are is the result of forty years of back and forth debate and negotiation and eventually compromise.
I can't speak for all of them, but I think that people that are bothered by systemd are upset that all of history has been brushed aside to make place for the preferences of just a few influential developers.
Whether a feature like logout is "logical" or not, is besides the point. Operating system design isn't just about logic, it's about serving users.
You build your software the way you want and like. If others don’t like that it breaks POSIX they should stop using it instead of complaining. Or fork it.
When you run your screen or tmux below `systemd --user`, you still would have to `loginctl enable-linger`, no? I remember having to do that when I set up a PulseAudio server on a headless machine where I don't maintain an active session.
so, unix has been running for 20+ years laden with this security flaw? strange that nobody has been screaming out to plug it all this time.
this feels like you have a bee in your bonnet that it is not a very 'pure' logout by some interpretation of what a "logout" should be. imho, "logout" should mean what it has always meant in the past.
This is standard from you. You knock the glass on the floor and blame the maid service for not cleaning up after you.
It's everyone's faults but yours.
>And yes, I still consider it a weakness of UNIX that "logout" doesn't really mean "logout", but just "maybe, please, if you'd be so kind, i'd like to exit, but not quite".
Oh how hyperbolic. Nuances and caveats in terminology is not a weakness.
I don't see why you're splitting hairs over this but can't be bothered to care about your UID numbering bug.
Or he fact systemd-resolv is responsible for DNS leaking on VPNs.
But yes, tell me more about how a functionality that enables terminal multiplexes is a "weakness"
>Now I am used to taking blame for apparently everything that every went wrong on Linux,
It's because of your smarmy, arrogance.
You break POSIX compliance, which has a real world effect in multiple areas and you accept bug reports with the humility of Donald Trump being interviewed by MSNBC.
Then when you retreat into your safe space, you play victim to the situation you created.
You talk of Linux culture toxicity, smearing the likes of Linus Torvalds, while essentially being the metaphorical sibling putting your finger in people's face repeating "I'm not touching you" over and over. Then you acted attacked when someone claps back.
You're a cry bully hiding behind a vaneer of professionalism acceptable for Red Hat's HR department which enables you to mark one more bug as "wontfix"; your attitude, your arrogance, your conceits that things not broken in fact, are so you can provide solutions no one asked for and no one benefits from.
You're just kind of yelling, and it diminishes any point you may have made.
What makes it worse is that he's often not completely wrong. Linux did need something like PulseAudio, something like Avahi and something like systemd. But his reach exceeds his grasp (which probably applies to us all, as I've found on my own projects), which leads to the well-known problems of PulseAudio & systemd.
I don't actually want him to quit the Linux world. But I wish he would scale back his ambitions just a tad, and consider that maybe — just maybe — other people have some good points, and valid concerns.
And also Windows/DOS are not terribly good design exemplars.
if which loginctl > /dev/null && loginctl >& /dev/null; then
if loginctl show-user | grep KillUserProcesses | grep -q yes; then
echo "systemd is set to kill user processes on logoff"
echo "This will break screen, tmux, emacs --daemon, nohup, etc"
echo "Tell the sysadmin to set KillUserProcesses=no in /etc/systemd/login.conf"
fi
fiTurning this on to true, for me it does no make sense to a user service (yeah, I run emacs as a user's systemd service) to keep running after I logout of my system.
P.S.: And the fact that for some people this behavior makes sense is why I think Lenart decision to put this as an option makes sense.
POSIX is nice, but rather lacking in certain aspects, such as security anf administration-friendliness. cgroups help with both, but people have to understand them and use them well.
1. Edit: Not literally closed by the reporter. Lennart Poettering closed it, "closed by the reporter" as in "the issue was resolved to the reporter's satisfaction".
Are we reading the same bug report? The one I'm looking at was closed by the creator of Systemd.
How else do you propose to make sure that when I log off my ssh-agent is really terminated and not just locked up with my keys still in memory? The POSIX approach is insufficient, there's no way to know if a process received a signal and chose to ignore it and keep running or if it received a signal but it was deadlocked and kept running.
If you're not going evaluate each individual program to determine whether the new behavior is appropriate then it should be opt-in rather than opt-out. Then ssh-agent and anything else that knows it should be forcefully killed can opt-in without breaking other innocent programs.
I think some people sometimes lack any perspective on the topic.
I’m not being emotional about it, just irritated.
Systemd has tangibly caused me to lose work with tmux; I appreciate there are root causes for this, but frankly, if some piece of someone’s code does that, for whatever reason that is beyond my control to immediately stop using it...
...it feels justified to be annoyed.
How do you suggest an alternative meaningful response would look?
Create my own distribution?
What tangible and meaningful alternatives do I have other than encouraging people not to use systemd?
Apparently you think Linus is one of those who "lack perspective"?
http://lkml.iu.edu/hypermail/linux/kernel/1711.2/01701.html
I get that systemd isn't the kernel, but it's close enough. There are many who would agree that breaking existing behavior in the name of security isn't wise. I have also not yet seen anyone point out specific security issues this solved. Unix has worked this way for a long time.
or better yet, read the release notes, it likely mentions this breaking change. (if not, that's a bug.)
Breaking compatibility is generally avoided to the utmost. Even security-sensitive things like TLS continue to support older, less secure versions to retain compatibility with peers that haven't been upgraded yet, much to the chagrin of everyone when they screw up the version negotiation, but better than the chicken and egg problem where nobody can upgrade until everybody has.
But the other point is that the claimed security improvement doesn't actually seem to be there in this case. They haven't made it so you can't have a program continue to run after the end of the current session, they've only changed what you have to do to make that happen, thereby breaking everything that did it the traditional way.
Perhaps with a signal handler?
I disagree that POSIX says that processes should expect a SIGHUP when a user logs out (SIGHUP means the controlling terminal was closed). I am not at all a POSIX expert, so please correct me if I misunderstand, but afaict POSIX explicitly does not specify what happens to the controlling terminal when a user logs out (http://pubs.opengroup.org/onlinepubs/9699919799/xrat/V4_xbd_...):
> POSIX.1 does not specify how controlling terminal access is affected by a user logging out (that is, by a controlling process terminating). 4.2 BSD uses the vhangup() function to prevent any access to the controlling terminal through file descriptors opened prior to logout. System V does not prevent controlling terminal access through file descriptors opened prior to logout (except for the case of the special file, /dev/tty). Some implementations choose to make processes immune from job control after logout (that is, such processes are always treated as if in the foreground); other implementations continue to enforce foreground/background checks after logout. Therefore, a Conforming POSIX.1 Application should not attempt to access the controlling terminal after logout since such access is unreliable. If an implementation chooses to deny access to a controlling terminal after its controlling process exits, POSIX.1 requires a certain type of behavior (see Controlling Terminal ).
See the enable-linger option for loginctl and KillUserProcesses for logind.conf. KillUserProcesses was set to default enabled on 4/9/2016, prior to that it didn't happen, but was configurable if desired. So you were always able to change the config to restore the previous behavior from the moment the default turned it on.
Edit:
Here is the commit where it happened
https://github.com/systemd/systemd/commit/97e5530cf2076a2b4f...
No, you were not.
The thing that people are missing here is that neither of the systemd-logind behaviours, with KillUserProcesses=yes or KillUserProcesses=no, is the long-standing behaviour of kernel login sessions all of the way back to 7th Edition that nohup, tmux, screen, emacs --daemon, mosh-server, deluged, and more all interoperate with.
The behaviour of kernel login sessions is that end of login session is a HUP signal to the session leader, and that termination of the entire TTY login service (such as at system shutdown) is a TERM signal to everything followed by a KILL signal to everything then remaining.
The systemd-logind session behaviour with KillUserProcesses=no is no signals at all at the end of the login session, and at termination of the TTY login service both HUP and TERM signals together then KILL signals, to everything.
The systemd-logind session behaviour with KillUserProcesses=yes is both HUP and TERM signals together then KILL signals, to everything, both at login session termination and at TTY login service stop.
As I pointed out years ago, the fix is to make systemd-logind use KillUnit at hangup and StopUnit at service termination, actually providing the conventional behaviour which it currently does not in any mode and addressing the original problems (with some background GNOME utilities in a login session that were never being sent a HUP signal at logout and would have exited had they been) that motivated this whole mechanism in the first place.
* https://news.ycombinator.com/item?id=12335128
* https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=825394#221
So how can it be the default?
This is why we have distro vendors, to build a system that works in the real world with software from developers with opinions that... differ to say the least.
Systemd is basically SMF, done poorly, because NIH.
* https://freedesktop.org/software/systemd/man/daemon.html
IBM was explaining what to do back in 1995.
* http://jdebp.eu./FGA/unix-daemon-design-mistakes-to-avoid.ht...
killing user processes on logout
By "killing", do you mean some other signal than (or in addition to) SIGHUP? Does it send SIGKILL?* https://news.ycombinator.com/item?id=12335128
* https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=825394#221
TIL what nohup(1) is for.
By "use systemd's new demonization API" you mean, instead of
$ screen
systemd asks you to write
$ systemd-run --scope --user screen
instead. Annoying to have to learn a new thing, but hardly the unbearable burden.
On the other hand, when you're an impacted user who's lost work, and researching the bug leads you to a years-old discussion in which someone is actively denying that the bug exists and refusing to fix it, that's infuriating.
Because it's a bug for some, and intended behavior for others. Look, you make it as if they introduced a bug on purpose to screw with some people. It's clearly not the case, there was a specific tradeoff involved.
They broke userland.
It doesn't matter what tradeoff they made - they went against POSIX behaviour, and as a result, broke numerous utilities, both past and future.
Let's say that again - systemd introduced breaking behaviour on userland, against POSIX, and instead of backing down and allowing for expected and specified behaviour, they said it's everyone else's problem.
That is neither professional, nor responsible.
When you make a mistake, a mistake that breaks the behaviour of POSIX, and POSIX utilities like _cron_, you apologise, and fix the problem.
You don't turn around and say that all the sysutils should incorporate your new idea.
Moreover, this doesn't affect cron at all. Cron creates its own PAM session for each job it runs which means those jobs are independent from any real login session (i.e. ssh, graphical, tty login), and thus also don't get cleaned up by them.
This affected stuff that is forked off a login session and then stays around as "orphan" if you so will, i.e. with all session resources released, except for these processes that try hard to avoid clean-up (usually by double forking + detaching explicitly from any TTY/ignoring SIGHUP).
I tend to agree with the idea that the choice of defaults belongs to the distro's. If the distro's are deferring to the upstream project on default settings for a critical system component then they need to be more thorough and validate what they are shipping.
That alludes to kernel development, which systemd is largely uninvolved with. A userland program chosen by various distributions failed to support conventions from a different userland program. That's all. Were the programs involved fundamental and highly important to many users' experience? Sure. Is busting out "you broke userland" like some magical shibboleth useful as a means of your conveying your unhappiness that your distribution maintainers chose to replace a widely-depended-upon program with a different program useful? I think not.
> they went against POSIX behaviour
Which? There's "tradition" and "specified behaviour". Both are important in different situations and in different degrees.
> You don't turn around and say that all the sysutils should incorporate your new idea.
Why not? They're no more privileged by the POSIX specification, or by the user/kernel -space divide than any other program.
Intel, the kernel, even Chrome broke my userland by mitigating Spectre.
It happens.
CRON was and is run as a system service, in its own scope. If you run your own cron instance, but forgot to set it up as a system service, yeah, it gets cleaned up as you exit your shell/session/scope.
So? "We don't break userland" is a Linux kernel thing. Systemd is not kernel, it's userland, and userland things break other userland things all the time. They already broke lots of existing stuff when they replaced /etc/init.d/ scripts with systemd definition files, should systemd also have not done that?
> It doesn't matter what tradeoff they made - they went against POSIX behaviour, and as a result, broke numerous utilities, both past and future.
Linux is not POSIX, so I don't see how that's relevant. For what it's worth, I don't even know what part of POSIX it broke. Care to enlighten me?
If you ruin everyone else's day, and change behaviour everyone else is expecting, then it's probably your own fault.
Approaching it as if everyone should simply change and do what you want, is the height of arrogance. You are generating work for others. And in this particular case, not only are you generating work for others, you are eradicating a category of software.
When a distribution adopts systemd, they let everyone know how things are changing, and slowly transition things over, releasing when stable.
We know systemd replaces init.d. It was difficult, but distributions using systemd got over that hurdle, but it did take time.
However, this is not the same.
Yes, systemd is userland, however it is also PID 1. It is a layer between most userland and the kernel, and so needs to reflect the responsibility of it's position.
Ignoring how NOHUP is supposed to be interpreted, is a _bad idea_, and yes, a violation of POSIX, specifically signals (SIGHUP and nohup), and how they are supposed to be handled.
Moreso, it greatly heightens the difficulty of many utilities that are expected to work.
Why should cron (all implementations of cron), suddenly need to rely on another userland library to maintain it's function?
You just broke most Linux automation. Across an entire industry.
Why should screen (all implementations of screen), suddenly need to rely on a userland library much bigger than most implementations, to continue it's base function?
You just broke an entire category of background systems - including systems communicating with embedded hardware. You might have caused a factory-floor fault. Which could cause injury, or worse.
A breaking change of this level can cause industry-wide ramifications that are not just limited to the digital. Unexpected behaviour is exceptional, and should take time and considerable thought before occurring.
Systemd has responsibility that no other userland system has. It's PID 1.
If they're going to require a massive change in process behaviour, then they are going to require consultation, awareness within the industry, and transition time. They should be working with distributions, aware of the man-hours they're generating, before they put something in place.
The problem is now your scripts won't work on systems that don't use systemd. Shell scripts work on FreeBSD, but now you can't use them because they require systemd-specific code.
I am not necessarily anti-systemd in most respects (I like a declarative definitions of services and less shell script hell), but the fact that they keep trying to get people (including container runtime developers like myself) to use _their_ API rather than the preexisting ones is fairly "anti-social".
I am not trying to get you to use our APIs. You talking about the cgroups APIs again, if I am not mistaken? As I tried to explain again and again: if you want container runtimes to manage their own cgroups then just set Delegate=yes in the unit file of your manager, get your own cgroup subtree, and you can do below it whatever you want, you do not have to call into systemd ever. Not a single API call, no C call, no D-Bus call, nothing. You get your own kingdom if you set Delegate=yes, and systemd won't interfere with that. This is extensively documented.
I wished you'd actually listen to what I keep repeating to you. We tried to be really nice to container managers, knowing that they disklike systemd APIs, so we put a lot of work in making the delegation boundary clean, so that they can be entirely systemd agnostic beyond setting the Delegate=yes boolean in their unit file, but alas, we just keep hearing the same nonsense.
The LXC/LXD people btw did get this right: they manage their own cgroup subtree now, and systemd doesn't interfere, and they don't link to or do dbus calls into systemd either.
In runc we don't have a dedicated manager or long-running daemon. Yes, Docker and cri-o use Delegate=yes (so I am quite aware of this option) but that really doesn't help people who are using runc in their own user sessions or wrote their own wrapper and aren't aware of Delegate=yes.
I get that we are quite odd, and don't fit into a system-service model. After all of the back-and-forth with both you and Tejun (especially when it comes to "rootless" delegation -- which systemd only offers if you get a privileged user to delegate for you), I'm not sure that there's much I can do on this topic. I get that what I care about is not something you care about, but I would hope you accept that I'm not just being obstinate for the sake of it.
> Not a single API call, no C call, no D-Bus call, nothing.
Right, unless you need to set this up for someone else. And we have code that does this too -- I don't really recommend people use it, but it is necessary (and I'm pretty sure some folks at Red Hat use it based on how many bug reports they submit related to it).
Since systemd is managing the entire cgroupv2 tree (and the fact we can get around that for cgroupv1 appears to be seen as a design flaw by both you and Tejun), obviously we have to talk to systemd to do this type of thing. I just wish this wasn't the way it was done (and if cgroupv2 had a named cgroup concept -- which is what systemd needs for tracking services -- I would think that this wouldn't be such a pain-point).
I guess I'm just annoyed that we can't use "better rlimits" with "rootless" container runtimes because of all of this.
> I wished you'd actually listen to what I keep repeating to you.
I am listening, and I am aware of Delegate=yes and all of that history. But as I outlined above, I don't necessarily agree with it entirely. And unlike a lot of people around here, I don't think any of these pain-points are coming up because of malice or something stupid like that -- I just think we disagree on our priorities.
> We tried to be really nice to container managers, knowing that they disklike systemd APIs, so we put a lot of work in making the delegation boundary clean
Don't get me wrong -- I do appreciate that we have Delegate now (there was a period of several years where "systemd decided to reorganise the cgroup tree, un-containing my containers" happened on several occasions -- and Delegate solved those issues).
And from what I've heard from the LXC folks, you were quite reasonable about getting systemd to work inside LXC. Which is good to hear.
> The LXC/LXD people btw did get this right: they manage their own cgroup subtree now, and systemd doesn't interfere, and they don't link to or do dbus calls into systemd either.
We do basically the same thing. We just don't support cgroupv2.
Rather, a breaking change to everyone's scripts and processes for zero benefit.
EDIT: My reply was supposed to be to xyzzys's post below, not the one I apparently replied to.. sorry about that.
I agree that it might not be the most desirable default, but if that's the case, then the guilt also falls on the distribution maintainers, who either ignored the big bold letters in the changelog, or didn't bother to test the everyone's standard workflows before pushing to stable.
Based on Lennart's behavior, yes I do.
Not to appeal to self-authority, but I have been maintaining production Linux systems in large-scale environments since the late 90s. If there were a benefit that outweighed the unnecessary breaking changes, I would see it, even if I didn't appreciate it. There isn't.
You should stop and think before you assume that other people are incompetent, both because it would make you a better interlocutor, and as a bonus it wouldn't violate HN's principle of charity.
Can we please stop misrepresenting the complaints against systemd? The only time I ever hear this "monolithic binary" argument is from systemd advocates. The actual complaint is about tightly coupling important features together. Not only does this make it difficult (often impossible) to replace individual components, when tight coupling happens at the (internal) protocol level, any replacement component ne4cessarily hast to implement a bunch of (sometimes unwanted) systemd baggage.
Busybox implements all of its features in single monolithic binary, but it isn't a monolithic design that tightly couples those specific components together. Replacing one of busybox's components is often as simple as removing busybox's symlink and installing the replacement. This isn't even a "Unix philosophy" issue. Even inexperienced designers shouldn't have as hard time Understanding why systemd is a monolithic design but busybox isn't.
systemd is a PID 1 program, it means it have to raise bar higher. When troubles begin, you would need tools to fix them, and if PID1 is crashed, you are out of luck. If system cannot boot into shell, you'd need to fix it from initrd shell. Or to boot other system, to fix this one. It sucks.
Linux kernel chases very high standards of reliability, because when kernel panics it is even worse than PID1 crash. Init system should follow the same standards as linux.
The bar is higher for pid 1 - if I were designing systemd I would have made a tiny pid 1 that just did message-passing to a more complex secondary process that could be restarted, or something, just to be safe - but I think systemd has empirically cleared the bar.
3AM, deep slumber, called out to look at a stricken server. Its problems included that systemd was frozen. Reluctantly I came to the conclusion that a restart was the only route forward. Cept, that is when you discover that the commands that have served you well for 2 decades don't work, as they are all wrappers for systemd, which has keeled over.
To this day, the `shutdown` man page, which I was checking in, makes no mention of how to resolve, tho in fairness the other commands (poweroff, halt, init) do. I discovered this after stumbling across https://github.com/systemd/systemd/issues/3282
If you find yourself stuck in the middle of the night, reading through docs to try and figure out how recover a machine with a crashed systemd, then `systemctl reboot -ff` or equivalent is what you are now looking for, the `-ff` being the key to "JUST £&*(ing RESTART THE MACHINE!!!".
Experiences like that, don't win you friends.
If this happened to me today with systemd I'd be up shit creek without a paddle.
EDITs:
there's the classic case of the linux "debug" parameter: https://bugs.freedesktop.org/show_bug.cgi?id=76935
and the even more classic case of firmware loading events: https://lkml.org/lkml/2012/10/3/484
and while "all software has bugs" systemd really has the most annoying bugs (by virtue of trying to do everything core to the system) and always insists that they are features and we are backwards whiny geeks for complaining.
Only recourse has been to reboot the instance from AWS dashboard.
I can’t get to the bottom of it because the tools don’t work when it’s down and there’s nothing there when it comes back up. I am not enjoying boiling to death in this pot of shit.
And then there’s the situation where it just won’t boot. I just fire up a new instance then because it’s easier than debugging it.
No, I have not. But I have seen how systemd gracefully failed to boot system to login, with good looking colorful error message. Something that reminded me "Keyboard is not found. Press F1 to run setup."
But OP was asserting that systemd crashes under normal operation because its pid 1 is too fragile, which is very different. At scale I already expect that there's a chance a machine won't come back if I reboot it - it's annoying if I can't ssh in, but, well, I already lost a disk I care about and it won't return to service and I need to fix it anyway. (And it's an easy fix, just add "nofail" to fstab.) At scale I don't expect init to crash under normal operation.
rm -rf --no-preserve-root /
https://lwn.net/Articles/674940/The configuration files should have set that to read only after boot.
The kernel patch where this was fixed can be found here:
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
* https://news.ycombinator.com/item?id=8384251
But it does happen to other people.
* https://unix.stackexchange.com/questions/440229/
And there was one crash that made the headlines.
(I have been hit by the same issue on my private notebook, but I have procedures in place to cleanly recover from failed upgrades on all systems, so it was not a big deal.)
You see, that's the argument I hear a lot from Systemd advocates. The problem with anecdotal evidence is obvious. When you hear people opposing Systemd, practically all of them have some real-life issues with it, often related to functionality that would otherwise be non-essential (i.e. doesn't really need to be handled by PID 1). Of course if you don't have a particular problem, you don't feel it's important. That's precisely the attitude people resent.
Yes, but a lot of people have real-life issues with it on their desktop of the form "It's too complicated." I'm asking specifically about real-life issues on production servers at scale. There will of course be tools that are poorly suited for a personal machine (even a personal server) but well suited for a team that wants to run a bunch of reliable servers.
For instance I would never be happy running RHEL on my desktop, but that doesn't mean RHEL is useless.
I just count my blessings that runit is widely packaged in every major distro because it can just happily sit on top of sysvinit, systemd, upstart, pretty much any init system and does things in a very simple shell script style, I really wasn't a fan of the weird ini-like format for systemd or several different tools I'm expected to learn just to read my (now binary) log files competently.
If you're sick of switching init systems constantly or don't want to have to write separate scripts for your linux box and your freebsd box even, I highly recommend checking runit out.
I'm sure I'll give it a serious shot eventually... in about 3 years once they work all the Poettering kinks out, just like PulseAudio. They're doing some cool things with cgroups and stuff, so I hope it gets there eventually.
If you've ever tried to use systemd inside docker to bring up a couple of services, you would know the hoops you have to jump through to get it working.
(I understand that docker wasn't invented to run multiple services in the one container, but sometimes it can't be avoided and simplifies app deployment vastly I.e, using CI to test your service actually starts up as per its definition: just run up a quick docker image with runit and a service definition file)
Good luck. If you need anything more than "I play a three minute song" on Linux audio you need both some type of real time kernel and jack.
I think I'd like systemd more if I had more confidence that he had done a better job.
I never really understood why Upstart didn't get more traction; was it just the Canonical backing made other distros avoid it, or did it have drawbacks I never ran across?
(systemd is hardly unique in being dependency-based instead of event-based - but it is a strike against Upstart in particular. Also, while one of the stronger claims of why Upstart should work this way is it matches the evented mindset of udev, systemd can handle events from udev just fine, by just having it add additional targets - and in the end, systemd consumed udev.)
The other was that it was hard to deal with its bugs. I had a job that wasn't activating. I spent days chasing it, including attaching a debugger to Upstart, and eventually gave up. I've used systemd extensively and have not had "Why isn't this service starting" bugs. (I've certainly had that confusion, but I've always been able to figure it out quickly.) systemd has in my experience been reliable and trustworthy. I'm not going to say that it's bug-free, or that I like its implementation, or that I have no complaints about it. I am going to say I haven't been frustrated by it - even when running very old distro versions with lots of known bug fixes upstream. That's very important for whether I'm happy with it or whether I'm going to switch back to sysvinit and find something else for service monitoring.
Systemd also solves the problem of unified ways to start services across many distributions. Packaging things is easier. I don't like the systemd ini target format and wish something else had won out instead. With something like upstart, in theory you could rewrite the upstart daemon and use the same scripts, since the format/standard was much simpler.
I feel like there must have been a lot of bullying to get systemd across so many systems. It just seems weird to see so many major distros all decide to accept something so complex if it wasn't for Redhat cramming it into everything they maintained and helped fund.
> Systemd also solves the problem of unified ways to start services across many distributions. Packaging things is easier.
> It just seems weird to see so many major distros all decide to accept something so complex
You answered your own question. Distributions adopted it because it solved many problems for them.
I followed the Debian discussions when they debated init changing.
It was not pretty, there was a lot of things, but I don't remember any bullying. It was simply a very long uphill battle from a sort of loose group against the majority of maintainers that wished to stop worrying about 99% of init script problems.
Systemd is not perfect, never was, but it gets shit done, whereas no other project does in this regard. (A lot of people wanted to tackle some parts of what systemd does, but that was late and insufficient.)
And, all in all, there is always room for a systemd2 in Rust, distros would switch to it in a heartbeat, if it would be better.
Even though Red Hat is a much larger organization, their contributions have never felt like power plays in the same way that Canonical's have. Perhaps it's just the way that Canonical tries to shove their stuff through while Red Hat seems to get more organic buy-in. For example, as controversial as Systemd was (mainly from sys admins and power users), it did have a level buy-in from distro maintainers and developers which Upstart never really did. Canonical just tried to say 'we're doing this.'
Since when? I remember some particularly contentious things like EGCS and various kernel doodads, and of course systemd.
(If there are distro maintainers / package developers who disagree, I'd be interested to hear about it. Again, this is just my impression)
https://web.archive.org/web/20140928104327/https://plus.goog...
Scott James Remnant
+
4
1
2
1
Reply
+Michael Hasselmann at the point that Kay, Lennart and
I sat down and discussed all this stuff, I don't think
Upstart was perceived as "shitty" at all. We'd had
on/off discussions for ages, but the big one I remember
was the LF Collab Summit in SF in April 2010.
Hindsight certainly lends a different perspective, and
I'd be the first person to say that Upstart doesn't
work as intended. +Lennart Poettering makes a great
point about mountall in a recent post, it was written
because Upstart couldn't do the complex filesystem
cases it was designed to be able to do; and I was very
aware even at the time that was a failure that would
need to be addressed.
Had the CLA not been in place, the result of the LF
Collab discussions would have almost certainly been
contributions of patches from +Kay Sievers and Lennart
(after all, we'd all worked together on things like
udev, and got along) that would have fixed all those
design issues, etc.
But the CLA prevented them from doing that (I won't
sign the CLA myself, which is one reason I don't
contribute since leaving Canonical - so I hold no
grudges here), so history happened differently. After
our April 2010 meeting, Lennart went away and wrote
systemd, which was released in July 2010 if memory
serves.
So I don't think I can claim that the perceived
shittiness of Upstart spawned systemd, because at the
time it wasn't seen that way. I don't think I can even
claim that it provoked Lennart in any way, init was an
area all distributions were fiddling with, so it was
inevitable anyway.
I entirely agree with Kay and +Greg Kroah-Hartman that
it was the CLA that caused systemd to be written
instead of Upstart.
But I don't need that self-affirmation anyway :) I
wrote Upstart, I got paid for it, I moved on to do
other things, something else came along and replaced
it. If Upstart hadn't been under the CLA, and systemd
hadn't've happened, all my code would have long since
been rewritten by now anyway.
That's the nature of the software world, there's no
point getting precious over things. Do your bit, have
fun doing it, move on and let others do their bit,
etc.Personally I’ve been using openbsd quite heavily for personal projects, for work my team transitioned from RHEL6 to FreeBSD- and while it has quirks (mostly on installation of software) it is incredibly stable.
It’s a shame my chosen cloud provider doesn’t treat it like a first class citizen, but that’s fine since the community projects seem to work on making good images.
I still use Linux on my desktop and laptop; since I feel like the design of systemd is more suited to those roles (systemd’s design in general feels like windows service activation) but I’ve been having issues that seem to be related to systemd. So it’s not winning me over.
I've switched to OpenBSD for nearly all personal work. We're still on RHEL/CentOS at work though.
I haven't figured out how to get xmonad running on OpenBSD yet or my workstation at work would be on it too.
FreeBSD is much better at just not screwing everything up (particularly on OS updates). Sometimes I have to manually configure a new piece of hardware, but once I configure it it stays configured, and the way to configure it doesn't change from version to version.
On my new laptop I just didn't bother replacing windows. With "Bash on Windows" I can run all the unix programs I wanted to, but windows is handling all the session-management type stuff that systemd would do. I've found the system more reliable and better at responding to hardware changes. As much as there are horror stories of windows update breaking everything, I haven't experienced that myself (whereas I have had systemd updates leave systems non-booting).
1. They changed the defaults for systemd-logind to disallow network access
2. When systemd-logind attempts to send a message to the NIS server, it instead hangs (or has a timeout greater than the systemd watchdog anyways)
3. The watchdog timer causes all systemd-logind services to be restart, which then tears down all of the child processes. My X session is a child process.
However, I believe the implementation is not great.
A couple specific observations:
* lots of binaries. /lib/systemd, the logs. This more than anything else is a tragedy. Binaries get in the way of viewing, diagnosing, understanding and changing things. It seems like premature optimization to me - I don't think speed requires it. I think the majority are completely unnecessary.
* copying from launchd? Commercial software vendors ship binaries. They are aligned with them. They resist disclosure, which protects from reverse engineering, lawsuits and user modifications. Does that align with linux?
* the config files are a disorganized mess. It is like /etc/init.d but poorly understood and then /etc/rc.d was lumped into the same directory and organization became chaos. I looked at one system here and in /etc/systemd there are 16 directories, 23 symbolic links and just 11 files. Try unraveling the interdependencies.
additionally the config files (which mimic windows config files) lack depth. Simple config files can have benefits, say if a GUI read and wrote them to change settings, but I don't think that applies here. What I noticed is that they cannot do anything, so additional logic requires an intermediate script or compiled binary elsewhere.
* it is responsible for too many functions. I bet it didn't solve so many in its first iteration, but sucked in more over time. It reminds me of the accounting program that grew to rule everything in Tron.
I wonder if these problems could be resolved. It's hard to remove complexity and add elegance later.
would you prefer sh/python/perl/lua scripts?
some (the suckless group) view C and make && make install the perfect way to observe and change things. (I like config files and runtime configurability.)
These dependencies between services have been there the entire time. At least now they're clearly expressed.
Lennart
If I’m writing a piece of code and I start wanting to rewrite the world around it, that’s a sign that I’m probably doing the wrong thing.
In the case of systemd, you actually have a decent case that a lot of things in Linux and the rest of the POSIX world could stand to be improved. A good, clean, portable, systems-focused approach could tackle all those things head on.
> We can't win
Juxtaposing two unrelated issues usually isn't a winning strategy.
It's that the plethora of binaries are too tightly coupled and are not at all modular or interchangeable, which is materially the same as being a monolith.
I think you are smart enough to know this, and are just feigning ignorance.
Sure I have to learn new things, sometimes stuff breaks, but that has always been the case. For me most problems were with upgrading Linux distributions.
Anyhow it's great to have less distribution specific stuff and less shell scripts. Thank you and all the other systemd devs.
It's got tons of really great functionality.
systemd feels like it is well thought out and modern, like it's an integrated rethink and rebuild of lots of various software utilities that have evolved over many years.
I'm investing all my learning in systemd where appropriate instead of things I used previously like cron and supervisor.
Case in point: I deal with several clusters that use Stanford Central LDAP for account info, and our UID numbers have gotten pretty high. So much so that it overlaps the range systemd uses for dynamic service UIDs.
The biggest annoyance is that the UID range is hard-coded, so I can’t provide an alternate range.
More details: https://github.com/systemd/systemd/issues/9843 (but please don’t spam the GitHub issue!)
https://github.com/systemd/systemd/issues/9843#issuecomment-...
If you are a user reporting a bug, and you say "I am using the RPM package of systemd from CentOS 7.4", then I know which RPM to download to get exactly the software you are using.
But, if you say "I am using the RPM package of systemd from CentOS 7.4, plus my own patch", then things will be harder for me. Not only would I need to get your patch, to be completely safe, I would need to get your build environment. For example, which compiler you are using, and which version of the compiler.
Working off of a common set of pre-built packages makes problem reproduction alot easier.
Unfortunately, unlike before, you can't easily just cut systemd out of the mix.
I don't even understand the rationale for some of the integration. Like systemd-resolve. Why did it need to suck in the resolver? Now I can't edit resolv.conf because it overwrites it.
I can't even find a reasonable list of everything it does. Init, mounts, login, pluggable hardware, nspawn containers, logging, and...whatever else I'm forgetting.
Systemd is written in C and relies on kernel headers. Which means that they need to have a layer that deals with different versions of the kernel headers. I remember it was headers about network.
At the time I was working on this, I had to contribute to systemd to add support for my kernel version. Got a few patches accepted in their master. And things were working fine for a while.
Until they suddenly decided to drop <= 3.10 kernel support. Yep. Afterwhat I had to maintain patch sets for systemd, which is a great amount of job, given that my upstream was updating systemd quite often.
Now my good people, explain me how such a mess could have happen with a previous init based on shell?
Anyway the root cause here is bad vendor BSPs being stuck on old Linux versions.
This is the person behind PulseAudio, Avahi, and systemd. In any sane world, these software projects would have been stillborn. Sensible alternatives would have outcompeted them. Instead, Red Hat did their best to foist them upon the community. They promised that after the initial adoption, all the wrinkles would be ironed out. All three times this was not true.
I hate to appear so cynical, and I'm certain that Red Hat didn't intentionally pursue this strategy, but it sure is convenient for them that the harder Linux is to use, the more they can sell support contracts. Not so convenient for people who actually want to use Linux.
Considering that Midgley Jr. is (indirectly) responsible for the suffering and potentially death of many humans and other beings, I submit that this (in spirit) triggers Godwin's Law. The debate is therefore over.
Poettering is directly responsible for the suffering of many human beings, as indicated by threads like this. It'd be difficult and unfair to accuse him of being even indirectly responsible for anyone's death, though.
Of course, and they didn't, because there were no sensible alternatives on linux. All these projects were clones of decently engineered macos projects, before they were created, linux people used some truly silly staff like sysvinit.
I would use something better than pulse and systemd, but there is nothing better available.
systemd is the third world copy of Solaris SMF, is it not?
They manage better performance in real world use cases with a fraction of the manpower invested in them. I wonder what will happen to the suite of projects supported by redhat now that they are owned by IBM however.
Jack as a regular sound daemon? are you joking? Have you tried to run jack with several audio sources and audio outputs (including bt headset) without pre-configuration? It's not a pulse alternative, it's a low-latency daemon for professional stuff. Consumer oriented daemon should just work.
>runit
Does it do anything beyond service managing? Login, timers, bootloader, containers, logging facilities?
>runit starts /etc/runit/1
Wait, does it simply run sh scripts? No, thank you.
I'm not aware of any audio daemon that is as feature rich and flexible and PulseAudio. None of the alternatives fit the same general purpose use case that is expected in a modern desktop that needs to compete with Mac OS and Windows 10.
Things like hotplug, Bluetooth devices and as such shouldn't be managed by your desktop manager, unless you want for things like hotplug and Bluetooth to only work in KDE.
Funny story:
During KDE 4.0 development, KDE introduced Solid library. Which abstracted HAL.
HAL was Linux'es "Hardware Abstraction Layer".
So HAL developers mocked KDE and got some tshirts that said "KDE Abstracted my abstraction layer".
1-2 years later HAL was deprecated. And Solid got a new upower/udev backend and no KDE developer had to port away their code from HAL to anything else.
Phonon's situation is similar. Xine, gstream, vlc, these technologies come and go. KDE apps don't care.
Also Phonon can get a QuickTime backend when it's built for Mac or something else for Windows.
Huh, I didn't know he was behind Avahi. That... explains some things. Thanks.
(And for the record I wouldn't call Torvalds an inventor either, even though he has more sense than Lennart.)
DevOps (yelling): no more shell scripts
Devs: systemd fstab is all kinds of fucked for nfsv4
Devs: Writes super complicated script to mount
Me/DevOps: Makes ansible script to make systemd mount paired with service that is basically a wrapper for... a shell script...
I think the only Unix system to get this right was likely Solaris via SMF.
https://en.m.wikipedia.org/wiki/Service_Management_Facility
WindowsNT also did it correctly.
systemd reminds me of the efforts in some cities at clearing 'slums' to replace them with 'modern' apartment blocks.[1]
sytemd seems like a similar technocratic, modernist effort aimed at UNIX.
Operating systems are kind of like cities. One the surface they are often inefficient and messy, and people often wish they could come in and make them logical and rational, but people live in those cities and neighbourhoods, and sometimes the postbox is where it is for a reason.
If Jane Jacobs[2] were around, she could write a book called: "The Death and Life of Great American Operating Systems"
[1] See: https://en.wikipedia.org/wiki/Pruitt%E2%80%93Igoe, https://en.wikipedia.org/wiki/Quarry_Hill,_Leeds, https://www.youtube.com/watch?time_continue=862&v=Dxr2tEGlWU...
[2] https://www.citymetric.com/fabric/against-modernist-nightmar...
My favorite systemdisms:
* (already mentioned several times above; sorry) even detached and nohup'd processes will be killed if the systemd context that transitively led to the creation of the process is killed: https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=825394
* mounts being in a separate "mount" namespace that is not accessible normally: http://www.volkerschatz.com/unix/advmount.html
Now I've been using Docker. And there is no systemd in containers. But of course, why would you want that? But the problem is that many packages now depend on systemd. Or at least, useful features do. So I'm back to fighting with scripts. It's funny, no?
I've been wondering about all those zombies.
I play a lot with shell scripts, and ps gets littered with grep and awk.
What I mean is that systemctl isn't available in containers:
root@foo:~# systemctl status sshd
bash: systemctl: command not foundI knew a guy who invented an orchestration system that involved shoving an init process into docker containers. Much sadness ensued.
Is there a better way to reap zombies?
Nope.
Util-linux-ng 2.14 Release Notes (09-Jun-2008)
==============================================
mount(8) supports new "nofail" mount option.
https://github.com/karelzak/util-linux/commit/abe3d704b6aeb6...The fact that previous init scripts gladly ignored these failures does not reflect negatively on systemd but on those init scripts.
Since 2011, FreeBSD has the same options with the same behavior, except that "nofail" is called "failok".
https://svnweb.freebsd.org/base?view=revision&revision=22283...
Add a special mount option "failok" to indicate that the administrator wants
the system to proceed to boot without bailing out into single user mode,
even when the file system can not be successfully mounted.
By your argument, FreeBSD broke UNIX.Ever since I had an RH AS server boot without /var mounted. It was truly lovely with a gigabyte of trash filling up / then trying to decide if it was worth merging data or just throwing it out.
Yes, it is faster. But when it breaks because bad dependency, it is hard to reason the breakage on a initrd shell on a small remote console.
I value debuggability, reliablity and manual fallback options on server.
I'm always confused with the Type (should it be simple? forking?) and my program ends-up starting but not stopping or things like that.
I know I lack the knowledge to write a good .service file in one go, as I write one every other month at best and forget all the details in between. It would be the same with a traditional init scripts. However, with init scripts, the way to debug is obvious: just add a '-x' somewhere and here we go, you have a detailed execution. I don't know the equivalent with systemd and I end-up bashing my head against the wall trying everything blindly.
Wow, just wow. Do you have any idea how insulting this is to the folks dedicating their lives to writing this GPL software?
You should try having an actual conversation with the authors some time. These are real people and long-time GNU/Linux users and advocates you're implicitly accusing of being NSA collaborators.
Linus himself semi-openly admits he's been approached...
Devuan (Debian without systemd):
More Linux Distros without systemd:
http://without-systemd.org/wiki/index.php/Linux_distribution...
22:10 "Unix is dead."
25:28 "Change is awesome when we are the ones doing it.
27:20 "No one needed to send death threats over a piece of software."
27:30 "Contempt isn't cool."Reminds me of the saying, "and if your friends all jumped off a cliff, would you jump off too?" which exasperated parents use to council their children that emulating others needs to be done responsibly.
If systemd didn't cause problems for people, virtually no user would even know what it was, and consequently, virtually no user would hate it. If you ask your average OSX user what launchd is, they'll come up with a blank because launchd isn't causing trouble for them. If you ask your average linux desktop user what systemd is, a large portion of them will recognize the name, and many of them will chew your ear off out of frustration.
At the end of the day it's not about "adherence to philosophies" or any similarly dubious concepts. Whether or not software is hated is ultimately determined by whether or not it causes problems for people.
Benno Rice's explanation is actually ahistorical here. What actually happened is that the world spent a lot of time "cloning" existing Unix softwares during that time period, for reasons that we all known, and a lot of the time the clones were behind the times and did not catch up. By the time that Miquel van Smoorenburg cloned AT&T Unix init and rc for Minix, for example, what xe was cloning had already become years out of date.
Yeah, I suppose I also admire the willpower and determination of Genghis Khan.
Corbet is almost as dismissive as Lennart himself of the life-disruptive bugs real people encounter due to Lennart's approach to software engineering and project management. "It has bugs? It's software" This attitude is positively correlated with the amount of hate systemd (and pulseaudio) have gotten over the years.
Generally speaking people like things that work well and hate things that don't. Any other justification citing 'philosophies' or similar nonsense are retroactive attempts to cast their feelings in a more intellectual light.
When somebody says "I dislike systemd because it violates the Unix philosophy" what they almost certainly actually mean is "systemd has caused me a lot of grief." You don't hear complaints about firefox or chromium violating the so called "Unix philosophy" because generally firefox or chromium will satisfy the overwhelming majority of users. That's why nobody uses Uzbl.
This is what systemd advocates (like the article's author) don't seem to get when they point out that systemd is a launchd ripoff so people who like launchd should like systemd as well. People like launchd because launchd doesn't cause them grief. People dislike systemd because systemd does. Design philosophies have jack shit to do with it.
Anyway the only real 'philosophy' Unix ever had was "worse is better" and systemd certainly seems to pay homage to that.
We can go back and forth with dumb comparisons all day but I don't really think we're advancing the conversation.
Also this: the equally widespread SQLite seems to be programmed with a much better level of discipline and comes with much better testing. It isn't nearly as bug-ridden.
I do get frustrated with sqlite a lot because it's missing so much of what I expect in SQL (no ALTER tables, having to rewrite a lot of right joins into other types of joins, etc.) but I have never had a situation where I got frustrated with sqlite because it lost my data.
They don't dismiss concerns - they validate them.
We can agree on the problem and then on the solution, but any proposed solution can't just pass through 'without scrutiny' because a problem exists.
Systemd proponents since the beginning have positioned any criticism as 'motivated', 'greybeards', 'haters' or anti-change which tells you that there is no space for discussion here, only 'acceptance'. Can there be any legitimate criticism of systemd without this kind of politics, hypersensitivity and now posing as victims?
This seems to imply that those criticizing it didn't have alternatives, which is just not true. However it is so monolithic that as soon as a tiny part is depended on by something else (e.g. gnome) alternatives quickly become a non-starter.
systemd seems to be getting rewarded by not playing well with others which is something that plays strongly into a lot of people's dislike with it.
sysvinit and typical distribution init scripts couldn't implement the behavior seemingly desired by the user, either.
The ticket follows the sensible approach: Document it, and close it as expected behaviour, and it did lead eventually to an LVM issue that did require a fix [0].
systemd did good here, above board, but there was a faff about it, because the expected behaviour of the audience differed from the reality of the situation.
I mean without that behavior linux is not a multi user system. The default is sane, however the lack of "permissions" in this kind is also a little bit bad, however since it is runtime configurable this should be enforced by the distribution, i.e. when installing the distri there should be a switch "this system is only used by a single user". it's basically the same as windows uac. while it looks bad for a single user, it's not that bad when you see computers as a tool which can have more than one user, with more than one "privilege".
system just shows that linux is by far not ready for the masses cause a lot of behavior is undefined, like the uid/username stuff. usernames can be created as you wish, but most programs will behave broken if you create usernames that start with a number, it's a security nightmare.
and the problem that there is no real process to actually get ALL people together who can solve this, it will never be fixed. there is only the posix standard but not everything applies to linux and not everything is a good idea in the sense of linux.
systemd actually makes a lot of things in the right way, but of course it's not perfect and probably at a certain point in time people would come up with something better.
I've seen the BSD talk on this and I agree, having a system layer is helpful. It'd be nice if it was plugable, NetworkManager (or others that have some standard messages you can send/get via dbus), consolekit OR logind, etc.
systemd does make it nice that I only have to write startup/shutdown scripts once for each distro, but I'm not happy with the layout of target files, the way mounts are handled, some of the weird race conditions I've found between systemd mount targets and fstab, etc.
systemd is modular, but the modules are still all part of the whole and are not easily replaceable. The same can be said when Docker went to a modeler refactor, but there are alternative implementations of the entire docker engine. Every attempt to create alternative implementations of systemd have eventually gone unmaintained because systemd keeps getting more and more complex and engulfing more systems.
If it wasn't for distros like Void, Gentoo, Alpine, Slackware, et. al, we'd no longer have a choice at all. There would be some things that simply couldn't be deployed on embedded systems because all of the dbus shims just wouldn't exist.
It's not that people are opposed to change, it's that there are legit concerns about some of the ways systemd works and is implemented, and the way it's been ham-fisted as a political move in a lot of ways.
Honestly, I don't think it will matter in a few years. I think the way things are going, eventually all services will be hosted via docker containers and it will be much easier to make Linux distros that have a tiny init layer that just launches a docker daemon and services. RacherOS already does this, with the init process being a container, which can be uses to start up shell environment containers and other service containers.
I personally think the industry needs a lot more resistance to change when it comes to interfaces and other things humans have to understand.
I mean, I'm not talking about systemd in particular; I'm talking about in general about how interfaces change over time and people don't seem to take into account the cognitive costs of that change. Sure, ss is better than netstat and IP is better than ifconfig... but how much of that 'better' could you have done in a way that didn't toss away the historical knowledge so many people have of those tools?
And really, sysadmin tools are the least of it; I mean, they are operated by professionals, so if you want to pay for retraining (or pay the costs associated with there being fewer of us)
People change customer facing interfaces to no benefit all the time, forcing people who are trying to do other things to put effort into re-learning their interface.
I mean, my point is that interface changes are expensive, and should not be undertaken without a really good argument that they bring more benefit than the cost of retraining.
netstat? /proc files? ss? parsing text? wtf?
I mean, sure why not, but at least don't call them interfaces. they are userland apps people like to script, because they are lazy to use libnetlink (or libwhatever thay uses the right kernel interface, if it exists at all).
That said, the recent gmail ui change made me reconsider Thunderbird again. And android looks different every year. sometimes it's better, sometimes it's worse. iptables, nftables. http1, http2 (and now 3 over UDP). change is the only constant.
Text processing is not harder than figuring out what library to use this month.
These things change a lot... but they don't have to, and running things on computers would be easier/cheaper if they didn't.
The idea that you are searching for is coupling. Modular systems should aim to have low coupling and high cohesion.
Additionally this concept of stateless container design and state kept in containers there are opposing implementations
But that said, does anyone know where on earth they came up with the command line ux? Like the names of the commands , and the parameters? I mean, they are like an April fool's joke...
Like, previously to have any logs at all you had to have a syslog daemon. This is not a new situation.
What's new is that now more and more-difficult-to-debug pieces are required. For what benefit?
I remeber seeing someone post one of the simplest init process one could write. It was only about 100 lines of C.
Of course, if you do that then the question is why even bother with an init process. Patch the kernel to not treat pid 1 specially. If a process gets orphaned, don't reparent it to init, just reparent it to nonexistent pid 0. Have the kernel deal with the thing that the 10 lines of C would be doing (waiting on processes to terminate).
I think there are meaningful complaints about systemd but the fact that it's a complicated pid 1 is not one of those. If it were actually a problem someone would have written and productionalized one of the above two approaches.
Really the biggest issue with systemd is is scope creep scope creep means more and more packages will have systemd as a hard dependency. Like there some aspects of gnome that will not work without systemd, but at least it's not entirely a hard dependency just means you lose features unless you patch gnome.
They want to push the cloud, to better control us, take away power from us, leave us at their mercy.
The next step is, when everything that matters is on the cloud, they are going to replace Linux with something else that only they know how it works on the inside, and that's going to be the end of it, we'll be left only with paywalled APIs and services, but they'll say it's better for everyone and all of that. It's sad.
I feel like as an industry we're wasting tons of cycles solving the wrong problem.
> Nerds have a complicated relationship to change; it's awesome when we are the ones creating the change, but it's untrustworthy when it comes from outside.