The importance of Devuan
blog.ungleich.ch
blog.ungleich.ch
Opensource is about alternatives. Forks are a good thing,I wish devuan great success so people like myself won't have to whine about systemd all the time. If you like systemd and it's philosophy,I wish you the best of luck with it. Just keep in mind that others don't have to like it and different people or organizations have different needs.
Whike problems are surmountable, rehearsing for failure is advised.
There definitely is bad too.
The real issue with systemd is that there simply aren't enough volunteers (both individuals and companies) out there willing to contribute. In short, since gnome is pretty much developed by RedHat and systemd is developed by RedHat, gnome is tied into systemd.
If there was enough of a community around gnome that RedHat wouldn't be able to pull it wherever they want, gnome wouldn't have been able to be that coupled with systemd. The issue is that debian simply doesn't have the manpower needed to de-systemd gnome.
* http://jdebp.eu./FGA/debian-systemd-packaging-hoo-hah.html
When I ask my fellow colleagues (mostly sysadmins), some of them like Systemd, some don't, but we all agree that in this particular case the freedom of choice is much needed.
Sincerely asking, I never tried running a non-default init system.
Right, but did they support using any other init system (besides sysv) before?
* https://packages.qa.debian.org/s/systemd.html
* https://packages.qa.debian.org/u/upstart.html (it's gone from stretch that's true but still supported in jessie)
* https://packages.qa.debian.org/o/openrc.html
* https://packages.qa.debian.org/s/sysvinit.htmlAll my Debian systems run sysvinit, and they work great - certainly better than the one system I co-administer which runs systemd.
I frankly do not understand why Devuan exists.
In any case, the pain was real. So finally I switched to Devuan, even though I wanted to stay with Debian, and now those problems and the time I had to spend fighting them are gone.
That's why Devuan exists. YMMV, of course.
Package: systemd
Pin: release o=Debian
Pin-Priority: -1
Package: systemd-sysv
Pin: release o=Debian
Pin-Priority: -1
Package: systemd:i386
Pin: release o=Debian
Pin-Priority: -1
Package: systemd-sysv:i386
Pin: release o=Debian
Pin-Priority: -1
The i386 parts are only needed if you have it setup with multi-architecture.Unfortunately, I have also had to recompile some packages to remove systemd dependencies as well. Fortunately that hasn't been too much effort so far but eventually I may have to switch to Devuan. The most ridiculous one was policykit and neatly illustrates the creeping "infection" of systemd into Debian - policykit-1 hard-depends on libpam-systemd which hard-depends on systemd. Why neither was marked as "Recommends" instead of "Depends" I can't understand (policykit-1 -> libpam-systemd or libpam-systemd -> systemd) - either would have been fine and allowed both options to work perfectly well as far as I can tell.
Sounds like a bug report is needed, so that gets fixed. :)
"Fill a bug report" automatic reply gets old quickly and it's especially absurd when it's said by a person totally uninformed on the issue.
And I think that's exactly what Devuan does - they recompiled/rebuild the packages with funny dependencies that caused systemd to creep back in. The rest are just the normal Debian packages. They even forward you to the Debian download servers for those.
The advantage of using Devuan is that they have spent tome to do this work, once, so now not everyone who wants to use Debian without systemd has to spent all this work again.
Yes, the Debian devs could have easily done this themselves. For some reason they decided not to.
If you install certain desktop packages, then a package named "systemd" might come back, but systemd won't actually run as PID 1. Since the pain from systemd comes from it running as PID 1, I don't see a problem with this.
If you want a system that's 100% free of all systemd code, even if it's not doing anything, then you need Devuan, but personally I don't see that as a compelling reason to fork Debian.
Frankly, it's an ugly hack, not a proper solution. I hope to run my servers for years to come and the upgrade process should be as smooth as it can be.
I thought the same at first but after trying to maintain a Debian installation without systemd, I was convinced otherwise. The "infinite cost" was not caused by Devuan developers but by Debian developers.
I felt the Debian transition to systemd was handled very badly and was rushed in order to get Jessie out of the door. Debian still provides the sysvinit-core package as if it were a supported init system however if you use it you'll find the operating system has many quirks and many things are broken. Before Jessie was released, I commented on one bug report related to a hard dependency on systemd for NetworkManager. Instead of fixing the regression, the package maintainer gave the excuse that they don't have the resources and would need more maintainers. I sympathise with the package maintainer but this is the sort of thing that led to this mess. I remember there being a lot of hostility in that period from Debian developers towards those who did not wish to use systemd.
Based on my experience, I think the right decision was made to "fork" Debian to create Devuan. It's a shame that they had to do this. I still hold hope that maybe one day all of the work put into Devuan can be reintegrated into Debian and it can return to be the "Universal Operating System" that they still claim to be.
I don't know whether that means it's "supported" or not.
A Debian developer built a deb package for it and added it to the "main" repository. Everything under "main" is normally considered to be part of the distribution. As you can see from the link I provided, the package was maintained for several years and received several updates until it was dropped for Stretch. That is exactly what I would expect from a supported package in Debian.
runit-init had been removed in 2010. Felix von Leitner's minit was later removed by Debian member Iain R. Learmonth in November 2015. upstart was removed as a consequence of a bug report (#789524) filed by Kyle Amon in June 2015.
One of the contentious events during the Debian Hoo-Hah was one of the package maintainers leaping to remove support for an init system from one package almost as soon as the Technical Committee had made its decision favouring systemd over upstart, OpenRC, and van Smoorenburg rc. People opined then that Debian package maintainers should not blithely remove support for init systems like that.
* https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=746715
More recently, however, this has begun to happen. In August 2017, a group of Debian people got together to organize the removal of all upstart support from all packages; and that is progressing.
* https://lintian.debian.org/tags/package-installs-deprecated-...
As for systemd, it wasn't as a Debian package, so it's hardly relevant for this discussion. It was only even packaged in 2012.
You have a lot to learn about the rather sad history of runit and daemontools in Debian, as well as about the ways that Debian people decide what is packaged. This is only some of it.
* https://tracker.debian.org/pkg/runit-run
* https://lists.debian.org/debian-vote/2014/11/msg00059.html
* https://unix.stackexchange.com/a/284453/5132
* https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=766187#78
* https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=861536#44
* https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=752075#35
Runit-init-the software was not removed from Debian. Only runit-init-the-package was removed and that was long before systemd-the-package (which, yes, was what I was referring to) appeared, and it was by a ROM from the author himself, after a bug report classified as RC (justifiably, in my opinion).
So I fail to see any changes in behavior by the Debian team, since I don't believe in this conspiracy theory that Debian Developers were finding bugs in other init systems just to push for systemd, despite not even caring enough to package it.
Could they be more open to other init systems? Probably. But that was never my question.
After propounding a non-existent concept of "obsolete as a Debian package" you are proceeding to propound an equally nonsensical concept of removing runit-run whilst not removing runit-run. https://tracker.debian.org/pkg/runit-run is really quite clear and unambiguous.
And that conspiracy theory is a straw man of your own making, too. As is indeed the commentary on "the ways that Debian developers make their decisions" which are your opinions not anyone else's. You have a lot to learn and you're getting a lot of things quite wrong.
Systemd allows actual resource management on a level that simply didn't exist with traditional init [1]. Yes we can argue about implementation, but the necessity of better resource management in more dense and shared environments can't just be ignored, especially when building a shared VM hosting platform like the author is talking about.
Those are two somewhat recent examples I can think of, but there have been many more showing the same kind of attitude.
They are ignoring hard-earned knowledge on how to do things securely, safely and reliably, they ignore the intentions of abstractions, and they pretend that obvious bugs aren't bugs because there is some weird way to reinterpret the bug into being correct behaviour (that noone expects and that causes harm).
By domain knowledge how to do this securely in implementations tested oer time, do you mean something like https://access.redhat.com/articles/2161461?
It doesn't look like systemd's resolver fares worse in comparison... In addition it can be sandboxed (and is), an advantage over having a resolver part of libc.
Would you mind explaining how exactly sandboxing prevents cache poisoning?
I have used LXD (with ZFS storage) in production but have not played with systemd-nspawn...
Thanks!
I usually debootstrap into /var/lib/machines/something and do "machinectl enable something; machinectl start something", that's it. Then I attach to the machine using "machienctl shell something" and configure networking (host0 interface) inside the domain, that's it.
For drop in configuration systemd-nspawn parses a config file /etc/systemd/nspawn/something.nspawn which usually just contains network configuration on my hosts:
[Network] Bridge=br-int
Systemd-nspawn enables and user namespacing by default and chowns the machines's root filesystem on first start. If that's not desired (Things like Samba fileservers don't work well with user namespacing) just disable it in the .nspawn file:
[Exec] PrivateUsers=no
Everything you need to know is in the manpages systemd-nspawn and systemd.nspawn. I usually install systemd from stretch-backports because running a fairly recent systemd version helps as it still gets new features, but I never had problems with stability.
One thing I somewhat miss from what you are explaining is all the aditional things that LXD gets you (snapshots using ZFS, image publishing/sharing, migrating containers between LXD hosts...)
But maybe some of those things are still doable (e.g. mounting a ZFS dataset as storage for /var/lib/machines/containerX)...
Thanks for your answer!
Just drop a .mount file in /etc/systemd/system and set RequiredBy=systemd-nspawn@something.service and StopWhenUnneeded=true and the filesystem should be mounted before the machine starts and unmounted when the machine is shut down. See the manpages systemd.unit and systemd.mount for details.
It's a bit like ISA vs. ISA PnP vs. PCI. Jumpering ISA cards to make sure resources didn't conflict was a bit of a chore and sometimes difficult to get right, but essentially there always was a way. ISA PnP tried to automate this, which was great if it worked, but more often than not just failed, and then you had no jumpers to fall back on to just fix things up manually (though sometimes you had special config utilities that with some luck you could use to fix things up with some cards ... maybe). PCI, though, was an actual reliable abstraction that actually worked essentially all the time, so there actually was no use for jumpers, so it is fine that PCI cards don't have IRQ/IO jumpers.
Systemd seems to me like the ISA PnP of init systems.
> the reason to use Devuan is hard calculated costs. We are a small team at ungleich and we simply don't have the time to fix problems caused by systemd on a daily basis.
They lament
> servers that don't boot, that don't reboot or systemd-resolved that constantly interferes with our core network configuration
Using systemd was costing them too much. Moving back to the previous init system looks like a rational choice for them.
My experience is different but I only manage a handful of virtual servers and not on a daily basis, plus my laptop. Systemd configuration files are not difficult to write and they restart the daemons if they crash. For complex stuff I make systemd run a bash script that eventually executes a daemon, kind of cheating. My laptop still runs well. Booting time is definitely not an issue, it went from fast once per month or so (kernel upgrades), to fast still once per month so. I didn't notice any difference after the change of the init system. A good thing but maybe it means that the return on investment was dubious for this use case.
Binary log files are objectively worse than text ones. cat, less and tail were good enough and shorter to type than journalctl. Ok, I could alias it to a four characters word but it was still a lot of work that could have been invested on some other goal. Instead I've got servers with possibly compact binary logs made of very few lines and large text logs from web applications.
I'm also puzzled by the philosophy of bundling more things together and tighten dependencies. It's somewhat disconcerting and I'd like a system where we could swap components out more freely, but this is an opinion and not facts.
>> servers that don't boot, that don't reboot or systemd-resolved that constantly interferes with our core network configuration
Which is interesting given Debian doesn't use systemd-resolved (unless manually configured). So they made up problems that don't exist by default?
This I have a problem with, if it crashes, I'd prefer it stay crashed, instead of crashing over and over.
At least when it crashes, a human can come and see why it's crashing, rather than doing the caveman thing and just restarting it over and over.
https://www.freedesktop.org/software/systemd/man/systemd.ser...
Truth be told, I'm fine with monoliths. They just have to be good. The Linux kernel is a monolith.
The real problem to me is how systemd makes the rest of the system more complex and integrated to systemd. Dbus is also not a great IPC layer, and yet it's the current state of the art.
We need better APIs and foundations. Systemd doesn't provide those; it just lets everyone depend on systemd. Good for Redhat, bad for innovation.
I have to get our systems people to write a paper on it. My experiences with systemd are pretty much perpendicular to what some people claim.
Anecdotal evidence is useless.
What you're suggesting is basically spreading FUD. Your article didn't point out a single actual issue with systemd.
I run large scale VM infrastructure on a systemd distro (RHEL, in particular) and I've yet to encounter a single issue that was caused by systemd or, for that matter, any issue with systemd that wasn't easily resolved.
That doesn't mean that there aren't any. Also, FUD (as I understand it) is people spreading unfounded rumours, not people documenting their own experiences and problems with a piece of software. The intent is incomparable.
If you are actually interested in tangible structural flaws in systemd, look no further:
https://suckless.org/sucks/systemd
https://www.theregister.co.uk/2017/07/28/black_hat_pwnie_awa...
Also, see how Poettering handled this zero day (Which grants root to any user with a numeral): https://github.com/systemd/systemd/issues/6237 HINT: He doesn't think it's a bug...
The fact that this even fucking exists in the first damn place: https://latesthackingnews.com/2017/06/29/a-systemd-vulnerabi...
See also: http://without-systemd.org/wiki/index.php/Arguments_against_...
This formulation may be misunderstood. The bug did not grant users root rights, it ignored the user= setting on a service if the service was intended to run as a user stating with a numeral. So you already needed to be root to create a service. The bug allows a service to run with more permissions than intended though.
* https://twitter.com/Serianox_/status/922050013655650305
* https://twitter.com/gibbon_/status/935786202724208641
* https://twitter.com/ManChicken1911/status/931820892291608576
* https://twitter.com/hawko2600/status/931082871187505153
* https://twitter.com/RPaschedag/status/930860506256113666
* https://twitter.com/BoolKiRool/status/930422520074919936
* https://twitter.com/infosecabaret/status/929069203193200640
* https://twitter.com/sirzeitgeist/status/925535140947922944
* https://twitter.com/saruspete/status/925182958012715009
* https://twitter.com/iMilnb/status/928260238322651136
* https://twitter.com/patrick_mooney/status/936374590070185984
* https://twitter.com/patrick_mooney/status/930929899686182912
* https://twitter.com/patrick_mooney/status/930929899686182912
* https://twitter.com/patrick_mooney/status/925026104972165120
* https://twitter.com/patrick_mooney/status/925023617674461184
* https://twitter.com/patrick_mooney/status/923815696894607360
At the risk of igniting something, the only serious recommendation I hear about systemd is improved boot times. My computers are rebooted only for kernel updates - this is the entire point of running Linux. I want uptime. I can stand a boot time of a minute or so if it's once a month or less. If I have to reboot regularly enough that boot times become a serious advantage, someone has done something very wrong with Linux.
In my opinion, the Debian developers were wrong to fully embrace systemd; it goes quite against the Debian principle of being able to adapt your machine to do anything. Yes, Red Hat held a lot of sway in convincing the community to adopt it, but Debian is a big distro as well - just look at the install base of Ubuntu. They could have held out and pushed back, but instead caved and now we're in the systemd mess. It intrudes so massively into userspace that it's impossible to get away from, and we wind up with the Windows model as this article points out - a massively complex spaghetti model of processes, all interlinked, all hinging on everything playing nicely. Just like Windows, it's a neat house of cards that could topple if any one of those crashed. This isn't Linux.
I'm not defending sysvinit, there's no denying it's showing its age and harks back to a much simpler time, but systemd is not the answer.
That's far from the only advantage systemd has.
Even then, boot times are extremely important for modern VM and container infrastructures.
Can you expand on that? I work at a cloud provider, and I never heard of such issues before. Our default images are Ubuntu 16.04, RHEL 7 and SLES 12, so systemd all around.
This could be an honest mistake. On my system, systemd-fsck-root.service and systemd-fsck@.service have `TimeoutSec = 0`, which I guess to mean "no timeout", but the systemd-fsck executable could have its own timeout. In any case, filing a bugreport sounds much easier than switching distributions.
And if the fsck is just a scheduled one (after x reboots) instead of being triggered due to filesystem weirdness, then getting the server back up and serving can be the higher priority.
That's just an example because you asked though. Wouldn't want that same fsck to be skipped if it was actually started due to filesystem weirdness being detected. :)
Mind you, there are plenty of other complaints (stability, back-compatibility, etc.) that are useful/valid regarding this software, but the above litmus serves to eliminate rather a lot of FUD and hand-wringing in my experience.
Paradoxically, this improves the signal of discussions about systemd's flaws, and there are plenty. With the obviously irrational/haven't-used-it/axe-to-grind stuff filtered out, it's much easier to learn about real features, problems, fixes, benefits, and drawbacks of any project, systemd included.
Actually, I replace systemd with *BSD. You know, the OS which has blessed core packages which are not replaceable?
SystemD has opaque behaviour and binary interfaces. It's the antithesis of the BSD model. Despite BSDs being quite coupled groupings of software.
Anecdotally we're attempting to move to BSDs instead of upgrading to RHEL7 due to SystemD.
... and for deployment, the Golang debugger doesn't even run on BSD at this point. So, going to have to setup new Linux (likely CentOS 6 - no systemd - or 7) boxes soon now that we're getting closer to first initial real world deployment for the project I'm on.
It's just the "UNIX is supposed to be decoupled" line that gets repeated every time systemd comes up that BSD (which, ironically, those people want to switch to) disproves.
So yes, the same centralization/reducing-choices arguments apply to glibc.
The problem with standards compliance is that, once something becomes ubiquitous enough, its compliance begins to "drift", and replacement parts get harder to find. Is it tricky to swap out the init component of systemd with something else? Not terribly. Is it tricky to swap out its dbus messaging setup with something equivalent that can still be used by unaltered software that was previously talking via the systemd setup? Yes, a bit. And so on. For equivalent hassles regarding libcs, check out some of the FOSDEM talks from the musl folks, or the alpine folks, where they talked about the difficulties in porting software that had come to depend on the quirks of glibc.
Or just, you know, google "quirks mode" and get simultaneously nostalgic and sad :)
I am running Void with musl on a Raspberry Pi 2. I love it.
It seems dead.
For me it is not only a Sunday morning essay, but also important that people understand, why the Devuan movement is so important to all of us.
From a business perspective, the most important thing -- indeed, the only thing that truly matters -- isn't the software itself. It's support. When it comes to open source, you either use a supported configuration or you assume ALL the responsibility for supporting your systems yourself. Who would you rather handle the support on the off chance your machine goes tits up? You? Or Red Hat?
No wait, scratch that. Who would YOUR BOSS rather handle the support?
The supported init system for most flavors of Linux (that aren't in LTS) is systemd. Systemd makes the most sense from a business perspective, and "but muh Unix philosophy" doesn't hold when there's money on the line. May as well get used to it.
That's just it, some of us aren't interested in the 'business perspective' and have other uses for our systems.
This is a good analogy to remind Devuan why they rejected systemd in the first place.
I've run large-scale infrastructure on both Debian and RHEL, with and without systemd. I've been involved in many distro decisions in the past. The init system was never a major consideration, and I've always been able to work around whatever issues I've had with both sysv and systemd.
Here's some considerations which were important for my team the last time we had to make a choice:
- Vendor stability and long term support
If you're running a business, you'll want long term stability guarantees - both from a technical, and a business point of view. Running a small community distro with few - if any - commercial users is a huge liability. Yes, you always have the option to fork it or take it over, but unless you're in the business of building a Linux distro, it's almost certain to be a bad business decision since your competitors are spending their time on their products instead.
Anyone who doesn't appreciate long term vendor stability hasn't yet been burned by the lack of it. Your Gentoo wizard left the company? Too bad, good luck maintaining your infra now (I've actually seen that exact scenario play out not once, but twice!). A few core maintainers left for another project and you're left building your own packages? D'oh.
Having a mature and well-funded organization or a large community of commercial users (in the case of Debian) supporting your distro is extremely valuable, even if you aren't the ones paying for it.
- Ecosystem
The ecosystem is also really important. It's often the main motivation for using a particular distro. Development tools, third party package repos, 3rd party enterprise software, troubleshooting resources, documentation and much more depend on a healthy ecosystem.
With Devuan, this isn't as much of an issue, but it's still sufficiently different from Debian in sometimes subtle ways that it will nullify some of the advantages of being in a common ecosystem.
This also includes hiring - you'll have a harder time staffing your company if you're using exotic stuff and you'll unnecessarily spend your time training them. Ever wondered why companies like Google publish so many papers, talk about their infrastructure and even publish books about it[1]?
It's because it means they can - by advancing the state of the industry - hire people who are already familiar with their architecture and concepts. This is a real problem for them - Google, for instance, is in some cases so far ahead of others that they have to spend a significant amount of resources to bring their new hires up to speed. I don't know about Amazon and Facebook, but I'm sure they have similar problems.
[1]: https://landing.google.com/sre/book.html
- Security
Security follows the supply chain. Your security is only as strong as your distro's - no matter how good your security controls and processes are, if your distribution vendor is compromised, you won't stand a chance unless you have a world-class security team.
Likewise, your customers fully rely on you - if you're compromised, they will be, too.
Security is much more than signing your packages and publishing security advisories. In fact, those are merely the results of a properly implemented security management framework - risk assessment and mitigation needs to be an integral part of your vendor's (and your) processes. Where do they keep their PGP keys? Who has access to it? Is is stored on a HSM or just sitting on a server somewhere? Is there a process for revoking it? Is there a change review process? Does the build system verify source code integrity? Is there two factor authentication for production? Do the team members keep their SSH keys on a smart card? This list could go on for miles.
- Quality Assurance
Even with the main distros, QA quality varies wildly. In my personal experience, Red Hat has by far the best QA, followed by Debian and then (with some distance) Ubuntu. I've had particularly bad experiences with Ubuntu, a sentiment shared by many operations people I've talked to (the worst one being the grub timeout bug that made me spend a few days pressing the return key a lot - basically, they removed the default timeout in a minor update and all my machines sat there waiting for someone to manually continue the boot process).
Attention to details, even if it's minor and seemingly insignificant, really helps productivity if you sum it up at the end of the day - each bug your vendor catches and each quirk they document is time you won't have to spend with debugging.
With smaller distros, this is even worse - in some cases, they don't even have proper QA.
Now, let's assume that the issues they had with systemd were as significant as they say they are (I'm doubting it, but unfortunately, they didn't talk about any particular ones), it still might very well have been a better business decision to stick with Debian and spend their time fixing the issues they had instead of migrating to Devuan, which is a HUGE liability compared to plain Debian.
Debian does a particularly good job with security patching (better than Red Hat and Canonical, IMO).
For example-- if you do it today on a core app/lib without consulting upstream is that enough to get your credentials revoked?
I ask because Debian famously patched a library to quiet a valgrind memory error which resulted in their userbase generating predictable key material for a few years[1].
I trust them not to do that again (at least with libs that are obviously security related). But is there explicit policy in this regard?
[1] https://www.schneier.com/blog/archives/2008/05/random_number...
Mounting/unmounting wasn't the biggest problem with systemd. Actually, once we figured out the right ordering of units, it went pretty well.
The greatest pain was the networking.
Systemd is a highly opinionated system. And it is opinionated in ways it really shouldn't be. Happily some aspects are configurable, and in much of our system setup, I reconfigured some of the more egregious settings. Some aspects were simply painful, such as networking.
Years ago, I dealt with other highly opinionated, and similarly broken systems. These systems insisted on doing things in their order to bring up their services, even if I didn't need them, because they maximized dependence radii. Which, for the life of me, I did not understand for my use case. But I could see it for other use cases.
My approach to dealing with this was to provide the absolute minimal basis for that system to operate, and then exit its configuration as rapidly as possible, given our experience with its bugs. That is, have it execute for the bare minimum possible time, before we transfer control to something we've developed, that actually works.
I adapted that to systemd. While I had to put up with all sorts of timeouts for systemd services that I could not adequately control, and could not remove due to these insane dependence radii, I could tune those timeouts way down. Which enable me to escape the systemd startup within reasonable timeframes. And then allow my code to take over.
Customers didn't notice unless they looked at bootlogs. They simply saw a reliable service. Which was made reliable after working around systemd's myriad of shortcomings. I could not get systemd to do what I needed, my opinions were different than its, and its control plane couldn't fathom what I needed to do. So the idea was simply push it out of the way as rapidly as possible.
Last time I checked, systemd - the init system - wasn't involved with networking at all.
There's systemd-networkd, but that's an optional component and no major distro is actually using it as a default. RHEL 7 - which runs systemd - supports both NetworkManager and their legacy networking scripts.
You said there were "myriad of shortcomings" and "egregious settings", some examples would help.
On egregious settings, google is your friend. However, here are just a tiny smattering of what I had to deal with. These were dealt with over a few years of delivering and supporting systems that had to work reliably and predictably, in a supportable manner. Each line has often a significant amount of debugging time invested behind it before we came up with the line you see. $TARGET is the target install directory for the image.
This is just a small sampling of the open source build system, that I grabbed from some of the configs we used.
Yes, systemd is broken, but not irretrievably. It is fixable.
# fix some systemd timeout brokenness
sed -i 's|^#DefaultTimeoutStartSec=.*|DefaultTimeoutStartSec=15|g' ${TARGET}/etc/systemd/system.conf
sed -i 's|^#DefaultTimeoutStopSec=.*|DefaultTimeoutStopSec=15|g' ${TARGET}/etc/systemd/system.conf
sed -i 's|^#ShutdownWatchdogSec=.*|ShutdownWatchdogSec=2min|g' ${TARGET}/etc/systemd/system.conf
# fix systemd journaling. Yeah, really sed -i 's|^#Storage=.*|Storage=persistent|g' ${TARGET}/etc/systemd/journald.conf
sed -i 's|^#SystemMaxUse=.*|SystemMaxUse=250M|g' ${TARGET}/etc/systemd/journald.conf
sed -i 's|^#RuntimeMaxUse=.*|RuntimeMaxUse=250M|g' ${TARGET}/etc/systemd/journald.conf
sed -i 's|^#ForwardToSyslog=.*|ForwardToSyslog=yes|g' ${TARGET}/etc/systemd/journald.conf
# fix the INSANE logind.conf per user directory size ... hard code it to 256M sed -i 's|^#RuntimeDirectorySize=.*|RuntimeDirectorySize=256M|g' ${TARGET}/etc/systemd/logind.conf
#
# fix the INSANE logind.conf KillUserProcesses problem, which nukes nohup/tmux/screen ... sed -i 's|^#KillUserProcesses=.*|KillUserProcesses=no|g' ${TARGET}/etc/systemd/logind.conf
# mask off systemd-udev-settle ... yes it is broken chroot ${TARGET} systemctl mask systemd-udev-settle
[edited for formatting]Are you sure about that? To my knowledge, Debian 9 is still using the good old network scripts. I'm 100% about Debian 8 since that's why I use in production, with systemd-networkd nowhere to be seen.
I agree about it being less flexible. Either way, it's a separate daemon not tied to systemd.
> fix some systemd timeout brokenness
Why are you decreasing the timeouts? I actually increased them in my case to give services more time to exit (upstream default is 90s).
ShutdownWatchdogSec is rightfully set to 0 by default - having a hardware watchdog reboot your system may not always be a good idea.
> fix systemd journaling
Debian runs a non-persistent runtime journal by default and leaves the long-term storage responsibility with the syslog daemon. What's wrong with this? If anything, they're being too conservative.
> RuntimeDirectorySize
What's wrong with the 10% of physical RAM default? Hard coding this to a small value is potentially messing with things like flatpak's portals which store large-ish data there.
> KillUserProcesses
Off by default in Debian. I specifically turn it on, since it kills any rogue user processes when they log out and makes sure nobody leaves anything running in a screen instead of doing it properly (I've even seen a crashed vim instance max out CPU on a server)
> mask off systemd-udev-settle ... yes it is broken
Masking this introduces race conditions into your boot process. The systemd-udev-settle unit calls "udevadm settle", which waits for all pending udev actions to complete before any services are started which may depend on it. If anything is broken, it's the particular subsystem which takes too long to initialize/does not support async events.
Important examples are LVM and iSCSI. If you mask the settle unit, systemd won't wait for all devices to be ready. This may or may not work and the safe default is to wait.
The last job developed/supported Debian based high performance storage systems in production, starting with Debian 7, then Debian 8, and Debian 9.
root@localdns:~# cat /etc/debian_version
8.10
root@localdns:~# locate systemd-networkd
/lib/systemd/systemd-networkd
/lib/systemd/systemd-networkd-wait-online
/lib/systemd/system/systemd-networkd-wait-online.service
/lib/systemd/system/systemd-networkd.service
/usr/share/man/man8/systemd-networkd-wait-online.8.gz
/usr/share/man/man8/systemd-networkd-wait-online.service.8.gz
/usr/share/man/man8/systemd-networkd.8.gz
/usr/share/man/man8/systemd-networkd.service.8.gz
Its there, but I disabled it in this case.
> Why are you decreasing the timeouts? I actually increased them in my case to give services more time to exit (upstream default is 90s).
This comes from long debugging/tracing sessions where we were trying to understand why we wound up with effectively repeating timeouts.
[...]
On systemd-udev-settle, allowing it to run was actually the cause of a race condition on start. The only way to fix this that I found was to disable this. Once we did this, no more race condition. The system booted normally.
It sounds like your experience is significantly different than mine on this. We built very large high performance computing and storage systems for many users, where the things that people used simply had to work, without surprise. Every line I showed came from a set of long and often painful debugging sessions where we had to figure out what new change had broken a key bit of software.
The RuntimeDirectorySize issue actually tickled another bug, and caused some incredible paging problems when used with a specific software package. Enough end users on a system (we had many) could launch an effective DoS against the system by using these tmps.
Generally, my observation, and I've seen others make a similar set elsewhere, is that the default and general distribution settings for systemd parameters appear to be generally set for desktop users (single or very few users per machine). Not for large shared resources.