Walkthrough for Systemd Portable Services
0pointer.net
0pointer.net
Some say docker is the new "curl|bash" ;)
If you manage to tell me what's in /etc/os-release on the machine (not inside the container), I'll eat my lunch.
It's also the fact that it generally doesn't do anything to integrate well with your system, and just pull all its dependencies in a folder and never update it afterward.
To me that's what other technologies like rkt seem to miss or at least choose not to focus on.
They're not much different from nix expressions, or brew formulas or such.
Unlike nix expressions (or rpm specfiles or various other technologies), the dockerfile format never concerned itself with clearly defining the hashes of inputs and dependencies, reproducibility, working without network access on build farms, etc etc,
Sure, it ended up being popular and usable, but at the cost of taking the art of describing how software is installed back in time.
Could it be that these things are mutually exclusive?
Anecdotally, I've tried to learn Nix for several years. I learned to write Dockerfiles in hours.
Don't confuse expedience with usability.
Out of interest, what did you find difficult? I ask because I spent last weekend setting up NixOS, and the only things I found difficult were discoverability and documentation (I spent a lot of time just reading through the source code).
Edit: Thinking about it it's very possible that this is just a case of me not knowing what I don't know, since I'm not currently trying to do anything particularly complex.
The language is not too bad, but how it should be used is the missing part. Like a cookbook or a properly populated Stack Overflow site.
When I was on Linux, I liked to distro hop few times a year. Whenever I got back to NixOS, I felt that I had to relearn everything, and that there was nothing to grasp on to. I always take that as a bad smell.
The other solution is to implement a totally different mechanism ( ACL + some other thing, Capabilities ) but then you lose retrocompatibility.
A process has supplemental group membership. If you were only limited to the above, then, yes, you would need to share resources through world or utilize a special broker if you wanted multiple privilege domains for a service.
But with supplemental groups you can be as fine-grained as you want. Just create a group for each domain with a unique membership set and make that a supplemental group for each user with access to that domain. When you want to share a file, just chown and chmod appropriately.
To see how this works in practice, look at BSDAuth (OpenBSD's alternative to PAM). Rather than using dynamically loadable modules for authentication methods--which effectively means that any program that wants to use authentication services needs root permissions--BSDAuth uses setuid and setgid helpers in /usr/libexec/auth.
$ ls -lhd /usr/libexec/auth/
drwxr-x--- 2 root auth 512B Mar 24 13:23 /usr/libexec/auth/
$ doas ls -lh /usr/libexec/auth/
total 372
-r-xr-sr-x 4 root _token 21.3K Mar 24 13:23 login_activ
-r-sr-xr-x 1 root auth 9.0K Mar 24 13:23 login_chpass
-r-xr-sr-x 4 root _token 21.3K Mar 24 13:23 login_crypto
-r-sr-xr-x 1 root auth 17.2K Mar 24 13:23 login_lchpass
-r-sr-xr-x 1 root auth 9.0K Mar 24 13:23 login_passwd
-r-xr-sr-x 1 root _radius 17.1K Mar 24 13:23 login_radius
-r-xr-xr-x 1 root auth 9.0K Mar 24 13:23 login_reject
-r-xr-sr-x 1 root auth 9.0K Mar 24 13:23 login_skey
-r-xr-sr-x 4 root _token 21.3K Mar 24 13:23 login_snk
-r-xr-sr-x 4 root _token 21.3K Mar 24 13:23 login_token
-r-xr-sr-x 1 root auth 21.0K Mar 24 13:23 login_yubikey
In this scheme, for any user who you want to grant permission to test basic system authentication you add the auth group to their supplemental group. Et voila, they can now use the framework, but without root permissions. For mechanisms that you may want to grant separately you use a separate group (e.g. _radius).Like many bad ideas, Docker basically won because you can use it to isolate and package pre-existing, broken software. And because it's easier to find examples of how to [abuse] Docker for sharing files across security domains than on how to use supplemental groups to do this.
On BSD when you create a file the group owner is set to the group owner of the directory. From the BSD open(2) manual page:
When a new file is created it is given the group of the
directory which contains it.
But on Linux the [default] behavior is to set the group owner to the creating process' effective group. From the Linux open(2) man page: The group ownership (group ID) is set either to the
effective group ID of the process or to the group ID of the
parent directory (depending on file system type and
mount options, and the mode of the parent directory, see
the mount options bsdgroups and sysvgroups described in
mount(8)).
The BSD behavior is more convenient when it comes to using shared supplemental groups because you can often invoke existing programs with a umask of 0002 or 0007 and files they create (without subsequently modifying permissions) will usually have the desirable permissions--specifically, writable by any process which has a supplemental group of the directory they're located in. For example, with BSD semantics Git technically wouldn't have needed a special core.sharedRepository setting. But because of SysV semantics on Linux Git had to be modified to explicitly read the group of the parent directory and explicitly change the group owner of newly created files.Do we really need a non-OCi image format and tool chain here?! I get immutability benefits, district benefits, and isolation. having a completely separate toolchain makes no sense to me...
Containers will be in this role, but fragmentation will only make that process slower and more painful.
I for one salute our future k8s kubelet overlords.
This doesn't use a new competing image format, but rather uses disk images or tarballs, which have existed longer than OCI has.
The toolchain also predates OCI since the underlying technology is really basically nspawn (and various other systemd.service options), which have been a part of systemd since before OCI stabilized. This is simply a new coat of paint on something that has been there for a while.
> Containers will be in this role, but fragmentation will only make that process slower and more painful.
Sure, so we shouldn't have let docker do anything and instead used LXC, since LXC already did containers, already had its own format and toolchain, etc. That's the exact same argument you're making now.
OCI is also not really a good standard. It has exactly one usable implementation (and only in Go, not even C where it's easy to link against and use it from other languages), and since systemd has already defined various things for services (such as seccomp filters etc) which overlap with what OCI does, it makes almost no sense for systemd to begin using it.
With regards to image formats, I actually think that we do need to improve the image format significantly and not just stick to tried and tested stuff. Tarballs are (quite frankly) simply the most awful format to use for container images (disk images are a close second). It's a shame that everyone uses them. I am working on a blog post to better explain what I mean by this.
And yeah, runc being written in Go is probably one of the most frustrating things about working on it. I often say that one of the worst things you could decide to write in Go is a container runtime -- it's a little bit odd that people keep doing exactly that. runc actually isn't entirely written in Go -- it has to have a fairly substantial amount of C because the Go runtime cannot play well with the delicate dance required to set up a container properly.
All of that being said, containers really did need to be standardised otherwise the container wars wouldn't have ended as nicely as they did. But that's not to say that the OCI doesn't have issues. It definitely does, but I'm hoping they can be fixed over time.
[ I am one of the maintainers of the "exactly one usable" OCI implementation, and have been working on OCI stuff since its inception. ]
This is what I was referring to when talking about their breadth of work. I've collaborated with them quite a bit, and it's always amazing working with them on a hard problem.
It was not easy ;-)
While this is interesting, the containers that I'd appreciate, are: multiple versions of programming languages & packages to be installed by each user without affecting the system globally. Every language has their own separate answer (virtualenv/venv, rbenv, node version manager etc).
Just today I struggled with "brew install python3" which complained with "Error: python 2.7.14_2 is already installed" and the only option offered was to upgrade. Whereas what I wanted was parallel installations of python 2.7 and python3. :-/
It's quite frustrating to me, because after I've been using Nix & NixOS for about 3 years now, there's no way I'd go back to the insanity out there. I never had the amount of stability, control, and predictability on any other platform, from configuring my OS, editors, WM, shell, random ricing, or building VMs, Servers, Docker containers, production deploys, development environments, there's really not much Nix can't handle. But getting familiar with Nix took some serious effort I'm sure not many people can afford. I begin to understand how frustrated people in the Lisp and Smalltalk communities must feel while people slowly reinvent everything they've enjoyed for decades (and Nix is still in its teens).
Teaching people about the value of having just a single package manager, no matter what OS or language they're using, is mostly futile because there's so much value attributed to being "mainstream" that now we have hundreds of mainstream package managers with varying degrees of sophistication, security, predictability, and ability to _not_ try and take over your system.
I use it to build and develop Go, Ruby, JS, Crystal, Elm, Mint, Haskell, Perl, Bash, VimL, Elisp, Guile, and whatever else comes my way. All I have to do to get a working and isolated dev env is go into that directory and let nix-shell do the rest. On NixOS you also get nixos-container with all its benefits, so I can test whole networks or just spin up some DBs, all behaving exactly the same they'll be once deployed by simply reusing that configuration.
Of course it also manages Systemd, so I guess we'll just write yet another function to slap a checksum on those containers and control them as well, it's just a shame that humanity loses so much time reinventing the wheel every few months.
On the other hand, there's still a lot of opportunities for start-ups in that space, like https://nixcloud.io, https://www.tweag.io (they just hired Eelco Dolstra so he can work full-time on the Nix core), https://www.packet.net (they sponsor almost all our new CI infrastructure), or https://vpsfree.org/ (first one to offer their own NixOS based distro specifically made for VPSs), which are all doing well.
the problem is OpenGL and Nvidia. NiX doesnt want to deal with NVidia's closed source drivers, and i can't blame them, i got so depressed working on the issue, that i gave up looking into it. there is no good solution. there are only kludges and kludges of kludges.
I would like to use Nix just to manage my local development environment on Ubuntu, not to replace my entire operating system. Is that something the developers are still interested in? Is there documentation?
Multiple versions of programming languages don't require docker, containers, or Nix, that is utter insanity.
Maybe it’s nice to think of a common solution to this!
https://github.com/asdf-vm/asdf
Then, there is something wrong with either brew or MacOS or both because we can install both Python 2 and 3 in parallel on any Linux with the system package manager.
What language runtime do you have in mind? I haven't seen one that doesn't support this. A normal user should not be allowed to affect the system, so installing language runtimes under your home directory is normal and expected.
Most languages use an environment variable to point to the packages being used, so it should be possible to have multiple versions installed under your home. This is how the language developers work. They need to have many development versions installed for testing.
Containers are useful for a lot of things, but a much more complicated set up than just installing a couple of language runtimes. You should have no problems doing that the normal way.
https://coreos.com/fleet/docs/latest/launching-containers-fl...
No longer actively maintained is a problem, but I think it still has a quite a few users.
If I just want to download some semi-trusted source code. And Configure; make; run; it in an environment where it can’t access anything unless I explicitly white list it (in particular, it should not be allowed to access the internet or any network resources, but I should be able to send http requests to it)
Which variant of all these container stuff would make things simplest to set up?
Docker seems way overkill (and seems to have the wrong defaults for this need anyways) but handcrafting things with iptables or ip netns seems a bit to low level. (And, well, not “contained”... I would prefer something declarative with automatic setup/teardown)
Ubuntu snaps looks like an interesting middle ground. But it also looks completely dead.
Any tips?
Edit: should probably rtfa before asking. it actually looks like a good fit ;)
https://www.projectatomic.io/blog/2017/06/introducing-builda...
It's a command-line tool for building and running container images.
PS: log events shouldn’t be perceived of as lines of text but structured data messages, so that they don’t need special parsing.
The separate journald log server available as part of the systemd project can generate qrcodes. journald can cryptographically sign logs so that future log tampering can be detected. This requires two keys, a sealing key and a verification key. The verification key must be stored off server. When the keys are generated systemd can display a qrcode to allow easy recording of the verification key.
Having a QR code parser as part of something that's tangentially related to the init system in concept but deeply embedded into systemd in practice is exactly what the post you were replying to is complainig about. It's unnecessarily complex.
>They are more mature
what does that even mean? that they are older? Similar to "modern XYZ" this phrase doesn't mean anything.
You seem to repeat what other people say like the "QR code" without even looking it up.
https://stackoverflow.com/questions/32924335/using-systemd-w...
> The journal is the only mandatory component of systemd outside PID1.
Seems pretty deeply embedded to me.
> what does that even mean? that they are older?
They are older, better understood, and there are fewer substantive bugs (e.g. malformed blob = run service as root).
> You seem to repeat what other people say like the "QR code" without even looking it up.
https://lists.fedoraproject.org/pipermail/devel/2012-October...
Not sure what else you're doing with a QR library, except maybe building tetris (which seems like a bad idea for something purported to be an init system).
You used the same argument ("They are more mature") again. It's still not saying anything.
Sure it is. An init system should, first and foremost, be stable and well understood. By virtue of its size and youth, systemd is not (at least not by devs and end users).
I have noticed that systemd is "stable" and "well understood" because I have no problems with it. But would that convince someone else?
Bugs like this:
https://www.theregister.co.uk/2017/07/05/linux_systemd_grant...
https://github.com/systemd/systemd/issues/6077
https://github.com/systemd/systemd/issues/9449
https://github.com/systemd/systemd/issues/9079
Indicate that systemd interactions aren't particularly well thought out.
https://github.com/systemd/systemd/issues/8730
Non-deterministic behavior is exactly what I'd strive to avoid in an init system.
Stuff like this:
https://www.agwa.name/blog/post/how_to_crash_systemd_in_one_...
Indicates both poor understanding of how an init system should work (including basics like privilege separation) and poor stability.
Stuff like this:
https://threatpost.com/linux-systemd-bug-could-have-led-to-c...
Well, I for one do not want my init system to be network accessible.
Please help me understand this line of reasoning. Someone has a list of reasons as to why they believe systemd is not a good fit for them, and something else works better. They get told "prove it isn't a good fit for you". They prove this with some line of argumentation. They get told "that is your opinion only, and that doesn't count". They provide links to show many other users have the same issues, and they feel these issues have not been given due consideration. They get told "typical systemd haters, they always provide a link dump with broken stuff, it doesn't mean anything"
This has pretty much become the default template for the systemd pro/con arguments, and is unwinnable. FWIW, my problem with systemd is that it broke the concept of free choice in my environment. It was harshly shoved down everyone's throat, through politics rather then the usual meritocratic methods. I appreciate that this was a commercially advantageous position for the various distro maintainers, which lead me to cancel all my (thousands) of distro support agreements.
It is the political approach that was taken that gives rise to serious suspicions for me. If the system cannot stand on its' own feet from the meritocratic perspective, and needs to resort to all kinds of political games to gain a substantive foothold, my default position is not one of trust.
The "la la la i-can't-hear-you" systemd fanboy troupe response whenever someone attempts to make a well-supported argument against systemd only fuels my distrust of the whole systemd story.
As I mentioned, I vote with my feet and wallet. I use Alpine wherever possible.
Waving your hands and suggesting everything is going to have bugs isn't much of an argument. It's unlikely that sysvinit would suffer from any of these bugs as there's simply a much smaller vector of attack than systemd by design. You could knock me over with a feather if arbitrary network traffic could compromise your system via sysvinit. The other class of bugs are ones where sysvinit is the defined behavior and systemd is simply deviating in unexpected, unintuitive, and undocumented ways.
Put another way, as a couple of anecdotes at $day_job, the boot time benefits (the one thing people always seem to bring up in defense of this ever-growing behemoth) were eaten up months ago in time troubleshooting, transitioning, and working around its corner cases.
More humorously, the words "fucking" and "goddamn" used to be followed by the name of some internal program known for being clunky. After getting up to Ubuntu 16 and Cent 7 in the whole environment, those words tend to be followed by "systemd".
Same for timers. Debugging not running cronjobs is a pain in comparison.
Writing sysv init scripts is so fucking shit, people started to dump everything in /etc/rc.local, especially for earlier RPi versions of Debian.
Do you know about Kafka, Spark, and Jenkins? They're all Java but they each have their batshit fucking insane quirks on starting up.
One of them runs in foreground. One of them runs in background.
Jenkins (actually as Hudson) had it's own fucking Unix daemon-isation inbuilt. For a Java program!
Nuts!
Thankfully pretty much popular Java server has a systemd unit file written for it which avoids any shell and Exec's Java directly.
For me it’s basically magic that I can now write a simple declarative file that describes how to run an application, and from there it works with all of the expected features.
I don't know if systemd is better (I've never managed a server fleet with it), but Upstart certainly wasn't good enough.
I hated systemd before it was cool to hate systemd. I never had a problem with sysvinit as the root casuse, but I had many problems with systemd as the root cause.
the real choice was between upstart and systemd, and upstart was clearly superior, especially for common use cases.
Both init systems and container systems are concerned with dependencies, logging, and managing the lifetime of processes, and systemd is way better at the last two than Docker, so I'm optimistic.
clearly, this is a messy situation. lots of different pieces doing different things in different ways. in order to harmonize this, and simplify it, I believe we need to have systemdd, a daemon for systemd, so that all of it's various pieces can be in one, neat, centralized location.
begin 644 - M0F%S:6PZ($QI<W1E;BP@9&]N)W0@;65N=&EO;B!T:&4@=V%R(2!)(&UE;G1I M;VYE9"!I="!O;F-E+"!B=70@22!T:&EN:R!)(&=O="!A=V%Y('=I=&@@:70@ )86QR:6=H="X* ` end
H4sIAAAAAAAA/4vML1UvSlVIVEjJTM7WAwBgPKpvDgAAAA==
I've generally had a positive impression of the systemd developers' ability to design systems, and to communicate those designs. In most cases, I do not disagree with the choices they have made. There have been a few errors, and there is a great deal of bad blood. I think both that the software is generally good, and that there's nothing wrong with informed criticism.
Everyone brings their own here (e.g. http://phusion.github.io/baseimage-docker/) like supervisord, runit,etc.
What would have been ideal is for a docker-compatible systemd to run inside the container.
IMHO the maintainers are bent on creating a competing standard (like this one) and don't want to build anything that brings the advantages of systemd to the Docker ecosystem.
Please provide evidence of that. I will provide counter-evidence, proving that is not the case unless you have very compelling evidence.
1. https://lists.freedesktop.org/archives/systemd-devel/2014-Ma...
Poettering writes: > To say this explicitly: we are really interested in making sure that systemd runs out-of-the-box in containers, and docker is just one implementation of that.
2. A PR to make systemd-logind work better in a docker container is accepted with no issue: https://github.com/systemd/systemd/pull/4154
3. Running systemd as pid1 in docker (privileged or not) works, so clearly the maintainers didn't do a good job of preventing it from happening. (just google, there's plenty of info out there on how to do this)
I'm sure there are many other examples too, but this is with a quick glance around.
I'd be interested in your reference before you spread such strange FUD.
Lennart's reply:
We had support for this in systemd in the initial versions, but since nobody was using and testing this, and the semantics were different from the system version in subtle ways we removed support for it.
Also note that systemd requires kernel 3.7 as minimal version right now. You cannot run it on older kernels, and hence really old operating systems anyway.
Sorry, but this is nothing we want or can support!
I admit that it is 2 years old - and if it has changed, then that would be nice. But there is an explicit reply there.. and I'm not spreading FUD. There's nothing more that I would love than to replace all the supervisord stuff in Docker with systemd.
That bug doesn't appear to be closely related.
You wrote: "... it would be great if there was a way we could use systemd+journald in older OS to run applications".
That's what Lennart replied to. He didn't dismiss using it in containers as pid1, he dismissed running it on arbitrary older linuxes as pid != 1... those are completely different scenarios.
I do not find any support in that link for your original comment, so unless you have more evidence or can explain how I'm misunderstanding your bug, I still consider you to be spreading FUD.
Because I have explicitly mentioned this in the bug. Quote unquote
Docker is only making things even worse - the recommended way of using Docker is shifting gradually to be using poorly-suited PID 1 processes like supervisord.
Additionally, there are other places where I mention Docker as a use case. Granted that English is not my first language, but I do not think there was confusion on this being around Docker. In addition the bug title is " Standalone version of systemd to act as a process control system" ...Which was as generic as I could make it.
Running an arbitrary job-runner on older system is an implicit use case....But by explicit clarification, I brought in the Docker use case as well. I also explicitly mentioned that we are prepared to use newer kernels within Docker if that is important. P.S. Already, a lot of us work with upgraded kernels because of Docker and this is already acceptable.
I WANT to use systemd - but I stand by my claim that this bug is sufficient illustration that there is no intent to support the docker ecosystem. I do not wish to start a flamewar, but I categorically reject your claims of me spreading FUD as deliberately malicious.
Again, all of those things are entirely unrelated to it running as a pid1 in a container, and those parts of your original bug report are what Poettering responded to.
You're reading way too much into that one bug. In fact, you could already run systemd in docker before you posted that bug if you go back to my first link to a mailing list thread about the very subject from 2014.
Your main bug report made it sound like you were asking to run systemd as pid != 1, which is unrelated to running systemd in docker.
Your clarification here only convinces me that you hold a grudge against systemd and wish to interpret benign comments in a bad light because, no matter how many times I read that bug, I cannot understand your claim that the original posting and Poettering's rejection of it is in any way related to running systemd in docker.
Here's one piece of evidence though[1]. Effectively systemd enabled a security feature for the host's /dev/console and when we said it wasn't necessary for containers (because /dev/console is not a real console) and it actually broke runc, Lennart said that we should fix runc.
This wasn't a security feature that got enabled. This was an expectation of the long-standing semantics of /dev/console, which as M. Poettering pointed out long pre-date systemd, being broken by a container manager.
So yes, it's right that the container manager be fixed so that it is possible for every open file descriptor for /dev/console to be closed and then the device re-opened again. Such semantics have been around longer than Linux itself has; systemd is written to expect them. /dev/console is not supposed to magically vanish/become inoperable once all currently open file descriptors for it have been closed. Quite a lot of other softwares, including everything that uses openlog() with LOG_CONS, expect this of /dev/console too.
Indeed, /dev/console is one of the very few device files mandated to exist by the Single UNIX Specification (XBD part 10). The container manager was actually setting up an execution environment that is not POSIX conformant.
And I observe that indeed said container manager did get fixed.
Of course we fixed it, and of course you can argue that Lennart was correct (I still think that having SAK protections in a container is nonsensical but that's all water under the bridge). I obviously agree that our /dev/console handling was incorrect. In our defense, /dev/console doesn't actually make much sense in a container since the purpose of /dev/console is to access the physical console not the current PTY -- so anything we put there would still be "wrong" from the standpoint of POSIX. But you can't just ignore /dev/console because then a bunch of programs don't work. I could also go on about how it was also a Go stdlib issue because we'd assumed io.Copy "did the right thing" but it turns out it really doesn't handle any form of interruptions properly. But I'm sure you're not interested in that discussion.
The point is that I agree it was fixed, and I agree that fixing it in the container runtime was overall correct. But that wasn't the point I was making -- it was that there have been examples where systemd has made a change that broke running inside a container and they were not willing to make concessions for container runtimes. Which is what GP was arguing about.
I do have plenty of other examples (cgroups are particularly fruitful for systemd bugs that won't die), but they aren't really related to running inside a container.
[ I'm not a hater of systemd, or Lennart. I actually really like having a declarative service manager. My frustration comes from having to deal with it when developing system tools that don't want to be tightly coupled with it. That's where systemd really starts to get ugly to deal with. ]
> M. Poettering
What does the 'M' stand for? His first name is Lennart.
In French, "M." is short for Monsieur.
https://coreos.com/rkt/#features
What were the maintainers supposedly refusing?