HNHacker News
TopNewBestAskShowJobs

siebenmann

168 karma · joined August 20, 2011

http://utcc.utoronto.ca/~cks/space/blog/ among other places
submissionscomments
siebenmann··on How we sort of automate updating system packages across our Ubuntu machines
Putting safe in quotes is accurate, because our updates are not limited to truly safe packages at all for the simple reason that the Debian package format makes all package updates potentially dangerous. Any package update can decide to ask you new questions and there is no guarantee that the default answer is what you actually want, because it's up to the package author to decide.

As a policy thing, we also want to know what packages were updated when and broadly to control this, and we have enough machines that this information needs to be aggregated. Even if unattended-upgrades was totally safe, we'd want to only trigger it only on demand, see the package list in advance, and get an all-machines summary at the end.

Life would be a lot easier if apt-get updates could be more controlled, so for instance you could say 'only update this list of packages'. But apt-get doesn't want to do that; updates are global and using 'apt-get install' to update specific packages (normally) has side effects, like marking them as manually installed.

(I'm the author of the linked-to blog entry. For reference, we currently have 119 machines that get updated through this system.)

siebenmann··on The Go runtime scheduler's way of dealing with system calls
An OS thread M can run on any available P. While there are some caches associated with each P, Ps are fundamentally there to insure that only so many CPUs worth of Go user code is ever running at once, so the important thing is that an M that wants to run user Go code has some P, not a particular P. Ms claim and release Ps as they go in and out of running Go user code, but I believe they don't release and then re-acquire a P as they switch between goroutines.

(I believe the actual implementation treats Ms as a sort of secondary thing. For instance, I think that the local list of runnable goroutines is attached to the P, not to the M. At one level, the M is just a context for running things on Ps.)

In the optimistic case when the system call blocks for too long, the M is unpinned from the P it was using and continues to sit in the system call (the Go runtime doesn't attempt to interrupt the system call itself). If there is another runnable goroutine and there are no free M's, the Go scheduler will create another M to run the goroutine on the now-free P. I think that the runtime directly allocates the free P to the newly created M rather than letting the new M try to contend with other things for the P, but I'm not sure.

I don't think there's any limit on the number of Ms (OS threads) that the Go runtime will create, but I haven't checked the code carefully. Idle Ms are reclaimed under some circumstances.

(I'm the author of the linked-to article.)

siebenmann··on The Go runtime scheduler's way of dealing with system calls
This should work fine. The goroutine making the system call that touches the NFS mount will consume an OS thread (an 'M' in Go terminology), but it will release its hold on other resources. Go uses as many OS threads as necessary to cope with running user code and doing OS system calls and so on (and starts new ones on demand).

If you had lots of goroutines do lots of things that stalled on hung NFS mounts, you would build up a lot of OS threads (all sitting in system calls) and might run into limits there. But that's inevitable in any synchronous system call that can stall.

(I'm the author of the linked-to article.)

siebenmann··on How modern Linux systems boot
Belatedly: that was a very interesting read on the history of this in both Linux and Unix more generally. Thank you for the link.
siebenmann··on Linux Load Averages: Solving the Mystery
This is great work in general and excellent historical research.

As an additional historical note: in Unix, load averages were introduced in 3BSD, and at that time they included processes in disk IO wait and other theoretically short-term waits that weren't interruptible. This definition was carried through the BSD series and onward into Unixes derived from them, such as the initial versions of SunOS and Ultrix. At some point (perhaps SunOS 3 to SunOS 4, perhaps later), the SunOS/Solaris definition changed to be purely runable processes.

(I'm not sure what System V derived Unixes such as Irix, HP-UX, and so on did, and their kernel source is not readily available online for spelunking.)

As of early 2016 when I last looked at this, the situation on FreeBSD, OpenBSD, and NetBSD was somewhat tangled. FreeBSD load average only included runable processes, but NetBSD and OpenBSD counted some sleeping or waiting processes as well.

siebenmann··on Microsoft Is Now 'Open by Default', Says Xamarin Founder Miguel de Icaza
If people are interested, the specifics of our current OmniOS environment are here: https://utcc.utoronto.ca/~cks/space/blog/solaris/ZFSFileserv...

We have been broadly happy with it and it has been quite stable after some initial teething problems.

siebenmann··on Why improving kernel security is important
The general SELinux issue is a complex subject. My short form take is that regardless of its potential in theory if implemented nicely, in practice SELinux as deployed has consistently prioritized mathematical perfection (and yelling at people) over practical usability in the field. The real result of this has been less security than would have been achieved with a less perfect but more usable system because in the field SELinux does not degrade gracefully and so many people turn it off entirely. Some number of systems are quite secure (assuming no leaks in SELinux itself); many other systems are not secure at all. This is a bad outcome (unless you decide that only people who are dedicated enough to use SELinux really matter and everyone else is 'unprofessional' or the like), and I don't like it. I want a better outcome, one with more security that I can actually justify deploying, one where more daemons and programs are hardened to some degree even if it's not a huge amount.

(At this point the OpenBSD pledge() work is looking attractive, although there are real organizational issues that would make it hard to do in Linux.)

Perhaps one can get to a better future with SELinux by having people build and ship systems for doing little SELinux configurations for daemons or systems that read daemon configuration files so they can automatically label directories and files for your system, or any number of other user friendly ideas. But we've had something like a decade of SELinux and its usability problems at this point and it hasn't happened yet. It's hard to avoid the obvious collection of conclusions.

siebenmann··on Why improving kernel security is important
I believe that what Matthew Garrett is talking about is mostly different from SELinux and AppArmor and so on. Those are all kernel features to harden user-level software in the face of vulnerabilities. Garrett is (mostly?) talking about internal kernel features to limit the damage of kernel vulnerabilities.

(Many of the grsecurity changes are kernel hardening, for instance; they don't directly affect user level code.)

siebenmann··on Why improving kernel security is important
Our viewpoint is ultimately pragmatic: at the moment, both SELinux and AppArmor appear to be too much work for the potential benefit they offer in our environment. We could spend a great deal of time configuring both of them in order to make things work, and they would still probably not do anything much for us in practice.

(Remember, both SELinux and AppArmor are secondary defenses, not primary defenses; they potentially limit the damage if your system is already partially compromised.)

In part this is because at least SELinux only really works easily if you put everything in what the distribution considers its standard location and run things in the standard way. The moment you deviate from this, you wind up having to research an increasingly large number of file and executable contexts and (re)label an increasingly large number of files. And I'm ignoring NFS here, which we use heavily (I doubt NFS files can easily have SELinux attributes, especially when they live on non-Linux NFS servers).

Security is always ultimately about pragmatics. You have X amount of time to spend on security in all its aspects, and you need to use this time efficiently, to gain as much security from it as possible. Our judgement is that use of our security time configuring SELinux does not have a particularly high payoff.

(For clarity: I'm the author of the entry that mwcampbell linked to.)

siebenmann··on The origins of chroot()
I suspect that the answer is 'badly'. V7's chroot() seems to be more than a little bit of a fast hack, one that was good enough for some things but not at all comprehensive or problem free.

(There are PDP-11 emulators and V7 disk images available through tuhs.org, so intrepid people who want to find out for themselves can actually try this out on a live V7 system.)

siebenmann··on The bad side of systemd: two recent systemd failures
Almost any time PID 1 segfaults, it's PID 1's fault. With PID 1 being systemd, that makes it systemd's fault here. Based on gdb stack backtraces in the Fedora bug report and looking at the systemd code, it seems like some sort of memory corruption or overwrite, perhaps a use after free issue. The segfault itself comes from dereferencing a clearly invalid pointer and said pointer was obtained by dereferencing a structure field through another pointer, so you'd get exactly this result if the structure was overwritten with other data at some point.

(In my grumpy sysadmin view, it is PID 1's fault even if the distribution is doing odd things around PID 1. Init processes need to be absolutely rock solid and extremely defensively coded, precisely because the world basically dies if they ever fall over.)

(I am the author of the original post.)

siebenmann··on Persistant and unblockable browser cookies using last-modified HTTP header
I've written server-side software that generated a meaningful Last-Modified header (it's very useful information) and also did an exact comparison instead of a time-based one. I did the exact comparison because I realized that it was very hard to guarantee correct results otherwise (correct being that I never 304'd a request unless the client already had the exact version of the page that I would have served). The problem is that there are a number of situations in a web server that can cause page modification time to go backwards. For example, several ways of doing rollback to previous versions of content will also roll back the page time, such as simply renaming an old version of the file to the current name.

To do correct time-based If-Not-Modified comparisons you really need to guarantee that all changes moves your Last-Modified time forward, no matter what. My view is that this is surprisingly hard once you start looking at corner cases. Certainly it's not something that a web server that serves general file content can ever guarantee; there are too many ways to shuffle files around behind the web server's back.

← PreviousPage 2 of 2