Critique of Microkernel Architectures – Is Linus Right? (2004) [pdf]
cse.unsw.edu.au
cse.unsw.edu.au
For those not familiar with the DEC Alpha, it always ran what was basically a hypervisor/microkernel in firmware (called PAL code). The OS kernel would actually run in normal user mode and make upcalls to the PAL code, and the PAL code could emulate an arbitrary number of protection rings (two for Ultrix, several for VMS). VMS and Ultrix required different versions of the PAL code to be loaded, and Linux on Alpha used the Ultrix version of the firmware.
It seems to me a shame that the DEC Alpha PAL code isn't the standard way firmware is done... ship a nanokernel/hypervisor and some basic drivers in ROM, and have some upcalls to replace drivers (or even the nanokernel/hypervisor). The DEC Alpha was a powerhouse in its day, and seems to have suffered very little from the PAL code abstraction.
Gosh, I never realized, as a former Keeper of Alphas for scientific work, previously dependent on the GEC 4000 series <https://en.wikipedia.org/wiki/GEC_4000> for physics. I wonder how it compared with OS4000 Nucleus.
The 4000s made competing VAXen look a bit silly for our sort of data acquisition and analysis. Since they were pretty important for my publication record, it always amuses me to see microkernel-ish systems rejected out of hand. A colleague disproved the conventional wisdom that you needed a "spectrum database" to manage the data by just trusting the filesystem to manage directories. More recently I recall userspace filesystems being called generally toys (?), notwhithstanding huge, high performance PVFS2 installations, for instance.
> The DEC Alpha was a powerhouse in its day
Yes, and tragically under-appreciated in my sphere. I tried in vain to get the maintainer of the principal data reduction program in a different field to replace the bottleneck of the disk-based sort (written for a PDP-11). Each image would normally fit in cache and the whole dataset largely in memory, but that somehow wasn't relevant, contributing to "the computer is slow". (I think the sizes were 10MB and 1GB.)
There are now a few secure microkernels. It really is possible to fix the Mess at the Bottom. If the kernel isn't secure, nothing else can be. Patch-and-release security just doesn't work any more. The serious attacks today are from organized crime and governments, not script kiddies, and they develop their own zero-day exploits.
Qubes OS is the hypervisor department's version of the "compartmentalize OS for security" idea. The microkernel dept should come with their own credible version if they want to win.
I wonder how feasible a modern, graphics-capable POSIX OS built on a core of sel4 is?
They let the CPU control where device drivers are allowed to write.
I don't know what the perf costs are like in practice though since you would need to either be able to reconfigure the IOMMU before requesting a device to write into a user process' address space, or you would end up writing it into the device manager's address space and copying it out via IPC.
Plan 9, A Distributed System Dave Presotto Rob Pike Ken Thompson Howard Trickey 1991 (I think)
'nuff said
On the other hand, it doesn't do all that much for you, there's a lot of kernel type behavior you have to implement above it. But you are starting out with a relatively very small Trusted Computing Base (TCB).
I do wonder if we could do the opposite of microkernels. A microkernel would be free if it was running a (pre/recompiled ?) VM and things like memory mapping/protection/... would be compiled into user programs before execution. After this these could run without any actual user/kernel space isolation and just eliminate things system call overhead entirely.
In theory this is possible, in practice, in low-level optimized piece of software, like the kernel, this is never true. This just can't work, at least for the next few decades, but potentially forever.
But to me, a microkernel isn't about performance. It's a complement to monolithic kernels. When I am writing source code, or typing up an important document, or archiving files for preservation ... I want to be doing this on a microkernel OS.
I have had my FreeBSD system crash completely due to the video card driver (from nVidia, of course) accessing unmapped memory. That's completely illogical. Drop me to a text terminal, or hell, force me to SSH in to unmount and reboot cleanly. But don't drop the entire system. Yet with a monolithic kernel, how can you be sure the video driver didn't corrupt some other part of the kernel space?
Likewise, if I want to play the latest video game, or I'm using it as a render farm, or to mine bitcoins, or I need to run a web server that needs scalability more than stability ... then I want to be doing that on a monolithic kernel.
We have most OSes that focus on raw performance, then we have NetBSD that focuses on portability, and then OpenBSD that focuses on security. Where's our (non-toy) OS whose primary focus is on stability? So far, Minix 3 looks like the most promising option, but it has no enterprise file system like ZFS, and is severely lacking in manpower.
For the latest video game, with a microkernel (and an IOMMU keeping the video card from trashing memory), when switching to fullscreen mode, your OpenGL library could request that your video game process become the video driver, with no loss of security or stability. That gives you lower overhead access to the video card than a monolithic kernel.
For compute-bound tasks such as bitcoin mining, a small realtime microkernel will give you better instruction cache utilization. (Back in university system architecture class, I ran cache benchmarks under Linux, QNX, and Win2k on my triple-boot desktop, QNX gave me the best cache performance.)
If you need to run a high-performance webserver, if you need absolute performance, run it in-kernel. If not, a microkernel (and a properly configured IOMMU) allows you to safely make your process the network driver as well as filesystem and disk drivers for your dedicated www data drive. Now, you'll want a second network port with its own chipset (perhaps on its own card) for being able to ssh in if the webserver crashes, but you'll have higher performance than a monolithic kernel.
In short, microkernels along with appropriate hardware protections allow you to safely make your device drivers just ordinary libraries that run inside your critical processes, giving you higher performance without sacrificing security or stability.
EDIT: I should also point out that Oracle, Postgress, and other high-performance RDMSes go through a lot of work to do "kernel bypass". A microkernel with minimum cache footprint would better get out of their way and more safely turn over more functionality to the application.
I don't think this is necessary with modern NICs and SR-IOV. I imagine you'd just configure one virtual function for the web server, and one for management. It's pretty common already.
Similarly you might have a RAID controller capable of LVM-like "partitioning" of the disk array so that you can present two virtual disks that can be dedicated to whatever purpose you need. I don't know if these actually exist yet, though.
Linux success is related to the ease people could copy stuff from commercial UNIXes into GNU/Linux, not the kernel architecture.
Neither iOS/OS X or NT manage to get the robustness, security or compartmentalization advantages of microkernels. I guess they get modularity and dynamic loading, but so does Linux. So Linux, X and NT are all mongrels, with iOS/OS X having microkernel roots and Linux having monolithic roots, and NT being a ground-up designed mongrel.
Both Mac OS X/iOS get it partially, by having kernel level RPC to communicate between modules and having moved a great part of their driver infrastructure to user space.
Let me know when your graphic card crashes don't require a reboot any longer on GNU/Linux.
So is there a the mechanism in OS X and/or Windows that prevents a graphics card hardware or driver from corrupting OS state and enables robust reset of card and driver? Or if it's a 90% solution, it significantly better because of some microkernelesque features missing from Linux?
https://msdn.microsoft.com/en-us/library/windows/hardware/dn...
https://msdn.microsoft.com/en-us/library/windows/hardware/ff...
A driver crash will just force it to be reloaded.
Likewise using kernel level RPC adds extra validation layers via the data marshaling than a simple function call would do.
As for OS X, only drivers for disks, network controllers, and keyboards are required on the kernel and use mostly Mach calls, not BSD ones. Anything else can be exposed via so called nubs to user space.
https://developer.apple.com/library/mac/documentation/Device...
https://developer.apple.com/library/mac/documentation/Device...
And third party articles: https://www.blackhat.com/docs/us-14/materials/us-14-vanSprun... http://bsodtutorials.blogspot.fi/2013/12/timeout-detection-a...
So there's a user-space part to the GPU driver and a kernel-space part, much like on Linux.
For the recovery functionality it sounds like the graphics card's kernel-side GPU driver just registers a callback that is used by Windows when it thinks the GPU or driver is stuck, but there aren't any special arrangements to make this robust against the driver corrupting OS state or other hardware. Same kind of mechanism could be implemented in Linux.
I'll save those OS X links for later when I have time to look into that one!
Specially since it requires a stable kernel ABI.
Although depending on how bad systemd is in the wild, a sufficient critical mass for a stable device driver ABI just might be come to be.
Indeed; we were greatly hopeful for it in the days of too many proprietary UNIX(TM) versions, and it began very well, but in truth it started going downhill after the 1st Service Pack for NT 3.51, and was indeed a "mongrel" in the decade of NT4/2000/XP when the graphics driver was moved into the kernel (per pjmlp's comment in this subthread about it being moved back out with Vista, and given how much people avoided Vista for unrelated reasons, you could add another 3 years until Windows 7).
The remaining ones have reduced themselves to copying those proprietary UNIX(TM) features into GNU/Linux.
As far as I am aware, IBM is mostly busy making GNU/Linux run on their mainframes.
Aix has lots of cool features, but it is still a fairly typical UNIX.
OS/400 kernel is quite cool, using bytecodes as executable format with a kernel level JIT and being object based. But there is very little more to innovate, most likely.
I left out General Dynamics, only because I didn't thought of them while writing.
I don't think that's right. IBM's main (no pun intended) mainframe OS is zOS, which is derived from OS/390. Linux can run on those mainframes (maybe on top of zOS, I'm not sure), but Linux is not the primary OS for IBM mainframes. And I don't think that IBM is letting zOS stand still while all the effort goes into Linux.
Most worthwhile OS research these days comes from universities. RPC in monolithic kernels is some 40 years old, I recall the old TRIX kernel at MIT that the GNU project was supposed to adopt way back, being based on that.
IBM is actually one of the more progressive companies, the GP is correct. Undoubtedly much more so than Google and Apple, and dare I say Microsoft (outside Research).
Well, WinRT (Ext-VOS), MIDL compiler for WP 8 (from Singularity toolchain), .NET Native, driver validation framework for DDK all started there.
WinRT is an application framework out of many historical ones, some impressive and others less so, you're well aware of that.
By driver validation, you mean the Driver Verifier tool, kernel-level signed drivers, or...? Those seem incremental, if useful.
.NET Native is just another entry in the currently popular line of optimizing language runtimes and VMs, I don't see why it's singled out as particularly impressive.
My point is you're overinflating things outside Unixland (and even inside, in the case of Apple - so it's more like free-Unixland) and selectively ignoring the lower profile work going on in free Unixes, even though in general I'd like to see an evolution beyond its paradigms myself.
MDIL is the Machine Dependent Intermediate Language, which was the format produced by the Sing# compiler in Singularity, given as input to Bartok, the native code compiler for Singularity.
The work was reused for Windows Phone 8 to generate .NET applications compiled ahead of time in the Windows Store, using just a dynamic linker at installation time on the devices
https://channel9.msdn.com/Shows/Going+Deep/Mani-Ramaswamy-an...
WinRT is the second coming of Ext-VOS, the initial idea that triggered .NET by making an OS ABI based on an improved COM. Then .NET happened and the idea kind of died.
With the going native wave gaining speed back at Microsoft, the project was reborn with some improvements like C++/CX and using .NET Metadata instead of the old COM type libraries.
Since Vista, the majority of new APIs are COM based and with WinRT, Windows is one of the few current OS with an OO ABI.
One of the driver validation tools makes use of the Z3 theorem prover as its verification engine
https://github.com/z3prover/z3/wiki
http://research.microsoft.com/en-us/um/redmond/projects/z3/s...
https://msdn.microsoft.com/en-us/library/windows/hardware/ff...
.NET Native might be just yet another compiler, however it is the continuation of the work that started in Singularity with Sing# and Bartok, followed by MDIL in WP8.
Also shows the direction of having .NET Native and C++ as the backbones of the OS, while moving away from C.
I am ignoring the work in free UNIX, because from the outside they appear to keep on copying commercial UNIXes.
Everyone is so high on Docker, yet I was using HP-UX containers in 2000 aka Virtualvault .
I concede these are all noteworthy tooling improvements, but none radical, nor all relevant to the subject of microkernels. C++ as backbone is quite clearly an incremental evolution and not even much safer.
Everyone is so high on Docker, yet I was using HP-UX containers in 2000 aka Virtualvault.
Basing your assumptions on Unix work over Docker woo is like basing Microsoft assumptions over Slashdot comments.
I admit it was a kind of low jab.
So what cool stuff is being done in mainline GNU/Linux that isn't present in Aix and friends?
I would suggest the ability to run on almost anything, from an embedded system or a phone, clear up to a mainframe. Aix et al never dreamed of that kind of scale.
Linux can also scale up to at least 4096 cores. Again, I don't think Aix and friends have any kind of precedent for that.
Of course, I could be wrong. And if most of those Unix flavors hadn't died untimely, maybe they could have gotten there. But I think that Linux broke ground that the other Unixes didn't, in at least those two areas.
z/OS has its own Unix Subsystem (Unix System Services, formerly known as MVS Open Edition or something like that). Linux can run either in an LPAR or on top of VM inside a virtual machine - I think I read somewhere that the latter is far more common, because people that run Linux on zSeries tend to use it for replacing lots of smaller servers rather than a few big ones.
But Apple, nowadays? They still ship that 1980's badly done micro-kernel research experiment Mach together with a whole BSD kernel in the same address space, same thing NeXT has done since 1988. All the innovation since then is in the high-level UX, not the base OS. And Apple nowadays does have the large cash pile that would allow them to waste some money on this kind of research.
Google too, their innovative stuff (GoogleFS, BigTable, MapReduce, cloud tooling) is running on top of a UNIX-like base OS (actually Linux, but it could have been any UNIX).
- Swift
- XPC Services / Sandbox
- Grand Central Dispatch
- IO Kit (C++ drivers)
- System Integrity Protection
- Metal
- Replacing C by Objective-C on the low level APIs
No part of the Darwin base OS is implemented in Swift, so it can hardly be considered relevant to OS research.
> - XPC Services / Sandbox
Kinda neat, but OS level isolation mechanisms have been a staple of UNIXes for decades: FreeBSD jails, Linux VServer, Linux seccomp, Solaris Zones, Linux grsecurity, SELinux, NT's security stuff (Windows Integrity Mechanism?)...
> - Grand Central Dispatch
It's a thread pool - does it have a system-wide idea of how many jobs there are? Neat. But I'm guessing it's just a user-space library + daemon?
> - IO Kit (C++ drivers)
Being able to implement drivers in an unsafe system programming language that was designed for such use cases like C++ is not particularly remarkable.
> - System Integrity Protection
From the Wikipedia article I get the impression that this is firstly about refusing to load unsigned kernel drivers, which NT has been doing since Vista (and even some Linux distributions now on UEFI secure boot enabled systems), and secondly about restricting root user with MAC based security mechanisms, similar to Trusted Solaris, Linux grsecurity, or SELinux.
> - Metal
AMD did it first, with Mantle. Now everybody is doing it.
> - Replacing C by Objective-C on the low level APIs
Guess nobody else has done this one. But wasn't it already the case in NeXTSTEP?
>> - Swift
> No part of the Darwin base OS is implemented in Swift, so it can hardly be considered relevant to OS research.
It is the only UNIX based OS moving away from C as their main language. All presentations at WWDC, except the new Objective-C updates, were about Swift.
Maybe in some future version, Swift gets to be allowed at lower levels.
>> - XPC Services / Sandbox
>Kinda neat, but OS level isolation mechanisms have been a staple of UNIXes for decades: FreeBSD jails, Linux VServer, Linux seccomp, Solaris Zones, Linux grsecurity, SELinux, NT's security stuff (Windows Integrity Mechanism?)...
You failed to understand that developing applications with XPC Services is similar to applications in micro-kernel based OS.
Also it is a container model imposed upon developers that wish to target Apple systems.
Whereas the other UNIXes container models are optional.
>> - Grand Central Dispatch
>It's a thread pool - does it have a system-wide idea of how many jobs there are? Neat. But I'm guessing it's just a user-space library + daemon?
No, it is implemented at kernel level.
It was good enough that FreeBSD adopted it.
Solaris is probably the only other UNIX that used to offer something similar.
>> - IO Kit (C++ drivers)
>Being able to implement drivers in an unsafe system programming language that was designed for such use cases like C++ is not particularly remarkable.
C++ is way safer than C. Those that code C with a C++ compiler should keep using a C compiler.
What other UNIX has moved away from C and required C++, none. Hence why it is remarkable.
>> - System Integrity Protection
>From the Wikipedia article I get the impression that this is firstly about refusing to load unsigned kernel drivers, which NT has been doing since Vista (and even some Linux distributions now on UEFI secure boot enabled systems), and secondly about restricting root user with MAC based security mechanisms, similar to Trusted Solaris, Linux grsecurity, or SELinux.
With the difference this is applied to a normal consumer OS. Again research is not only kernel code.
>> - Metal
>AMD did it first, with Mantle. Now everybody is doing it.
AMD doesn't ship OSes.
>> - Replacing C by Objective-C on the low level APIs
>Guess nobody else has done this one. But wasn't it already the case in NeXTSTEP?
No, it was all pretty much plain C at the lowest level.
Now some of those C APIs have actually OO Objective-C APIs.
On the one hand, yes, that would be nice.
On the other hand, I have been using GNU/Linux on my desktops for about fifteen years, and I have only seen a handful of kernel panics during that time. And most of those I strongly suspect being caused by faulty hardware. I remember the X server freezing on a few occasions, but that can indeed be restarted without a reboot, and the last time I saw that happen was years ago, back in the bad old days when people had to edit xorg.conf...
CPU cycles are cheap this days, yet cache misses are even more expensive (need to touch main memory? please wait 300 cycles).