The Helios microkernel
drewdevault.com
drewdevault.com
> Again, much of the design comes from seL4, but unlike seL4, we intend to build upon this kernel and develop a userspace as well.
I'm wondering why they didn't decide to build the rest of the system on top of that.
I'd totally respect answers like "because building kernels is fun" or "because I wanted to do that in my programming language". I'm just wondering if there was another reason seL4 was unsuitable.
In short, seL4 is cool, but not very practical, and I think it can be done better.
To be clear, I am specific about the "seL4" inspiration rather than the "L4". Naturally L4 plays a role, but Helios's API design has a pretty obvious influence from seL4 specifically.
- the seL4 team definitely works on user space stuff (see the rest of the repos in the seL4 organisation on GitHub), but it's very basic thus far. there is some design work going on to standardise some user space bits and pieces. but i don't think they're ever going to build an top-to-bottom opinionated operating system stack – that's for other people to do.
- it's not a solely academic exercise, it is being used in industry, but most of what everyone sees is the output of the research group, which gives it an academic sheen.
- the unfamiliar internals and implementation are driven by a fundamental difference between seL4 and almost all other kernels... it is verified.
Yeah, e.g. General Dynamics bought OK Labs. *waggles eyebrows*
All of these people are frankly better engineers than me, of course, so my opinion doesn't matter here, really. But I do have a worry about people in our industry seeking novelty instead of building on the proven.
(Except of course for research or hobby projects. Go for it.)
SUN, DEC, SGI and many others besides that aimed to dominate the server market were left by the wayside, and even Microsoft was running (very) scared. If not for their illegal behavior they too would have lost the server market.
I did a summer contract job at an IBM subsidiary in summer 98 or so, and I had a manager there fight me on installing Linux on some spare PCs, because he thought they should be running, y'know, SCO. A Real Unix. It already looked preposterous by then.
That is pretty much my own experience with sel4. Not very useful for real world
I was considering seL4 for a hobby project (having worked with L4Ka:Pistachio a long time ago), but didn't dig deep enough to spot these problems.
As soon as you start adding stuff or running it in a different hardware the verification will no longer apply, so why bother?
> Please don't use Hacker News for political or ideological battle. It tramples curiosity.
"What, this comment doesn't do either of those?". Taken in the context of your other comments in this section, it does.
https://www.theregister.com/2021/12/06/heliosng/
Very interesting OS that is FOSS now. By v3 it ran on multiple CPU architectures. With some of the manycore CPUs appearing in recent years, this is crying out for a revival.
Drew?
It does not look like the other Helios project is going anywhere, so I'm not too concerned with the naming conflict right now.
There are a couple of books about it from Prentice-Hall: ISBN 0-13-381237-5 and ISBN 0-13-386004-3, plus various academic papers, commercial software for the OS (mostly development tools), open source software (X11, gcc, other stuff), etc, etc.
It was the default operating system of the Atari Abaq/ATW (https://en.wikipedia.org/wiki/Atari_Transputer_Workstation), and also ran on various other Transputer, i860, and various other early parallel computer systems.
It too was a micro-kernel, with a very Plan9-like global name system implemented by applications plugging into the name server protocol.
If nothing else, it might be of interest to read up on it, but given the similarities between the projects, it seems to be an unfortunate naming collision.
Don't forget the original (for Linux anyway), SEL4.
Also, seL4 comes from a long line of L4 microkernels (https://en.wikipedia.org/wiki/L4_microkernel_family#History) which actually predates Linux (L3 was developed back in '88).
If you really wanted, could build a minimal kernel and move many devices to be inserted at runtime.
Anyway, wouldn’t mind if all that was made a lot smarter and made more accessible to people.
- Microkernels need more development effort for equal functionality, something that was a constraint when he was doing Linux on his own.
- The context switching penalties on 386s were ginormous - it's much more reasonable now, I'd even say it's not much more penalizing than thread switching.
- We live in a multi-master world with IOMMU - tons of devices have virtual address spaces on their own, with most of them doing DMA transfers.
OS X and Windows always had a hybrid architecture since the beginning, and with the push for user space drivers for third parties even more so.
Most commercial RTOS are microkernel based, like QNX and INTEGRITY.
Even Android imposes a fake microkernel like architecture since Project Treble, where since Android 8 classical Linux drivers are considered "legacy".
Then there is the whole irony that those that strongly argue for monolithic kennels, nowadays run Linux on top of a type 1 hypervisor, full of containers.
"Hybrid architecture" is a PR exercise. These are monolithic designs with a better marketing team.
Imagine if somebody tells you their new programming language has a "hybrid" numeric type. It has all the advantages of floating point numbers, yet none of the disadvantages of the integer types. Wow!
Wait, hang on, let's look a bit more closely, those are floats. They've just described floats, but used some semantic sleight of hand to claim they're a "hybrid" numeric type.
Others enjoy that there is some effort in place to have RPC comunnication between kernel modules instead of direct function calls, a hard push for third party drivers only to exist in userspace and a standard ABI for extensions.
I rather see the progress side, even if partial, than hold to legacy designs.
See I could also call a PR exercise to safe languages that depend on using unsafe everywhere on their lower layers, maybe they aren't safe after all.
And yet we know that isn't the case.
Although it's easier to look at what's technically different in a language like Rust, probably more important is a philosophical difference and the resulting community.
Let's use an example I gave recently to my line manager (who is technical but not an excellent programmer): Sort stability. C++ says, here's sort function, and in the documentation by the way this is unstable, if you need a stable sort that's named stable_sort(); Rust says, here's a sort function, if you know enough to know you don't need a stable sort, there's an unstable_sort() function which is faster [and exists in tiny embedded systems since it doesn't need an allocator]. Unstable sort isn't unsafe (well, it is in C++ but that's C++ for you, it's not a fact about unstable sorting) but philosophically it violates the principle of least surprise as your default sort() function, and that would be a foot gun.
Go has an "unsafe" library with the dangerous stuff in it. But unlike Rust it doesn't have the safe philosophy and so e.g. as we saw recently Go doesn't feel the need to even warn you that things aren't thread safe and some Go proponents think you're crazy if you expected anything to exhibit thread safety.
If I saw a clear RPC philosophy in an OS like Windows I'd mostly go along with what you're saying. If when I looked at Windows I saw a bunch of loosely coupled systems communicating using a well defined protocol, I'm on board that's really not just a monolithic kernel. But what I see instead is mostly layers smeared on top of a monolithic kernel, some of which pretend to be RPC but are really just doing a syscall to some kernel code where the work happens. The language of RPC is used in some places, but the philosophy is largely absent.
Or having been a former Microsoft or Apple employee working on kernel stuff, then I could eventually value the opinion how fake they happen to be.
Yet somehow I get the feeling, from both of us, only I have delved into such internals.
C++ indeed has std::ranges::sort() and std::ranges::stable_sort()
However Rust's [T].sort() is paralleled by [T].sort_unstable()
That is, I just got the Rust names in the wrong order. The orthography is different because how these work under the hood is very different, but that's not important here.
But this is OK, as long as Unsafe is a monad!
This is similar to what I heard from other people about NT. IIUC, NT was called hybrid because it allowed something similar to loadable modules.
https://en.m.wikipedia.org/wiki/Local_Inter-Process_Communic...
Additionally, since Window 10, all critical modules are sandboxed, and running under hypervisor control when hardware support is present.
https://www.microsoft.com/security/blog/2020/07/08/introduci...
This is something that Windows 11 requires, secure kernel hardware requirements aren't optional.
Calling it a marketing exercise only reveals a lack of understanding about the subject, and possibly some bias as well.
Definitely my case.
So far, you can write drivers for USB devices (https://libusb.info/), simple I/O devices (https://www.kernel.org/doc/html/v5.14/driver-api/uio-howto.h...), and filesystems (FUSE) in userspace. There are probably more, but those are the ones I know about off the top of my head.
The performance really depends on your application. There’s a lot in the kernel network stack (eg iptables/nftables) that not every application needs and that might be costly. But then again for usual TCP based applications with larger dataflows the kernel path is also already decently optimized with hardware offloads - and running userspace networking stacks comes with its own issues.
If networking functions like sendmsg make up > 30% of an applications profile it can be worthwhile to investigate it.
[0] https://drewdevault.com/2020/08/13/Web-browsers-need-to-stop...
[1] https://drewdevault.com/2020/03/18/Reckless-limitless-scope....
But hey having a browser and ext2 should be good enough ;)
EDIT: Well, ignore, I didn’t realize we are discussing from scratch browsers :)
Must feel nice to have an article written about you!
Given your karma, I assume you already know this, but from the guidelines:
> Be kind. Don't be snarky. Have curious conversation; don't cross-examine. Please don't fulminate. Please don't sneer, including at the rest of the community.
9 months ago:
It's probably an additional huge workload for the production of the video though...
Is he rehearsing his coding videos? Does he pick moments that went well? Is he a superhuman coder?
I wish it could be a successful OS and maybe some day we can replace Linux for it, since now Linux is controlled by big corps.
I know Drew is a fan of Plan 9, so that may not be out of the picture - though it looks like there's already a non-negligible amount of stuff planned that may or may not be gotten to.
So, not. Amusement value only. (Amusement suffices, sometimes.)
A new language is a full-time project, if you are at all serious. A new kernel is another full-time project, if you are at all serious. A new kernel coded in an absolutely immature language starts it out at a major handicap, even if you were serious.
Nobody is obliged to be serious, of course, on a hobby project, but there is no hiding the fact. Wishing doesn't make up the difference. Taking the question at face value, it deserved an honest answer, even at risk of resentment in, well, the peanut gallery -- which delivered.
You are the peanut gallery, by the way.
> Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something.
Waited until this thread is hopefully quieter to post this.
Inspiring because he's clearly chasing his passion.
Confusing because it comes across as a lack of focus.
Did you start with an assumption of UEFI, long mode only?
The comment in question is here: https://news.ycombinator.com/item?id=31764243
However I just want to add a personal comment. In my view, programs will come and go. They can be written and rewritten. But what is really important if we want to ensure free computing for all is not gradually eroded is to work on as many open standards for interoperability as possible to serve as alternatives for proprietary protocols. Also open, community developed languages. However the hardware situation is out of the hands of software developers, yet critically important. Open software and protocols on closed hardware is unsustainable. Lets hope someone like Musk takes the initiative in this area.
I don't think that's a good idea. There's the xkcd 927 concern, since there are already long-established standards for a lot of things; and that brings me to what I think we really need, and that is stable standards. The longer a standard has been around unchanged, the more implementations will arise. On the other hand, the browser situation is a great example of what happens when some entity takes control of an "open" standard and then churns it endlessly. I think this is particularly true of programming languages, since they are foundational for everything else.
A number of standards is still patent-encumbered. Back in the day the algorithm to encode a GIF was patented, do you remember that? Do you remember how long Oracle and Google spent in court debating whether it's acceptable to describe and implement public interfaces?
A world where you can but may not interoperate is very inhospitable, no matter how stable the closed standards are.
What?
$ git clone https://git.sr.ht/~sircmpwn/helios
Cloning into 'helios'...
remote: Enumerating objects: 1820, done.
remote: Counting objects: 100% (493/493), done.
remote: Compressing objects: 100% (444/444), done.
remote: Total 1820 (delta 242), reused 0 (delta 0), pack-reused 1327
Receiving objects: 100% (1820/1820), 415.45 KiB | 455.00 KiB/s, done.
Resolving deltas: 100% (867/867), done.
$ cd helios
$ git shortlog -s -n --all --no-merges
174 Drew DeVault
30 Eyal Sawady
1 Sebastian
Drew is pretty good at getting folks involved.Fuschia is an extraordinarily complex kernel design. Hell, it's a "microkernel" which is bigger than Linux! It's fairly typical of Google's inwardly-focused engineering culture, using questionable tools (with large shadows) such as Bazel and gRPC, which were likely chosen simply because they plug into the Googler engineering culture. At the same time, these decisions give way to a very complex system with heaps of moving parts, in this rube goldberg design which is endemic of Google engineering.
Helios is much, much simpler. The kernel itself will probably clock in at under, say, 20,000 lines of code (for x86_64, at least), and it has a very small syscall API, less than two dozen in the final design. Many of the things Fuschia does in the kernel will be done in userspace on Helios, in the Mercury component, such as service discovery and capability allocation. Finally, Helios is written in Hare, which itself is a much simpler language than C++, and kernel hackers at the very least should be convinced of the argument that the complexity of your implementation language contributes to the complexity of your implementation.
Bazel and gRPC, whilst both saddled with problems stemming from Google's inwardly-focused engineering culture (and also just some sub-optimal historical decisions) are not "questionable tools" in the sense that they solve the wrong problems or something else solves the same problems much better.
Blaze/Bazel (and its various clones like Buck) are basically the only general purpose build systems out there that even attempt to do the basic tasks of a build system, namely actually figuring out what has changed and needs to be rebuilt. So of course they are gonna use it, nothing else comes even close (and the main downsides don't apply to them).
Similarly, gRPC has a lot of warts in both how its encoding, type system, API and transport work. But anything in that space that doesn't completely suck (such as cap'n proto) is basically a clone of the core design. Again what else even attempts to solve the core problem of having some backwards/forwards compatibly reasonably efficient rpc, messaging and data storage?
Extraordinary claim. You will need to explain how Bazel does that and, say, CMake + Ninja don't.
Perhaps the OP means full incremental compilation, which requires "cooperation" from the compilers, really (or the build tool actually parsing the language's AST like Gradle does, I believe). Or in the case of Bazel, the build author explicitly explaining to the build what the fine-grained dependencies are (I don't use Bazel so I may be wrong, happy to be corrected).
Sure you do. Any build system which has a "clean" command that you'd use for anything other than freeing up space.
If you think this claim is extraordinary and CMake is a counterexample, I don't know what to say. Have you really not have had to manually fuzz around with stale builds (by e.g. issuing "clean" commands, removing build/ directories etc) when using CMake or had to figure out why something worked on your machine but not a colleagues?
> You will need to explain how Bazel does that and, say, CMake + Ninja don't.
CMake makes no serious effort at all at "figuring out what has changed". Even trivial stuff doesn't work. If you have a wildcard pattern in a CMake file and you add a file that matches the wildcard, you will generally get a stale result (which is why it's "not recommended"). So not even explicitly specified inputs work correctly and CMake makes no real attempt to to detect unspecified implicit dependencies. Does your CMake build setup correctly detect when you updated gcc or some system library? That some build artifact implicitly depends on another?
By contrast Bazel goes through some fair amount of effort not only to detect changes to specified dependencies correctly but also to sandbox build steps to check that they don't depend on stuff that has not been explicitly specified. As the blaze docs say:
> When given the same input source code and product configuration, a hermetic build system always returns the same output by isolating the build from changes to the host system.
You can get pretty close with bazel; good luck with CMake.
System libraries are intentionally disregarded as possibly changed dependencies, that is a design decision that most build systems make. System libraries are not supposed to change in binary incompatible ways without a major version upgrade.
Your last point is about reproducible builds, which is a different topic.
I disagree, and there is a huge philosophical gulf of understanding between my view and those who feel differently.
Let me ask you two simple questions:
1. I work on a project in a git repo and do a git pull to get the latest changes to the branch. I do the equivalent of `make` with my build system, and encounter a problem (either a build error, or some unexpected bad behavior from the built artifact). Should I be allowed to conclude that someone has messed up, or do I first have to engage in some gyrations to make sure I have a "clean` build and the correct dependencies (git clean -dxf, issue commands to manually check and update dependencies etc.)?
2. I push some code after running the equivalent of `make test`, which passes successfully. A collaborator informs me that after pulling their build is broken or there is some unexpected failure, which does not appear to be a flake. Should I just expect this to happen every now and then and live with it?
If your answer to 1. is "No" and your answer to 2. is "Yes" than there may indeed be a gulf of understanding, but probably not a "philosophical" one. If your answers are "Yes" and "No" respectively, then I'd certainly love to hear what philosophical more aligned general purpose build tools have this property, because I'd be interested in investigating them. The only other tools in this space that I'm currently aware of of even making an attempt in this direction are either different granularity (nix) or unreleased/more of an academic POC (redo/shake).
1. Caching, while maintaining hermeticity and correct dependency tracking.
2. Remote execution - yes, contrary to the author's general focus on small, specialized, C (or similar) tools, products out there have to deal with thousands of dependencies that take forever to build and are nice to offload to a beefy server. They also need to support several different devices and architectures.
3. Actually being able to run tests, use the same caching mechanism on tests, surface those results to CI.
4. Being able to plug in arbitrary languages into the build system, with a sane (subset of Python) DSL to write rules in. Because not every project is written in C/C++, or in a single language.
5. Having very clear separation of build phases so you can't shoot yourself in the foot.
6. A lot of hooks to provide better integration with CI, as well as to collect profiling info from your end users, so you can actually make life better for your engineering org as a member of the developer tools/infrastructure team.
In general, my gripe with Bazel complainers is there is a loud community of them on the Internet whose primary software engineering experience is either in languages that build fast, projects that are small and with few dependencies, or just have a small number of people.
There is a world of software out there beyond the Web and C UNIX utilities. Try doing any high end computer vision, robotics or HPC work and you've to deal with a bunch of dependencies trying to solve very complicated problems. One can discuss whether some of those dependencies are well designed or not, but at the end of the day they are very good at their job and don't have much competition.
Bazel solves a lot of problems for a lot of these people, so calling it "questionable" is quite ignorant.
The closest alternative is Thrift. It's not as popular, alas, but its design is a lot better: you get to combine various encodings (where protobuf mostly has its binary encoding, json, and the underspecified text for configuration); and you get to pick a transport instead of being forced to use http2 with questionable features. You could run Thrift RPC over tcp with basic framing, or over 0mq, or http, or anything else. I'd also argue that Thrift's data representation is better than Protobuf (nesting) but I'll write about it later.
So protobuf is yet another "worse is better" from google: a ton of warts, choices that might make sense for google itself but not for a lot of other situations (http2!), and still winning over superior alternatives :/
Oberon was going for the 'Make it as simple as possible, but not simpler' mantra. Oberon is about 10.000 lines of code for the complete OS, including its UI, compiler etc.
I love the architecture of Oberon, but today it would be more fitting to use within an E-ink e-reader device.
But Helios could eventually evolve into an OS that runs a mail server or web server efficiently. Or even run an existing UI setup like Wayland/Sway.
"Mainstream": has more than 5000 users on its best day ever.