Real hardware breakthroughs, and focusing on rustc
redox-os.org
redox-os.org
Honestly, I am always excited about systems like Redox that are going for the integrated and consistent user experience approach into building OSes. The microkernel approach on modern hardware sounds interesting here given that they skipped 32 bit entirely. The impression of most developers here keep referencing the Tanenbaum-Torvalds debate of 1992, as a justification for adopting monolithic kernels, which made sense at the time.
Fast forward into 2019, given that the CPU-level vulnerabilities are now alight, the need for a microkernel OS written in Rust could not have been greater. The reasons against a microkernel was commonly associated with the IPC and 'context-switch' performance impacts, but I find that the security implications here can help with countering these CPU-level vulnerabilities and the performance-concerns are in-directly solved by hardware advancements/implementations given for free or optimising the OS to run on multi-core systems.
Systems like Fuchsia have done both and I hope Redox follows and does this too.
In the case of the Zircon microkernel in Fuchsia, unlike Linux, almost all system-calls are asynchronous/non-blocking with a small number that are blocking [0], which is interesting since the OS is also fundamentally optimised for multi-core systems. Combined with both features, it makes it very suitable for real-time based applications, something that Linux is not optimised for and requires fundamental tweaking and changes for Linux or even some other traditional microkernels to achieve.
Thus, I doubt that on a system like Fuchsia/Zircon, the context switch would be notably 'very, very slow', even if Fuchsia was running on modern-hardware optimised for microkernels such as Zircon if that were to happen.
[0] https://fuchsia.googlesource.com/fuchsia/+/refs/heads/master...
They can probably do some smart stuff with having each syscall just add stuff to a work queue so userspace can resume as quickly as possible, but it's still fundamentally more context switches than e.g. Linux would have, which if side-channel attacks keep coming out might be a problem for them.
If mitigation is starting to amount to ~20%, that's roughly a CPU and soon to be 2-3, that can be told to never process directly for Userland and never tell Userland when exactly output queue contents of sensitive tasks become available. All that work essentially being free compared to running with Intel's fixes and full CPU flexibility.
They still require a context switch to access the protected memory in order to enqueue the async request.
Can you expand on your reasoning?
You might make submission lazy and ond only actually submit all queued operations on the poll call, but it is not much better.
The part that the Intel fixes makes slow is the mode switch, since you need to flush caches and trash data that may leak. Making it async doubles (at minimum) the number of mode switches needed to do work.
Additionally there are quite a few production quality microkernels, that people keep forgetting, because most only look at desktop OSes.
I don't get it. If these vulnerabilities involved leaking data across process boundaries, and the fixes make processes slow, how do more processes help?
Rust is supposed to be safer. There are still plenty of type system unsoundness issues or LLVM unsafety issues that leak through every so often.
Hardware is simply Yet Another Language.
i’m curious why?
i understand if a “safe language” provided unsafe capabilities (like marking a section with unsafe keyword to access registers or raw memory) then of course that’s a problem... but if you don’t have global variables and memory access, shouldn’t it be possible to cover it with just software?
maybe i’m missing something? (and apologies if the answer is obvious ^^)
Can you index out of bounds and catch the exception? You're attackable via Spectre. Can you run multiple threads and time operations? You're vulnerable to MDS.
The attacks don't rely on memory unsafety -- they rely on the state of the CPU being changed by doing fairly normal things, while an observer probes the state by doing other fairly normal things.
> Can you index out of bounds and catch the exception? You're attackable via Spectre.
ah, i see... ok assuming you could never access out of bounds (all arrays must be typed by length) would that still apply? or do you mean any kind of exception handling that can be deterministically invoked? what if a hypothetical safe language had no exceptions?
> Can you run multiple threads and time operations? You're vulnerable to MDS.
would that more or less depend on wether or not hyper-threading is enabled or not? assuming you can choose which cpu you are using(that is not vulnerable like intel) would that negate needing hardware memory protection when used with a “safe” language?
> The attacks don't rely on memory unsafety -- they rely on the state of the CPU being changed by doing fairly normal things, while an observer probes the state by doing other fairly normal things.
yes, thanks. i was just wondering what the limits to that re in terms of “safe’ languages vis-a-vis hardware memory protection/os etc...
Cache coherency is optimistic, message passing is pessimistic.
Cache coherence protocol are all designed to keep everybody in sync. So if you scale to thousands of cores, you better sync selectively and application specific. Thus it makes sense to use message passing (actor model).
It is unclear where the it makes sense to switch. Somewhere between 2 and 1000.
Hybrids are also an option. Supercomputers often use OpenMP (shared memory) and MPI (message passing) in a two level way.
Not some, almost all supercomputer nowadays use an hybrid model.
For three main reasons:
- Memory is a precious resource and separated process space + messaging is more memory hungry than a shared hybrid process space
- Hybrid machines with CPU + GPU on every node are everywhere and they requires a shared memory model.
- Debugging a pure MPI model at scale is a huge pain in the ass that take some of the most complex debugger in the planet.
They don't?
In practice cache coherency is a correctness mechanism; actually using shared memory concurrently R/W is problematic, performance wise. That's how we get scatter/gather-style systems, specifically to avoid concurrent R/W.
Wait, what? I'm looking at an image of an unlidded Zen 2 CPU right now and it's a bunch of discrete compute devices surrounding a switch. Inside these devices are multilevel caches, just like real world network systems. The central "chiplet" multiplexes access to external buses over multiple serial PCI lanes; any network engineer would would recognize this pattern as trunking. All of this looks exactly like a network.
Not sure if anything happened there after the FOSDEM talk.
https://archive.fosdem.org/2019/schedule/event/hardware_soft...
https://archive.fosdem.org/2019/schedule/event/hardware_soft...
There may be a way out of the loop exploiting hardware parallelism, or if they bring some improvements for memory locality leading to orders of magnitude improvements. But even exploiting memory locality, that would bring some very high and obvious gains is locked at the same kind of loop.
On the performance question, I have a new 64 bit system coming with 64 threads[2], and I'm really curious to experiment with non-relocating microkernel architectures (architectures where kernel features have a fixed address that doesn't change while running). This is something I dreamed about at Sun using a Dragon (64 thread SPARC machine) and Spring (the research OS) but it was too expensive to allocate one of those machines to our small group for OS research.
[1] Yes, there are OSes written (mostly) in Java, but at Sun we tried a completely Java OS and the safety features of the language prevented it. The 'controlled relaxation' of safety features in Rust here (as opposed to 'jump into native code, all bets are off!') are awesome.
[2] In theory I'm in the queue for the first arrival of an AMD TR-3990 for my sTRX motherboard. We'll see how that holds up. :-)
So if you can write a multi-user OS that only used the JVM (so it was 'Java all the way down') then you could build a Java chip for embedded systems, use a JVM for running it on a commodity processor, etc.
I had a lot of fun building the PiDP11/70 kit last year which is a PDP 11/70 built on top of simh on a Raspberry Pi. You can run BSD 2.1 on it which is decidedly Old School, but it (and RSX-11M+) are classical multi-user OSes. And that system has all the pieces of what we were trying to do with Java which is a virtual machine, compilers that converted code to run on that virtual machine, and an Operating system that can be compiled to run on it.
One of the standard tests of compilers is generally that they bootstrap, compile themselves with bootstrap, compile themselves with themselves, compile themselves with themselves again, and finally compare the binaries. Any differences at that point are considered to be a bug.
Dynamic linking seems to be at odds with this.