Fomos: Experimental OS, built with Rust
github.com
github.com
> The argument that a cooperative scheduling is doomed to fail is overblown. Apps are already very much cooperative. For proof, run a version of that on your nice preemptive system: [...] Might fill your swap and crash unsaved work on other apps.
The difference is that in a preemptive system, a `while (true)` might slow down the system, but in a cooperative system, the machine effectively halts. Like for good. If you're caught in a loop and don't yield control back, you're done.
In terms of security, this would make denial of service attacks against such systems trivial. You could exploit any bug in any application and it would bubble up to the entire system.
Or maybe I'm all wrong. I'm not an OS dev, so please someone correct me.
For very specialized (probably embedded, but maybe other) use cases, I can see the value of specialized operating system which dispenses with time slice scheduling within a core and just assigns waiting tasks to a free core when/if it becomes available. On systems with very high core counts, and with lots of short lived tasks, there could be some value to this.
Even back then, on many machines and operating systems it's not like the whole world stopped when a process would not give up time slice. Things like mouse cursors, keyboard buffering, and often even some I/O were able to proceed because the hardware assisted in making that possible.
The thing is, I don't think going cooperative simplifies much.. you still have to handle concurrent access to shared resources, and re-entrancy etc. By the time you've made your operating system able to handle that (memory, VM subsystem, I/O, etc.) from multiple cores... you may as well just go implement timeslicing as well.
The opposite could work: OS is cooperative, unless some threshold of resource usage is triggered (a timer interrupt of instance). It then context switch to enter a, hopefully rare, failure mode, thus turning preemptive. Kill the app, and get back into cooperative mode. Let's call it optimistically cooperative & pessimistically preemptive.
That's literally the definition of preemptive scheduling.
Also: this is more like a multi-threaded app itself than an actual OS, for it to be a proper operating system you'd expect at a minimum memory barriers between applications. In that sense the bar for what an OS is has been raised quite a bit since the times of CP/M where it was more of a 'handy library to do some I/O for me' (on mainframes and mini computers there were already proper operating systems back then but on micro computers we had to wait until OS/9 to have something similar).
What was it called? Do you have more details?
In the end though, all this really does is perhaps change how aggressive a pre-emptive system needs to be. It could wait longer to pre-empt. Fundamentally though, the design would still need to be pre-emptive.
In practice, I'm not sure how relevant pre-empting is any more. Software is much better behaved than it used to and I rarely see these runaway processes anymore. And when they do it's usually pegging a single core at 100% while other processes share the remaining core.
For the programmer, it's most convenient when the isolation between unrelated programs is maximal but isolation between related programs is minimal; but for the user it's the most convenient when programs are maximally isolated in terms of computation, but not for other things such as storage/permissions etc. Practically, this requires a complex scheduling mechanism that is more preemptive than cooperative, but also cooperative a little. E.g. in linux, it's not particularly hard for a rogue program to bring the system to halt with something like a fork bomb (especially if swap is turned on). Even if you have a preemptive scheduler, if your program has 99% of the threads running in the system, it's effectively cooperative because chances are you're scheduled most of the time anyway.
Certainly not, cooperative multitasking is used all the time in userspace programs to great effect, and OSs can do the same thing if the limitations are understood.
For the OS that's the focus of this post though, I think it does have the aim of being general purpose and running whatever you want on it. At that point the limitations aren't really acceptable (which the author acknowledges), and that's why the whole question of "what happens when a program misbehaves?" has come up. If the answer is "preempt the process to run it later" then it's not actually a cooperative system and application developers need to keep that in mind.
Nope, you are completl right! A while (true) might slow down the system, but even that is not necessarily the case: I wrote a program to test this (based on the pseudocode in the readme) and my system is totally usable, with basically zero lag! This is the power of a well-written operating system that includes a mysterious concept called priorities. There is actually no discernable different when running the program (other than battery use and temperature going up). In fact, I am running that program as I right this because the difference is so negligable.
program: https://paste.sr.ht/~pitust/47dc80ff09243b4bf41a08ae9434a32c...
> In terms of security, this would make denial of service attacks against such systems trivial
Denial of service attacks where you have access to the target system are not super hard to mount AFAIK.
Not because it isn't a resource disaster. But because unless you measure something, you might not notice.
Oh! You did measure something. Battery use and temp. There you go. It's a disaster. Some kind of management of such irresponsible threads is definitely a good idea.
> Oh! You did measure something. Battery use and temp. So I didn't do any scientific measurements, but the temperature difference doesn't seem to be very significant.
> It's a disaster. But cooperative scheduling isn't any better. In fact, it's even worse! Since if an app doesn't yield (and let's face it, no app is perfect), you can easily get a deadlock. Developing for such a system without a VM sounds like a nightmare too to be honest.
> Some kind of management of such irresponsible threads is definitely a good idea. What kind of managment do you propose? I mean, if there is nothing else running on the system, it's probably okay to just let them keep running (it's not like they are harming anything except battery life, and you might want something that can max all cpu cores, like if you are running a compilation or a video game).
Particularly enjoyed:
In Fomos, an app is really just a function. There is nothing else ! This is a huge claim. An executable for a Unix or Windows OS is extremely complex compared to a freestanding function.
I can't even begin to imagine how cool a kernel would be written this way.Correct me if I'm wrong, but I think this is how smalltalk/squeak works?
I hope the author continues with this project. File system, task manager, safe memory stacks, nice resource sharing, ...
... and of course, running DOOM as a minimum proof of concept requirement! /s.
/* Create function pointer of appropriate type, initialised. */
int(*foo)(int, int) = NULL;
int main()
{
void * ha;
if ((ha = dlopen("A", RTLD_LAZY|RTLD_NODELETE)) != NULL) {
foo = dlsym(ha, "foo");
dlclose(ha);
}
/* if "foo" is non-NULL you can call it now. */
...But I don’t think that’s really that significant… you could, e.g., enumerate all the implicit things the OS provides, express them in a structure, and now you’ve got your explicit context. The “fomos” context is only as simple as it is because the OS provides a small number of simple things. Expand those to full OS capabilities and the context will get full OS complex.
The odd thing about this “OS” is that a running app is just the start function called repeatedly in a loop. This makes one app call a single coherent slice of app execution. That’s kind of interesting, but makes this pretty limited. E.g., it looks like apps are entirely cooperative. It’s maybe more of a cooperative execution environment.
Maybe it would be good for some embedded uses?
So you get the benefits of simplicity of a static binary and the size benefit of a dynamically linked binary.
Well that's the way I understood it anyway.
Your premise is wrong: in, say, Windows or GNU/Linux an application is, say, a PE/COFF or ELF executable, both non-trivial file formats. The int main(...) function in C is just a very leaky abstraction of all this.
There is no such thing as "program is just a function" unless it's compiled with the OS, like on some microcontrollers. For programs residing on storage, the OS needs a way to load them into memory along with its resources which means it needs to read some kind of format, even if it's trivial.
The "app is a function" distinction is simply about the `Context` struct passed it, otherwise it loads and runs executables as expected (load ELf memory segments in, jump to entrypoint)
Edit: I guess I'd also add that on Linux or other OSs it's possible to construct and run process completely from memory with no on-disk executable file. JIT compilers do effectively this, the produce the executable code completely in memory and run it.
This is rather likely true. But since writing such code is much harder than, say, writing FizzBuzz, I assume that somewhere in the OS code, a "wrong" abstraction is used which makes this task far too complicated.
I'm not really sure what you're looking for or expecting unless you want the OS to compile your code for you (which this OS doesn't do either).
In Smalltalk those would be classes' methods rather than freestanding functions, but yes: all objects within the image directly send messages to each other, i.e. invoke each other's methods. Lisp machine operating systems are somewhat closer, since initially they had no object system and used freestanding functions calling each other, though later they became generic functions specialized to their arguments' classes.
Virtual memory and vtables are now more about access control than about managing the scarcity of pointers.
I would imagine it could greatly speed up composable things like that...
To me this just screams the curse of greenfield development, where the architects are yet to discover what led other OSes to require all things they missed.
> As long as the OS stays compatible with the old stuff in the context, it can add new functionalities for other App by just appending to the context the new functions
Welcome to backwards compatibility hell. You're boxing yourself in this way because you're forbidding yourself to remove old / defunct entries from the Context struct.
I think a better approach would be to introduce some kind of semantic versioning between OS and app. The app declares, somehow, what version of the OS it was built against or depends on. The OS then can check if the app is compatible and can pass a version of the Context struct that fits. This doesn't get rid of most backwards compatibility problems, but you keep the Context struct clean by effectively having multiple structs in the kernel, one for each major / minor version.
struct Context{
padding: [u8;256], // old stuff
ctx: ContextV42
}
Just typing that felt a bit wrong though. On the other hand the app declaring what version it is feels like what executable formats (like elf) already solve. I am trying alternatives.This is a bit strange. I would think async I/O in the style of io_uring would be fantastic but this kind of model seems to rule out anything like that. That’ll make it hard to get reasonable perf. It’s also strange to not support async as it’s a natural suspension point to hook into but you would have to give up a lot of the design where your application state has to be explicitly saved / loaded via disk if I’m not mistaken. Seems pricy. Hopefully can be extended to support async properly.
In a similar vein, I suspect networking may become difficult to do (at least efficiently) for similar reasons but I’m not certain.
You give up the language-level support for coroutines and async, though...
The example is just too contrived. On a preemptive OS, apps typically hang in ways that don't turn the whole thing cooperative (thread deadlock, infinite loop, etc.). Also, a preemptive system could kill an app if it creates too many threads, files, or uses too much RAM, long before it gets effectively cooperative. Our systems are just more permissive.
> [Sandboxing] comes free once you accept the premises.
and yet
> any app can casually check the ram of another app ^^. This is going to be a hard problem to solve.
So no, sandboxing doesn't come for free.
That said, it's a cool idea and I wish the author success!
And even then, I think modern browsers still isolate each tab in a separate process just to be safe. I don’t think they share memory.
But the basic idea of using a managed language like Java or something to eliminate the need for hardware process security goes way back. Microsoft's Singularity project is I think the best developed effort at this.
Any system competently designed for robust sandboxing would have limits for all resources and reject requests when the limit is reached.
I'm curious what Fomos uses as a distinction between "process" and "executable."
On Linux a "process" is the virtual address space (containing the argv/envp pointers, stacks, heap, signal masks, file handle table, signal handlers, and executable memory) as with some in-kernel data (uid, gid, etc) that determine what resources it is using and what resources it is allowed to use.
An "executable" is a file that contains enough bits for a loader to populate that address space when the execve syscall is performed.
One of the distinctions is that you do not need an executable to make a process (eg, you can call clone3 or fork just fine and start mucking with your address space as a new process) and while the kernel uses ELF and much of the userspace uses the RTLD loader from GLIBC you don't need to use either of these things to make a process in a given executable format.
And finally, a statically linked executable without position independent code is "just a function" in the assembler sense, with just enough metadata to tell the kernel's loader that's what it is. But without ASLR to actually resolve symbols at runtime, it's vulnerable to a lot of buffer overflow attacks if the addresses of dependency functions are known (return to libc is one of those, but it's not unique).
I'm the first to point out the flaws in glibc and want an alternative to the Posix model of processes (particularly in the world where the distinction between processes, threads, and fibers is really fuzzy and that is clear even within Linux and Windows at the syscall level), but I'm curious what is going on in Fomos. Most of the complexity in "executables" in Unix is inherent (resolving symbols at runtime is hard, but also super useful, allowing arbitrary interpreters seems annoying, but is one of the strengths of Linux over Windows and MacOS, providing the kernel interface through stable syscalls is actually the super power of Linux and a dynamic context either through libc or a vtable to do the same thing is not that great, etc).
Dirt simple model works until it doesn't. Without address space separation there is no safety between executables and cooperative scheduling is the same. Running faulty binary or simply bit flip in ram can crash the whole system instead of just that process.
"How does loading work" has a simple answer: you map the executable into virtual memory and jump to the entry point. But the "why does XYZ format do this to achieve that" has a lot of nuance and design decisions - none of which are documented. Particularly things like RTLD, which is designed heavily around the design of glibc and ELF, while the designer of the next generation of AOT or JIT compiled languages for operating systems with capability based models for security might want to understand before they design the executable format and process model that may deviate from POSIX.
There's space for design and research there, and a platform that makes that easy has a lot of value. While I would encourage the designer of such a platform to read the literature and understand why certain things are done the way they are, it's valuable to question if those reasons are still valid and whether or not there's a better way.
Foundation. The Galactic Empire makes use of atomic energy and other technologies that were created in the distant past and the technicians can only (sometimes) repair but not create. 'The Last Question' has a question that remains unanswered throughout human history, which isn't quite the same thing.
I thought that it was drivers? Linux isn't particularly unique for having a stable abi (and the utility of such a decision is highly questionable). The driver support however is extraordinary and undeniable.
I think a slightly more fancy compiler could do something similar to "enforce" coop-multitasking: insert yield calls into the code where it deems necessary. While the halting problem is proven to be unsolvable for the general case, there still exists a class of programs where static analysis can prove that the program terminates (or yields). Only programs that can't be proven to yield/halt need to be treated in such fashion.
You can also just set up a watchdog timer to automatically interrupt a misbehaving program.
They forked C#/.Net, taking the async concept to its extreme and changed the exception/error model, among other things.
There are several other OS projects based on Rust, relying on the memory-safety of the language for memory-protection. Personally, I think the most interesting of those might be Theseus: <https://github.com/theseus-os/Theseus>
As long as all of them have to use an authorised llvm equivalent and that one can enforce memory access and cooperation that could look from the outside like a quite normal user experience with many programming languages available?
For performance though, I don't think halting analysis is needed. Even if the compiler can prove a piece of code terminates, it doesn't help if the inserted yield points occur too infrequently. If a piece of code is calculating the Fibonacci sequence using the naïve way, you do not want the compiler to prove it terminates, because it will terminate too slowly.
I am aware of two ways to enforce security for applications running on the same hardware:
(1) at runtime. All current platforms do this by isolating processes using virtual memory.
(2) at loadtime. The loader verifies that the code does not do arbitrary memory accesses. Usually enforced by only allowing bytecode with a limited instruction set (e.g., no pointer arithmetic) for a virtual machine (JVM, Smalltalk) instead of binaries containing arbitrary machine code.
The author of Fomos doesn't want context switching, memory isolation, etc. And Rust compilers don't produce bytecode. Is there another way?
Sounds like a context switch to me :) (at least for the MMU registers)
1. native Rust modules
2. arbitrary code running in WASM sandbox
3. arbitrary code running in KVM virtualization, perhaps with a Linux kernel there to provide a backwards-compat ABI
4. one can compile the WASM+untrusted app into trustworthy machine code (still implementing the WASM sandboxing logic for the untrusted component); I recall Firefox did this with some image handling C++ code they didn't find trustworthy enough
I know the Theseus project is working toward WASM support.
I’d also assume that a function that is misbehaved (doesn’t return) could be terminated if the system runs out of cores.
My point is that cooperative multitasking doesn’t necessarily equate to poor performance. Time sharing was originally a way to distribute a huge monolithic CPU among multiple users. Now that single user, multi core CPUs are ubiquitous, it’s past time that we think about other ways to use them.
I’m really excited that this project exists.
> cooperative multitasking doesn’t necessarily equate to poor performance.
I meant to say “poor interactive performance”
The lack of context switches in this model is likely to actually improve performance. So now I’m curious what would happen if we turned the Linux timeslice up to something stupid like 10s :)
I suppose that I’m idealising a bit here but ISTM that the structure of FOMOS means that the CPU state doesn’t need to be saved, so the context switch involves only memory protection, register resets, stack pointer reset and little else. You don’t even need to preserve the stack between invocations. And unlike preemptive multitasking, there seems to be little or no writing to memory needed, which would seem to obviate a bunch of contention. (Noting that it’s 30 years since I fiddled with operating systems at this level)
Frankly I find this really elegant and exciting.
Theseus is a safe-language OS, in which everything runs in a single address space (SAS) and single privilege level (SPL). This includes everything from low-level kernel components to higher-level OS services, drivers, libraries, and more, all the way up to user applications. Protection and isolation are provided by means of compiler and language-ensured type safety and memory safety.
https://www.theseus-os.com/Theseus/book/design/design.htmlYou're right, but maybe this is the way to go and the tradeoff to accept. After all, the ideas behind Theseus feel so obviously correct, alike to those of Nix.
But in general I think these types of experiments show that operating systems could be improved with greenfield designs.
Reminds me a tiny bit of Mirage OS. https://mirage.io/
I'm not sure, but my guess is that the author might want to checkout out Barrelfish and that there might be ideas they might find worth stealing from there.
And to be fair, exokernels are probably one of the least understood forms of kernel. I feel like professors/textbooks/papers are legally required to explain the concept poorly.
Does this repo build a standalone OS the runs on a bare machine, or does it run in a VM like QEMU, or is it a Rust application program that runs hosted on a conventional OS?
Programming languages also have considerable variation when it comes to dependency management. There is a menu of options w.r.t. symbol lookup and dispatch, ranging from static to dynamic.
I’ve mostly been an application level developer. Something interesting about OS development is system function calls behave differently when called in different contexts (privilege levels, capabilities). One mechanism in play is dynamic dispatch. The end effect is that the underlying details of a system call can be quite different for different callers.
(The Unison language has some innovative ideas IMO.)
If you come across a good resource that covers these topics as a “design menu” across different OS styles and/or a historical look at security mitigations (as opposed to only a summary of what is used now in Linux), could you share it?
it's... an array of pixels? does that imply all apps must draw themselves?
Looking at the source code of the cursor app [0], we can see it draws the mouse cursor and skips the rest. The transparent console app [1] is doing something more complicated which I haven't tried to fully understand, but which definitely involves massaging pixels and even temporarily saves pixel data in ctx.store.b1 and ctx.store.b2 so it looks like some kind of double buffering.
[0]: https://github.com/Ruddle/Fomos/blob/cba0460af59e63f46c7646f8a2f29d574ff0d722/app_cursor/src/main.rs#L75
[1]: https://github.com/Ruddle/Fomos/blob/cba0460af59e63f46c7646f8a2f29d574ff0d722/app_console/src/main.rs#L384If apps draw themselves, and transparency is involved, this could need:
a) Save background which is overwritten ("damage areas" is a modern description, I think?).
b) Alpha blending - calculating a weighed average between background & what you're overwriting it with.
c) And maybe some kind of text-buffer -> bitmap conversion (if not done ahead of time).
Enough pixel massaging right there.
Games? Write them thar thangs in Rust.
The ideas of Theseus is akin to those of NixOS, one easily feels they are 'the one obviously correct way' of doing it. I recommend the founder's presentations on YouTube.
I do not understand how the emulation layer can defeat validations. As long as the emulator is conformant to validations, it should be okay. I’m probably ignorant of something here.
Regarding performance, in a presentation, the Theseus dev mentioned one study which showed the performance tax of context switching with attack mitigations (Spectre etc.) can be as much as %15-%30 in modern conventional OSes.
Don't need I can follow. But some apps will have dependencies -- some might even be tantamount to a standard library in other OSes. The Context provides all that too?
A good experiment, but nothing useful. Maybe the author can generate some new ideas from this.
Polling style scheduling will just be so slow if calling each executable always involves context switch. And then cooperative scheduling isn't really possible if some process doesn't play nice.
So many virtual machines run complex kernels and security features, for just one application listening on some port.
While this is a toy operating system for some standard Desktop scenario, it is already promising for another scenario, with the move of services to managed containers (serverless) in the cloud, with huge overhead (relative to the tasks performed) and slow startup time.
Then, this can be the AWS Lambda runtime for Rust (add networking and database libraries), and it will be faster and more efficient than other Lambda implementations for the other languages.
''' Exo-kernels are interesting, but it is mostly a theory. '''
is that _true_ ? what is the difference between exokernel, and hypervisor ?
Really cool idea!
...you've got to start with the customer experience and work backwards for the technology. You can't start with the technology and try to figure out where you're going to try to sell it....