A hybrid thread / fiber task scheduler written in C++ 11
github.com
github.com
This is the paper: https://www.linuxplumbersconf.org/2013/ocw/system/presentati...
I'm also really curious why they require modifications to the Linux kernel. My first guess would be stronger integration with the IO model at the syscall boundary (similar to io_uring).
Edit: is this the talk your referring to? https://www.youtube.com/watch?v=KXuZi9aeGTw
With the fibers implementation you can just do that. It doesn't kill your performance, and you don't need to go to a painful async model just for performance reasons.
If by “nowadays” you mean ~2000 when I first learned socket programming (using select!)? ;-)
Boost also has fiber.
How does Google's Fibers differ from other C++ fiber implementations? What makes it so wonderful?
Marl was originally written for SwiftShader¹ a pure-software implementation of the Vulkan graphics API, which needs to run on desktop and embedded devices, with CPUs ranging from 2 to many cores. SwiftShader executes a number of parallel rasterization tasks which have complex blocking dependencies on one another. Marl attempts to simplify the problem of running and synchronizing tasks. It shamelessly borrows quite a few concepts from golang, simplifying fan-out, fan-in style problems.
The main reason of why not just use pthreads / std::thread really comes down to blocking tasks. Using regular threads, you either need to know your dependency graph ahead of time to ensure tasks are executed in dependency order (blockless), or you likely end up spinning up many more threads than you actually need to allow things to block and wait on others to complete. SwiftShader has tight requirements on the number of threads we're allowed to create, so Marl was our solution. Marl is even capable of running single-threaded with no code changes to the tasks. It's worth mentioning we also evaluated other scheduler solutions using completion callbacks, but "callback hell" was something we were keen to avoid.
I'm still looking into Marl optimizations, but it is already pretty fast. Fiber switching is notably faster than OS context switching for many of the benchmarks we've done.
I wish you could create a version based setJmp/longJmp if instrinsics weren't available (say on a different processor, like AVR). That could really help!
There is an implementation that uses ucontext, which you can enable with defining MARL_FIBERS_USE_UCONTEXT. That said, ucontext is a bit broken on macOS, and I haven't attempted to maintain this codepath, so may be removed in the near future.
Assuming you know a bit of assembly for your platform and the ABI (specifically caller vs callee saved registers), adding new platforms is not too difficult. For example: ppc64 and MIPS64 support were both added by external contributors.
We're always open to contributions, so if you write a nice generic implementation using setJmp/longJmp, I'm sure it would be accepted!
Cheers!
Some more info here: https://github.com/google/marl/blob/cbef55d588bc28661bedb822...
In short:
- The key difference between fibers and kernel threads is that fibers use cooperative context switching
- When a coroutine yields, it passes control directly to its caller (or, in the case of symmetric coroutines, a designated other coroutine).
- When a fiber blocks, it implicitly passes control to the fiber scheduler. Coroutines have no scheduler because they need no scheduler
An interesting case is a different Google fiber library, not open-sourced but described in a public talk. [1] Real kernel threads but (mostly) scheduled in userspace via new syscalls [2] that suspend the current thread and unsuspend a userspace-chosen other thread. The idea is that most of the overhead of kernel threads is the scheduling and userspace has the ability to make a quick good choice. [3] Thus you can get most of the CPU performance of an async implementation but the relatively easy-to-read callback-free code of a threaded implementation. Anyway, people can disagree on what to call such a combination.
[1] https://www.youtube.com/watch?v=KXuZi9aeGTw
[2] the rseq stuff described in that talk has been upstreamed [4] but the switchto stuff still hasn't.
[3] I'm not sure if the performance argument is as strong since the Meltdown/Spectre mitigations added to the performance cost of any syscall.
[4] https://www.efficios.com/blog/2019/02/08/linux-restartable-s...
Here's[2] a fun/interesting 15 year old blog post explaining some of the use cases and pitfalls of using Win32 fibers.
[1]: https://jeffpar.github.io/kbarchive/kb/128/Q128531/ (CreateFiber)
[2]: https://docs.microsoft.com/en-us/archive/blogs/larryosterman...