Fibers aren’t useful for much any more
devblogs.microsoft.com
devblogs.microsoft.com
Response to “Fibers under the magnifying glass”: http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p086...
Response to response to "Fibers under the magnifying glass": http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p152...
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p086....
And Response to response to "Fibers under the magnifying glass", at
http://www.open-std.org/jtc1/sc22/wg21/docs/papers/2019/p152...
Gor so far seems to be ahead as stackless coroutines are part of the standard.
Memory footprint : It says fiber user stack is 1 MB and so fibers have comparable memory footprint to threads. This is not true for goroutines which typically use a 4K stack.
Context switching overhead : gives numbers for architecture, but goroutines do not use the expensive switching instructions listed in the paper. Instead golang basically saves just the PC, SP and DX registers, significantly reducing the overhead.
Dangers of N:M model : The dangers mentioned of corrupting memory etc is specific to C++ libraries and do not apply to golang.
Dangers of the 1:N model do not apply to goroutines either.
My conclusion from the paper is as follows : Fibers are bound to fail as an OS feature, or as a library. To make fibers work you need to do what golang does i.e. make it part of the language with compiler support to reduce the context switching overhead and the memory footprint. You will however, pay a price in higher FFI cost. That may be a tradeoff which may or may not work for you, depending on the nature of your application.
What is then the purpose of coroutines as opposed to just having more machines?
OS threads can scale (ipootace^), but then you need proper concurrent data structures and a complex memory model below to support them.
Does go have those?
Make it this one: https://blog.golang.org/concurrency-is-not-parallelism
It explains in detail how this is possible.
If so, that's a pretty weak argument. I use coroutines/fibers quite a bit in personal projects, but learned long ago for a variety of reasons to avoid depending on third party libraries I didn't have source for - especially ones that try to do too much magic like TLS behind the scenes just to save me the trouble of supplying an instance/context pointer to every call.
Usually when I'm using fibers it's so I can have many of them, which means I'm using tiny stack sizes, which means I'm not casually calling into third party libraries I can't easily audit and control anyways. If I weren't making many, I'd just use full-blown threads.
I really hate that this is the solution that Rust ended up pursuing. There are claims that you can work around TLS by just no using TLS if it is not available, but I have yet to see someone removing TLS and still be able to use a multi-threaded executor.
So we need to preallocate everything for all systems, but hey, if you preallocate everything for worst case - you are out of RAM (on PS4/Xbox/etc).
What we need is a mix: scratchpad memory that lives longer than stack, but costs as cheap as stack. Tagged heap with stack/arena allocator comes close performance wise but not ergonomics wise.
Ergonomics of writing code with such constrains is very painful. Stack is the most ergonomical/fastest scratchpad you can have, and as soon as you have async/await/etc in a middle of your function - you need to think about unwind/rewind of every stack variable.
Java had to rewrite the memory model of the whole JVM when they introduced the concurrent package.
Even C++ has a hard time making good use of concurrency and threads because of locking.
My hunch is that because of the way memory works there is no benefit for non GC languages when implementing threads and concurrent memory over many cores.
I've been pretty happy with Rust on multi-core servers, probably because Rust never allows mutable memory to be seen by more than one function.
I've been using async Rust at work, which turns out to particularly interesting, because most of the async executors for Rust can actually move suspended async functions between cores. So you might have dozens and dozens of async routines running across an 8 CPU pool.
There was definitely a bit of a learning curve involved, but we've only ever encountered a single concurrency-related bug, which was a deadlock. (Rust protects against memory corruption, undefined behavior and data races, but not deadlocks.)
If rust does not allow concurrent memory reads, then that's a big problem in my world.
It seems to me we have a collision between "data driven" and "parallelism on the same memory" in terms of progressing on performance and in that case since I can get parallelism with memory and execution safety with hot-deployment on a VM; the data driven approach does not fit my server side performance needs to the point where I'm able to give those other features up.
On the client, I'm all for C with arrays though.
I'm not informed enough to comment on how async works exactly in Rust. However:
> If rust does not allow concurrent memory reads, then that's a big problem in my world.
Note that this is not the goal of the borrow checking system of Rust. You can read memory concurrently just fine (given immutable references to something), you're just not allowed to write to it while you're doing that.
Basically, references in Rust come as immutable (like `const` in C) and mutable, and you're only allowed to have multiple immutable or one mutable reference to the same thing at a time. If you have a mutable reference, you can derive multiple immutable ones from that, but the borrow checker will prevent you from accessing the mutable reference as long as one of the immutable ones is still active (which Rust manages with the concept of "lifetimes").
It is true that many of the fiber issues presented here are C++-specific. However, what a lot of the comments here are missing is that C++ issues have a way of becoming your issues whenever you use an FFI, even if you aren't using C++. Go's solution is generally to try to avoid using cgo as much as possible, because of these performance issues. That can work for the areas Go is generally used in today. But, as the article points out, that does not work for all applications. For example, I would not want to write graphics code in any system with M:N threading due to FFI cost, including Go.
It’s also worth noting that FFI is not the only way to have Go and C++ interop. For many use cases a lightweight RPC layer between two apps will give better throughput, something that also is done in production to great effect.
Not really. WinForms is a lot of the reason for C#'s existence, and WinForms is just a wrapper around pinvoke'd Win32. You're crossing the boundary a lot.
> For many use cases a lightweight RPC layer between two apps will give better throughput, something that also is done in production to great effect.
I have a hard time believing that RPC can possibly be faster than cgo. You have the overhead of message serialization and deserialization, two message copies (into the kernel and out of the kernel), two context switches, and a trip through the OS scheduler.
That is an implementation detail. I also don’t know many who consider WinForms to be particularly high performance.
(Additional note: though I have not explicitly said it prior, I believe that PInvoke actually was quite slow for a long time, at least certainly during the WinForms era. For all I know, it might still be.)
> I have a hard time believing that RPC can possibly be faster than cgo.
I can’t find a solid reference, but the issue is that Cgo is simply not ideal for heavy applications. It makes scheduling slower. If you are doing expensive work in C++, such as phoning out to the network or decently heavy computation to the point where Cgo overhead is not the concern, then you are unlikely to have much issue with the cost of RPCs. If you are doing tiny amounts of work with no IO one must wonder why you would not just port those bits to Go.
(Example of scheduler issue: https://github.com/golang/go/issues/19574)
I continue to contend that considering this to be a show stopper to be unfair or at least not very honest.
I wonder if that's why WPF does so very much on the managed side. And I wonder if using UWP XAML from C# is less efficient than WPF in some scenarios because of this FFI overhead.
On their benchmarks comparing XAML/C++, XAML/C#, RN and Electron, it is hardly a few percentile more than C++.
It is Electron that goes sky high in performance loss.
From : https://codeburst.io/why-goroutines-are-not-lightweight-thre...
"In Go, this means only 3 registers i.e. PC, SP and DX (Data Registers) being updated during context switch rather than all registers (e.g. AVX, Floating Point, MMX)"
Perhaps a better title would be "Fibers require compiler / language support to be viable"
Linux seems to have gone the other direction, supporting thousands to tens of thousands of threads.
It's pretty clear that fibers cover goroutines, green threads, M:N threads, etc., this is just Microsoft-specific name for them.
See these pages:
https://docs.microsoft.com/en-us/windows/win32/api/processth...
(This feels like a weird autistic conversation, I'm going to step out now.)
An IO value is just something that the Haskell runtime is able to invoke somehow. Haskell functions can not directly run IO values (ignoring unsafePerformIO). A fairly elegant implementation would probably simply make IO values be asynchronous operations (procedures that take an "on complete" callback that receives the result)—again, since there's no way for a Haskell function to actually run the operation, all it can do is return such an operation to the runtime to be called.
This is a textbook example of why people shouldn't editorialize titles.
The right title here is "Fibers Under A Microscope".
The link comes via https://devblogs.microsoft.com/oldnewthing/20191011-00/?p=10... which summarized the PDF as """a fantastic write-up of the history of fibers and why they suck. Of particular note is that nearly all of the original proponents of fibers subsequently abandoned them [...] fibers are basically dead""".
By restricting itself to the TIOBE top 10, the paper also misses a discussion of BEAM which successfully offers N:M threading.
Changed from http://www.open-std.org/JTC1/SC22/WG21/docs/papers/2018/p136... above.
Experience has shown that switching to a better URL is generally better for discussion, though there can be a lag before the thread catches up.
If so, it is also a successful usage of Fibers/green threading in highly concurrent environment.
It's made to handle 1000s of 1-to-1 calls that all fit individually on one core, not one 1000-to-1000 call that needs to use all cores at the same time.
Parallelism without "intra cpu and inter core" forced scope has never been hard, just spin up more machines.
Or if I put it this way: "if you don't need parallelism and memory speed between the parallelism enough to encounter cache miss problems, could you as well use separate computers?"
If the answer is yes then fibers and coroutines are meaningless.
This is the problem of this decade, if we can't solve the bottle neck of cache misses on sharing memory (in parallel on one task across cores efficiently) we have no reason to try and scale the number of cores at all.
And if that's true we have hit peak Moore's law for transistor computers; in performance, years ago and in energy efficiency, probably around 8nm.
Just to illustrate the problem one last time: Naughty Dog converted it's engine to fibers to allow 60 frames per second, but the controller input to frame latency increased with many frames (because the multiple cores cooperating on each frame meant they have to push the frame back because memory is slow) so the benefit was actually reduced.
You get lots of smooth bells and whistles (that look good when you don't play the game) but the only meaningful metric (how fast the character acts on your reactions, which is the definition of gameplay) actually degrades.
That said I'm looking forward to TLoU 2 as much as the next guy, but I'm not expecting any technical improvements except visuals.
Submitted title was "Microsoft: fibers/goroutines inappropriate for scalable concurrent software[pdf]".
Submitters: "Please use the original title, unless it is misleading or linkbait; don't editorialize." This is in the site guidelines (https://news.ycombinator.com/newsguidelines.html).
The paper makes two references to Go: one to talk about split stacks, and one to reference the FFI (specifically, cgo). Cgo absolutely still has large overhead. That was true when the paper is written, and it's true today. It's inherent to the M:N small-stack design that your FFI calls that require big stacks require switching from a small stack to a big stack. You cannot get that overhead down to zero; it's a fundamental tradeoff of the design.
Go's solution here is to try to minimize the amount of FFI usage. As the paper points out, that may work for Go, but will absolutely not work for many other use cases.