Parallelizing the Naughty Dog Engine Using Fibers [video]
gdcvault.com
gdcvault.com
https://www.youtube.com/watch?v=X1T3IQ4N-3g
Ubisoft's talk spends more time getting into the weeds with atomic ops. Naughty Dog's is more of an architecture discussion. If you can only watch one, I'll recommend Naughty Dog's.
Me and a friend did this at university with the Ogre 3D engine and failed miserably, because I didn't know much abut C++ or thread safety.
We tried the actor model, every game object became an actor and sent messages around. The actors would then be spread over the amount of CPUs and the performance should rise with every CPU. In the end we got it running on multiple cores, but the message overhead killed the performance, haha.
Fiber safety is much more easy to assure than thread-safety. Under Windows using fibers is very simple (just use ConvertThreadToFiber on your current thread if you have not done already and then call CreateFiber) in opposite to POSIX systems where you will have to roll up your own implementation of fibers (or use some existing library which is not part of the POSIX standard).
I wonder if there is a good open-sourced C++11 project for this pattern? (a job/task queue)
Also how does this pattern compare with just using future/promises with a parallel executor?
https://code.facebook.com/posts/1661982097368498/futures-for...
or grand central dispatch from objective C?
That said, without proper language/compiler support, though, it's difficult to make them "safe" (for some definition). Even within fb, given these libraries, we are pretty wary of using fibers with C++ unless absolutely necessary. You gotta really need it. This is where the coroutines proposals working their way through the C++ committee can make lives better.
[1] https://github.com/facebook/folly/tree/master/folly/fibers
[2] https://github.com/facebook/folly/tree/master/folly/futures
Ironic that with deployment we're moving in the opposite direction - we've decided we really do need the complete isolation of separate VMs/containers and are willing to pay the performance overhead because process separation isn't enough.
http://web.archive.org/web/20080530073139/http://java.sun.co...
The question has always been: who deals out the work, the OS or the App? and as we can see, the question will continue to be asked, and un-answered, probably ad infinitum ..
There's also other resource overheads. For example, I would think nothing of spinning up 10000 greenlets/green threads in python. They cost maybe 1KB of memory each. But spinning up 10000 OS threads? That's a little more dicey.
[1] https://en.wikipedia.org/wiki/Fiber_%28computer_science%29
A library I personally recommend, which supports several platforms and back-ends: http://byuu.org/library/libco/
Superficially it was really easy, but turned into a goddamned nightmare. The problem is that pthreads gets on really badly with user-allocated stacks; some pthreads implementations find the current thread ID by looking at a magic value on the current thread's stack, and of course with a user-allocated stack this wasn't there.
But why was I using pthreads alongside coroutines, you ask? Well, I wasn't. But I was linking to a library which was referring to pthreads symbols, even though it wasn't calling them; which made the linker automatically pull in the thread-safe glibc, so malloc/free ended up wrapping in mutexes...
The net effect was bizarre and infrequent crashes that only showed up on machines which weren't mine and that I didn't have access to. I spent months debugging that.
Eventually I switched to simulating coroutines using pthreads with a big lock so that only one thread could run at a time. (Which actually turned out to be better, because now I could do some operations properly in parallel.)
Beware of setcontext! Bewaaaaaare!
[1] The author didn't like requiring assemblers to build his projects, so instead the coroutine switch machine code is embedded in the binary as a byte array, which is then copied to a page with WX permissions when the library is initialized.
As for assemblers, I just don't see how requiring gas or maybe NASM/YASM for x86 would complicate the build process so much. Heck, gas is a dependency of gcc. The only problem I see is for Visual Studio users, which could be solved with using NASM/YASM for Windows (which is not really a big deal when the co_swap_function has to be different due to the calling conventions anyway), or a separate MASM implementation for Visual Studio (giving you duplicate implementations with VS Windows and gcc Windows on x86 and x86_64 respectively).
I know you're consistent with your views on code clarity and that this is a hobby for you etc, but libco seems like a tiny, static enough library that it would make sense to accept a little ugliness in order to Do The Right Thing™.
(PS: Thank you for higan, I've learned so much from it over the years.)
I agree. I just feel there are exceptions, one of which is dynamic recompiling emulators.
> it could be a deal breaker for more security-conscious projects that might want to use libco (say, for a cothreaded server).
I was worried that since mprotect works on entire pages, that by clearing the write flag, I might be removing writability to certain variables. But, theoretically, static char arrays should always end up in the .data section, so it should be fine. Still feel uneasy about it.
Unfortunately, there's no POSIX way to read the current protections, so I can't just apply |X; plus that's probably not safe.
I could allocate three _SC_PAGESIZE blocks of memory (for alignment), mark the middle block R|W, copy the function to it, and then mark the block R|X. As long as OpenBSD allows those two calls, that should work.
> it would make sense to accept a little ugliness in order to Do The Right Thing™
I don't see how requiring Windows users to download yasm (64-bit) or nasm (32-bit) is doing the right thing. Every single additional step required to compile my software will result in more and more people giving up or not bothering.
0xe8a16ff0, /* stmia r1!, {r4-r11,sp,lr} */
0xe8b0aff0, /* ldmia r0!, {r4-r11,sp,pc} */
0xe12fff1e, /* bx lr */
If you do, send me a patch and I'll merge it upstream. * https://github.com/RichieSams/FiberTaskingLib
* https://swtch.com/libtask/
* https://github.com/halayli/lthread
* https://github.com/stevedekorte/coroutine
RethinkDB are using something like this:
http://rethinkdb.com/blog/improving-a-large-c-project-with-c...
http://rethinkdb.com/blog/making-coroutines-fast/