Split Stacks in GCC
gcc.gnu.org
gcc.gnu.org
If one carefully adds stack splitting and duplication capabilities to an ordinary procedural language, it amounts to a conversion to continuation-passing style. Not only does that cut down on some of the drawbacks of massive multithreading on 32-bit (address space exhaustion, though not the expense of context switches), but it also permits interesting implementation strategies (e.g. making call/cc work in C or languages implemented in terms of C, which can in turn be applied to implementation of stateless servers).
Nice summary and much more about that here: http://bizrules.info/conference/ORF2008DFW/GopalGupta_Buildi...
There are some other good presentations on C/LP in the same directory, too.
Random idea (feasible only where address space is plentiful): why not let the operating system's virtual memory subsystem worry about allocating stack space (reserve a large stack with mmap(), let the page faults be responsible for actually allocating the memory and periodically munmap() and re-mmap() the upper part of the stack.
Compilers targeting Windows need to generate stack touching prologues for procedures that allocate more than 4K of local variable data, or else they risk code first touching the page beyond the guard page, which will cause a much more unpleasant access violation.
So the solution is to dynamically allocate them, and at the end of stack routine free them. That's all good, but what if you longjmp() and have alloca() function still allocated. That would leak... Unless unwinding is done, and RAII kind of type for alloca is made. (Or maybe wrap alloca in C++ RAII, or maybe that's what the compiler would be doing).
http://ece.ut.ac.ir/Classpages/S86/ECE404/Papers/capriccio.p...
Edit: After RTFM, is the purpose of the paper, simulating threads handling blocking I/O with a epoll and/or select. Great paper.
If the minimal "reserved stack space" per thread is 4K, than a million threads would consume about 4GB of memory. That's not quite feasible in a 32-bit address-space. Would most threads run with even less than 4K of reserved stack space?
So, my question is why spend time on implementing this?
A modern OS can only allocate memory as needed with a page-level granularity.
For many millions of threads, this is not good enough and will waste GB's of memory completely unnecessarily. Maybe in a decade wasting GB's of memory will be negligible (but then we might want trillions of threads!). And pages seem to be growing too, so the OS granularity is too coarse for micro-threads.