I was initially sold on the N:M model as a means of having event driven programming without the callback hell. You can write code that looks like pain old procedural code but underneath there's magic that uses userspace task switching whenever something would block. Sounds great. The problem is that we end up solving complexity with more complexity. swapcontext() and family are fairly strait-forward, the complexity comes from other unintended places.
All of a sudden you're forced to write a userspace scheduler and guess what it's really hard to write a scheduler that's going to do a better job that Linux's schedules that has man years of efforts put into it. Now you want your schedule to man N green threads to M physical threads so you have to worry about synchronization. Synchronization brings performance problems so you start now you're down a new lockless rabbit hole. Building a correct highly concurrent scheduler is no easy task.
A lot of 3rd party code doesn't work great with userspace threads. You end up with very subtle bugs in you code that are hard to track down. In many cases this is due to assumptions about TLS (but this isn't the only reasons). In order to make it work you now can't have work stealing between your native threads and then you end up with performance problems and starvation problems.
Next thing you realize is that you're still spending lots of memory on creating stacks for your green threads. Then you realize pthreads just let you create small stacks for your native threads and then you realize that 8Mb stacks don't mater much due to delayed allocation. I think that both Rust and Go have back tracked on spaghetti stacks since they are a lot of work require the compiler to generate extract code and they have some bad worst case scenario behavior (where you can get in a look of growing and shrinking a stack reputably due to a function call in a loop).
The final nail in the coffin for me was disk IO. The fact is that non network IO is generally blocking and no OS has great non-blocking disk IO interfaces (windows is best but it's still not great). First, it's pretty low level, eg. difficult to use. You have to do IO on block boundaries. Second, it bypasses the page cache (at least on Linux) which in most cases kills performance right there. And in many cases this non-blocking interface will end up blocking (even on windows) if the filesystem needs to do certain things under the covers (like extend the file or load metadata). Also, the way these operations are implement require a lot lot of syscalls thus context switches which further negate any perceived performance benefits. The bottom line is that regular blocking IO (better yet mmaped IO) outperforms what most people are capable of achieving using the non-blocking disk IO facilities.
This is clearly based on my own experiences. It looks like the Rust folks had similar experiences. So my hope is that anybody thinks long and hard before going down the N:M rabbit hole. I ended up studying my mistakes and history is chock full of people abandoning the N:M model. You can read about the history of NTPL (which is the threading model in Linux 2.6+/ glibc) versus NGPT which was the N:M threading model purposed. The 1:1 NTPL model was simpler and performed better. Freebsd and Solaris moved from their N:M threading models to 1:1 models.
I think the N:M model is going to keep rearing it's head in academic papers about performance of highly scalable systems but in the real world it's benefits / performance will keep being elusive. The only counter point to this is Go that seams to be making a run for it with Go routines.
I should have titled this comment "How I learned to stop worrying and love plain old threads."