In the M:N approach you have to allocate stack space for each goroutine that you spawn. This requires that you either know the size of the stack up front (generally not possible without being conservative and requesting a large allocation) or that you start small and grow (resulting in a lot of memory traffic and pauses in the growth case, and much harder to do in C).
By contrast, with the zero-cost futures approach we statically know exactly how much per-goroutine size we will ever need, and we can allocate precisely that amount. Furthermore, we only save the data that's absolutely needed across blocking calls. This results in much smaller per-connection state, and as a result it's quicker to allocate.
It's the difference between static and dynamic control flow. Full M:N requires us to give up static knowledge of what a goroutine will do and try to do the best we can at runtime. With futures, we have a lot more static knowledge, and as a result we can optimize more aggressively.