A big penalty is due to (failed) branch prediction: a coroutine switch is essentially an indirect branch so the predictor might need help (for example using a different branch instruction address for each coroutine type); also call instruction used to call into the coroutine switch function is not paired with a ret instruction, which messes with the specialized call predictor. Changing the final jmp instruction in the switch function to push add; ret actually makes thing worse; the best solution is to inline the switch function in the caller via inline assembler (this also helps with the previous issue and, if the compiler provides the functionality, it allows only saving the registers that are actually in use)
edit: last time I did a synthetic benchmark of my one of my own coroutine implementations, the switch performance was only constrained by the number of taken branches the CPU could issue (one every other cycle): https://github.com/gpderetta/delimited/blob/master/benchmark...
[1] I'm talking about the typical x86 cpu.