Back in 1990, the Amoeba system revealed 2x improvement in throughput and 5x improvement in latency for its RPC over the SunRPC of the time. [1] Relative measurements for the V-System and Sprite were similar. QNX, a microkernel-based system with high commercial success and a long history (used by over 40 automotive manufacturers) has highly fast IPC by integrating message passing with the CPU scheduler. After years of research, Jochann Liedtke introduced L4 in 1996 [2] using only seven generalized calls with a 20x improvement in speed over prior art such as Mach. Mach, by the way, is the glaring exception to microkernel performance because of complicated message packing and port rights checking. Yet the Hurd developers are doing well in optimizing it, and to this day Mach is used to malign microkernels by people with no background on the subject. OKL4, in turn, has shipped in ~1.5 billion devices by 2012, powering the baseband processor behind nearly every mobile phone. [3] MINIX 3 further only shows a ~5-10% performance drop relative to monolithic Unixes, this for 2006. [4]
They never learn.
[1] http://www.scs.stanford.edu/nyu/03sp/sched/amoeba.pdf
[2] https://homes.cs.washington.edu/~bershad/590s/papers/towards...
[3] http://www.creativemac.com/article/OK-Labs-Software-Surpasse...