Trying to emulate Mach, which was a dud as a microkernel, didn't help.
QNX is one of the few microkernels to get it right. L4 got stripped down so far that it's just a hypervisor, on which people usually load Linux. L4 took out arbitrary-length message passing in favor of interrupt-like events and shared memory between sender and receiver. This simplifies the kernel, but now it's easier for one side of a sender/receiver to mess up the other, since they share a communications area. The QNX primitive set (MsgSend, MsgReceive, and MsgReply) work well enough in practice to allow full POSIX functionality. Applications can talk to file system servers, network servers, etc. through those primitives. All QNX I/O works that way. You take maybe a 20% performance hit for the extra copying, but you get robustness in exchange.
Most important thing for performance in a microkernel: the CPU dispatcher and the message passing have to be tightly coordinated. You must be able to call another process and get a reply back without trips through the scheduler or a switch to a different CPU. QNX gets this right, because MsgSend is blocking. The sender blocks and the receiver starts without having to schedule. The data being sent is right there in the cache of the CPU, ready for use by the receiver. Good test for a microkernel - put some CPU-intensive jobs in a loop, while also running something that makes short request/reply calls to another process. If the request/reply process stalls out, the microkernel is doing it wrong. If the CPU-bound processes stall out, the microkernel is doing it wrong. Message passing should schedule as smoothly as a subroutine call. If it doesn't, performance under load will suck.