Heck, I have worked on an algorithmic trading platform that in the limit of 5us receives market data, dedups it (multiple multicast streams for redundancy), uncompresses it (fricking zlib), parses it, analyzes it, sends to multiple algorithms which decide if current market situation matches certain rules, decides and fills market order, the order gets inspected by independent mechanism to stop the algorithm if it malfunctions, and only then it gets sent to market, over TCP, which is another form of IPC.
All in the span of fricking 5us which is 40 times less than the benchmark suggests for this simple task. Granted, the algorithmic trading world goes to great lengths to avoid overhead including kernel overhead any kind of task switiching, branch prediction fails, etc. But still, come on, guys...