There are some very fast microkernels out there, like the L4 family, which negate the IPC overhead of microkernels by being small enough to fit entirely in the L2 cache of most processors. Linux may only have the single IPC call per round trip, but it's a fucking huge kernel and there is typically a ton of cache thrashing going on.
Shouldn't the "hot" path fit in the L1 of a "modern" processor (100 kB+)?
I'm not sure this holds today. One example I can see is related to crypto. We used to have specific hardware for computing cryptography functions but it's now handled directly in standard hardware and the software has not evolved ( but it's been faster and faster to compute checksum functions )
Adding 20% to a 10 second operation is a lot longer on a wall clock than adding 20% to a 1 second operation.
If I could get a micro kernel with a better security profile than Linux, that doesn’t sound so bad. I could easily be using a ten year old chip for my day to day browsing and probably wouldn’t notice.
I suggest that you look at genode/sculpt:)
Making that fast is fundamental to any OS.
Most microkernals have methods to do IO in a different way. If they don't then likely they would pretty slow when doing I/O heavy operations.
Nginx is IO bound in general - for my headless go stateful server, i could get syscall6 up to 20% of the profile time. It was a fight between GC in the gRPC code allocating for headers and the syscall6 for top single profile footprint.