I wonder if this kind of approach could be generalized to all syscalls - pushing them asynchronously to shared memory, then receiving output in arbitrary order once kernel core takes care of it. Is this feasible? Does it make sense to expand this beyond IO? My reasoning is that we already usually have more cores than we need, why not dedicate some specifically for kernel stuff?