I agree that a well-designed system for most use cases won't have performance issues, since we should not just be optimizing context switches but also things like kernel bypass mechanisms, mechanisms like io_uring, and various application-guided policies that will reduce context switches. Context switches are always a problem (the essential complexity of having granular protection), and moving an extra 4KB is not negligible depending on the workload, but we are not out of options. It will take more programmer effort, is all.