> In production we’d need to think carefully whether we wanted to include a big chunk of non-hardened code in the kernel.
E.g. is the other end of that file descriptor you're writing to another process on the same machine? Intercept the calls and replace them with operations directly on shared memory using a futex to coordinate, for example. Or is the process doing read()'s that are smaller than what is most efficient on the current system, and it's the only reader? JIT buffering into user space.
A lot of what you might consider it for could be done just fine in user-space with sufficiently smart libraries that knows how to handle a bunch of special cases, but the caveat is that you then depend on users knowing to use them, and my experience is that there is enough code out there doing tiny read()'s that I have no faith whatsoever that people will do a good job at even the very basics. Meanwhile the kernel has a lot of information very easily accessible that a userspace implementation might have to jump through hoops to find out.
There is a lot of space between 'fully Turing complete' and 'useful' where you can allow certain patterns but not arbitrary code. This area gets discussed every ten years or so.
There was even something called Proof Carrying Code where you make the caller do some of the work up front.