That's why I said "It wouldn't be something you'd release for wider use.".
I wanted to point out there are use cases where a full-fledged kernel module would be an overkill, and some safety would be desirable.
Not safety and security from others, but from myself, so that small dumb mistakes are less likely to cause kernel crashes. For research, debugging, curiosity, and so on. Not necessarily for production use!
Although, now that I think about it, it would be nice to be able to sandbox trusted code as a part of a kernel module as well. Writing kernel code is hard work, because it needs to be truly correct.
There's often code in kernel drivers that's not performance critical, but must be done in kernel mode and must be crash-free and secure to execute. So why not take advantage of sandboxing for additional safety?
https://docs.microsoft.com/en-us/windows-hardware/design/dev...
For example, I don't remember the last time I cared about disabling bounds checking, even in C++ code (VC++ allows to keep it on).
And talking about performance With good compilers my secure memcpy_s is actually faster than the glibc or BSD libc memcpy, with compile-time constexpr.
No need to have a traditional socket API while still being able to do access it from multiple applications.
I have no interest to create such beast, but I'd be truly shocked if it couldn't at least beat a generic kernel based stack.
It'll lose some performance compared to the single application approach, but there might still be a niche for this kind of way.
Sure, you'll need to copy memory, but the data should be almost always in L3 cache anyways.
Yeah, it hurts if it's 400 Gbps ethernet. L3 bandwidth is like 50-90 GB/s. X86 just doesn't have enough bandwidth even to the caches! Better have pretty high CPU frequency, as I think L3 gets faster in proportion (but not completely sure). 200 Gbps should be somewhat fine.
Also pretty bad if the data travels over a QPI link... better have both processes in same NUMA region. And that ethernet PCI-e adapter... :-)
Regardless, I do think it'd still work way faster than anything a reasonably general kernel stack could do. Might be a reasonable compromise when process & permission isolation is required.