Jitk: A Trustworthy In-Kernel Interpreter Infrastructure
css.csail.mit.edu
css.csail.mit.edu
Now we know the implications of Spectre make kernel JITs a potential vulnerability. You should not run untrusted code in the same address space with secrets it could steal.
Take the axioms with a grain of salt.
Being super ascetic about what code is allowed in the kernel isn't an endgame here. We don't need to take cross domain VMs completely off the table; we just need better ways of describing those security boundaries to the hardware. It's been done before. Some of Azul's secret sauce was additional cheap security modes in the MMU that they leveraged for their GC.
And trying to fully bypass cache and speculation when doing userspace to/from kernel transfers is going to suck. Remapping pages also. And if MMU is untrustworthy, well...
Which part are you referring to?
> And trying to fully bypass cache and speculation when doing userspace to/from kernel transfers is going to suck. Remapping pages also.
Here's a demonstration of doing it remotely, with specially crafted network packets.
https://arxiv.org/abs/1807.10535
It's not as hard as you're making it seem, the root here is that hardware is currently broken.
> And if MMU is untrustworthy, well...
If the MMU is untrustworthy, it doesn't matter if you allow a cross domain VM or not.
So that would make executing any code anywhere on the same cpu (anything behind the same cache) potentially vulnerable. JIT or not, kernel or not, ...
Spectre is a CPU cache timing attack to derive data. Cache lines are cleared in between context switches, so it can only affect the same process.
The solution was kernel page table isolation. Don't map the kernel and thus all of physical memory into each process. However, this now has the overhead of mapping and unmapping the kernel pages for each context switch.
There are Spectre variants that work cross address space. Hell, there's a Spectre variant that works remotely, with carefully crafted network packets.
https://arxiv.org/abs/1807.10535
And don't give me that "oh, we'd have to give up speculation and OoO cores to in order to properly fix Spectre". No, we just need more than two security domains so we can describe to the processor which code should have access to which data. Make Rings 1 & 2 Great Again.
I would have loved an option like this for example when working with a power management issue last time (PCI-e surprise removal and hot-plug can be annoying!)... Sure, I can use remote kernel debuggers etc., but those can be sometimes a hassle to set up, not always the most convenient option.
Like when an issue occurs on the other side of the world, and kernel dumps and tracing leads nowhere. Having an extra option for easy injection of code in kernel would be a nice thing to have in the toolbox.
In other words, one-off scenarios. It wouldn't be something you'd release for wider use.
Knee-jerk reactions have been somewhat off-putting lately. All of us should keep in mind other people might have a completely different mindset, problems and needs than our own. We should celebrate new ideas and novel points of view and not to brainlessly use existing dogma or cheap talking points to immediately slam them.
(I've developed (among other things) commercial kernel drivers for Windows. And "for fun" for Linux just to keep my skills updated.)
For example, I don't remember the last time I cared about disabling bounds checking, even in C++ code (VC++ allows to keep it on).
And talking about performance With good compilers my secure memcpy_s is actually faster than the glibc or BSD libc memcpy, with compile-time constexpr.
No need to have a traditional socket API while still being able to do access it from multiple applications.
I have no interest to create such beast, but I'd be truly shocked if it couldn't at least beat a generic kernel based stack.
It'll lose some performance compared to the single application approach, but there might still be a niche for this kind of way.
Sure, you'll need to copy memory, but the data should be almost always in L3 cache anyways.
Yeah, it hurts if it's 400 Gbps ethernet. L3 bandwidth is like 50-90 GB/s. X86 just doesn't have enough bandwidth even to the caches! Better have pretty high CPU frequency, as I think L3 gets faster in proportion (but not completely sure). 200 Gbps should be somewhat fine.
Also pretty bad if the data travels over a QPI link... better have both processes in same NUMA region. And that ethernet PCI-e adapter... :-)
Regardless, I do think it'd still work way faster than anything a reasonably general kernel stack could do. Might be a reasonable compromise when process & permission isolation is required.
That's why I said "It wouldn't be something you'd release for wider use.".
I wanted to point out there are use cases where a full-fledged kernel module would be an overkill, and some safety would be desirable.
Not safety and security from others, but from myself, so that small dumb mistakes are less likely to cause kernel crashes. For research, debugging, curiosity, and so on. Not necessarily for production use!
Although, now that I think about it, it would be nice to be able to sandbox trusted code as a part of a kernel module as well. Writing kernel code is hard work, because it needs to be truly correct.
There's often code in kernel drivers that's not performance critical, but must be done in kernel mode and must be crash-free and secure to execute. So why not take advantage of sandboxing for additional safety?
https://docs.microsoft.com/en-us/windows-hardware/design/dev...