TLS in Linux 4.13 kernel
github.com
github.com
Everyone: the above gives a lot of very helpful context & explanation for this (from when the patch was first posted in Dec 2015)
Also, it seems that only particular parts of the stack are handled by the kernel. Per the README:
The socket does data transmission, the handshake, re-handshaking and other control messages have to be served by user space using appropriate libs such as OpenSSL or Gnu TLS.
This means that kernel is only doing some symmetrical encryption and simple framing (which are both the same kinds of things it was already doing in other application domains).
The paper reports a relatively small performance improvements. But I suppose if you're Facebook (they and RedHat implemented this) and performance improvements are measured in millions of dollars you care a lot about 7% less CPU utilization.
Why? If you have servers with millions of qps, the graph would show the 99% latency over, say, a second or ten.
I think this graph is pretty standard? But I'm happy for suggestions of better ones.
For example, chi-squared and if you can quantify the error distribution and it's normal, student's t-test.
For all we know, KCM/KTLS generates the highest peak latency, but fewer times per bucket. That difference would completely change the interpretation of these data. If userspace data frame handling produces higher 99th percentile latency, that does not tell us anything about its maximum latency. Further, if we're looking at the 99th percentile, 1/100 data frames have latency higher than that percentile, this will happen hundreds, thousands, or millions of times per second depending on your interface bandwidth and typical frame size!
You think 7 percent is a small performance improvement? IMHO it's rather large. I take the time for much smaller than that.
To understand why, let me ask a question do you think doubling performance is a good thing? In other words a 100 percent increase in throughput or a 50 percent reduction in execution time. I hope your answer is yes because a lot of people think so. So what's 7 percent? Well imagine there are 7 different areas of your code that could be improved with each one contributing 7 percent. That's 49 percent or about what we'd call a significant improvement. Fixing up a number of single digit things adds up. But there is a more sinister side to it. Suppose you wanted to let those 7 little things go, and look for the big fish - the one that's going to get you 50 percent all in one shot. Guess what? It doesn't exist. Why? Because these little things add up to about half of your time, so to find the big one would imply taking the execution time for the rest of the code down to zero, because that's the only place there's 50 percent left. Unless you're doing something really poorly, there isn't a big win to be had. Performance comes from compounding the effects of many small optimizations.
If just moving TLS from the application to the kernel gives 7 percent - presumably by eliminating overhead - that's a good thing. Even if it's just 7 percent of time in communication which is only a portion of application performance, it's still a good thing.
--- note, this is a comment about performance and says nothing about tradeoffs in security by putting TLS in kernel ---
I think a 7% reduction in CPU utilization for TLS is small given the vast bulk of systems that spend most of their time IO bound on network interfaces and/or storage and will see little to no improvement in throughput. I allowed that Facebook may see this as an important improvement given their scale; it could be worth millions of dollars in power and/or hardware.
>> that's a good thing ... it's still a good thing.
Sure, it's good. I didn't characterize anything as "bad," merely "small."
Mmmm, I'm afraid I've got some unfortunate news for you :)
Thanks for the heads up.
This is likely to end up being supported in systemd soon (I'm guessing) to get an SSL socket activation.
For handling encrypted secrets, this is popular as a HSM idea. You authenticate to a black box which does the crypto for you, but the key can't be extracted. Sometimes HSMs even have a physical tampering / self destruction protection.
this is how plan9 does tls and how it really should have been in the first place. adding tls support is one function call to wrap an existing file descriptor in tls and you get another file descriptor back.
Of course that is inefficient. But that is what KTLS is at its core: an optimization.
I'm almost done with re-writing it to use the same M_NOTREADY mechanism as async sendfile, and to do all the framing in the kernel, rather than doing the sendfile() framing in-kernel and sosend() in userspace. This removes the majority of the code, and makes it quite a bit simpler. The downside is that it depends on my vectorized mbufs. Hopefully we'll have something public in the next few months.
Kernel Connection Multiplexor
Facebook’s primary motivation was to gain access to the un-encrypted bytes
in kernel space. KCM is used to decode the framing, and make intelligent
scheduling choices, before sending the frames to user space. KTLS sockets
are mapped 1:N to user space sockets, where N is the number of user space
threads, which are usually mapped to cores. Using this scheme, KTLS + KCM
is able to reduce the total number of thread migrations of an individual
request[0] https://github.com/torvalds/linux/blob/v4.13/Documentation/n...
A kernel facility like that certainly makes an attack even more obvious and maybe harder to detect. But if this is a concern you can't use external hosting.
This applies to VMs too, the hypervisor can see and transparently modify everything going on. There is development going on with special CPU instructions that could prevent the hypervisor from reading VM memory, but that is not state of the art and I would not trust it for a long time.
You can't even be really sure that a bare metal server you are renting is not extracting the cleartext via a modified bios/efi, bootloader etc, or modified hardware at the worst.
Honestly sometimes I wish we had a library of these sorts of attacks so that it's easier to believe that they're easy, but there is a vague security advantage in not having easy-to-use, well-tested implementations of these attacks available for free on the internets.
eh? more obvious and harder to detect at the same time?
I will begin experimenting with this kernel this week.
In particular, my impression is that a large reason of why microkernels failed in the '90s - see e.g. OSF/1 - was the overhead of message-passing. The two most popular kernels today that vaguely resemble microkernels, namely Darwin (Mach-based) and NT (with its subsystems), do no isolation between parts of their kernel and just have ordinary function calls. TLS exists in userspace and works extraordinarily well, but the desire to move parts into kernelspace is specifically to avoid the overhead of copying data between kernelspace and userspace and to allow using things like sendfile() that wouldn't be possible across address spaces.