Improving NGINX Performance with Kernel TLS and SSL_sendfile
nginx.com
nginx.com
BTW, FreeBSD 14 supports cha-cha poly. But is far more CPU intensive than GCM, so I'd advise against using it.
Public key crypto systems seem like they would be much scarier to have hardware acceleration, though I’m sure if you broke it down to low level enough bits you could make it impossible for the hardware to “break” it (aside from, well, if it decided to just maliciously subsitute your code for its own. But it could do that without extension ISAs.)
https://www.intel.com/content/dam/www/public/us/en/documents...
Choosing ciphers for ease of the client makes more sense, IMHO, when the client is really constrained, like a feature phone or tiny IoT things.
I've been in the room with some web scale companies, but never was aware of any big excess TLS termination costs, but maybe I just didn't ask and they had some specialized hardware/software that I didn't know about.
My experience is that at small (think less than 700 bytes or so) packet sizes, TLS overhead can be a very real cost as you scale up, and it's easy to hit "walls" where really you just need to buy more servers because the engineering cost of getting your servers to really saturate their NICs isn't worth it. How much of a big deal that is for you will depend a lot on exactly what sort of a thing you're serving, though.
TLS termination costs are a large factor for Netflix because their CDN appliances want to have as much throughput as possible in a small space. At 'web scale', it wouldn't be a big deal to be happy with 100gbps througput and use 4 boxes; but getting ISPs to install more boxes is hard, it's better to get one box to work as fast as possible.
Personally, I've seen it be an issue; dual xeon 2690 v1 or maybe v2 couldn't push 10G via TLS, but could via plain text. But the v3 chips had better acceleration and could do 2x10G no problem. I never got access to more than 2x10G networking, so no idea what the limit was. That was under managed hosting (dedicated bare metal), so no ability to do networking beyond what was offered. At Facebook, they tended to use smaller servers and more of them, and I didn't deal with TLS as much.
And doing TLS in userspace, rather than kTLS is even worse, because it disables sendfile. That means you now have extra copies across the user/kernel boundary.
Also, while the tx side has seen lots of investment (from CDN companies/owners), the receive side usually comes later. For instance, it's not supported for TLS 1.3 in openssl (although there's an open PR).
TLS session management is rather hairy. Judging by their Linux numbers, I'd take the performance hit over pushing something that complicated into the kernel.
In general you have a point but it’s a judgement call like many things in engineering. Even in the absence of specialized TLS hardware, TLS operations are so common, there is a strong case for pushing it in the kernel if that improves efficiency by a double digit percentage.
SSL_sendfile in particular is an efficiency boon for large static site hosts, it could result in significantly less hardware waste and/or reduced power consumption.
Pushing complexity into the kernel only makes sense when it's not immediately thereafter offloaded to easily-isolated hardware.
Can you be more specific? Is this something that could be done with non-root privileges and without explicit coordination? Like normal TCP sockets?
Modern hardware interacts with software ("drivers") by ringbuffers & data areas.
To support this use case, the hardware generally provides multiple "logical devices". For example, https://en.wikipedia.org/wiki/Single-root_input/output_virtu.... They're defined so that handing an untrusted party control of a logical device limits what they can do, e.g. what VLANs or other network overlays they can interact with. Each one gets its own ringbuffers etc.
Another use case for the same hardware idea is virtual machines. Here, you can think of a userspace process as a virtual machine, minus all the overheads and pretense of being a whole computer.
A userspace process is given access to the memory areas containing ringbuffers & data areas for a logical NIC. A library acts as a driver, and controls the NIC. All interaction is just reads & writes to memory, after setup the kernel is not involved at all.
Bulk ciphering isn't hairy, and it's the same approach as IPSEC; userland negotiates sessions, kernel does the bulk ciphers.
Also, you're thinking of CPU-accelerated crypto, but you're missing two other use cases: here, coupled with sendfile(2), you can reduce the number of back-and-forth between kernel and userspace when you already know what will be written on the socket. The other use case is the (few, for now) network cards that do TLS in hardware, meaning going at line-rate, whatever your CPU speed is (not sure you can do single-thread 200Gbit/s crypto on x86_64).
The other way to eliminate protection boundary crossings is to push everything to userspace, a la Snabb, DPDK etc.
I wonder if this is still the case with 3.15?
Edit:
I figured I could check for myself. I don't know for sure what the default kernel package is, but there apparently is a linux-lts package. After installing this package, it leaves a config-lts file in /boot which, when grepped, returns:
# CONFIG_TLS is not set
The more I learn about Alpine (and musl), the more I don't want to use them. It appears as if I have an inherent performance penalty serving https web sites with nginx when I do it from Alpine.
>The following OSs do not support kTLS, for the indicated reason:
>Alpine Linux 3.11–3.14 – Kernel is built with the CONFIG_TLS=n option, which disables building kTLS as a module or as part of the kernel.
and even recommends building OpenSSL and Nginx 3.0 yourself anyway, so looks like it will be a while before this might be available out-of-the-box for most major dists. But of course everything is OSS so you can DIY if you don't mind getting some ./configure under your fingernails :)
I like alpine because it's simple enough even for me to wrap my head around and understand what's going on. None of what I'm serving is high traffic or complex enough for this to matter to my usecase - and I suspect this applies to many people's situation.
I appreciate musl/alpine for their stability and simplicity I suppose, and a bit of performance is an OK price to pay in my mind.
This is a weirdly alarmist take on this? If you're trying to use bleeding edge kernel features, which this basically is, you should probably feel comfortable using an alternative kernel because odds are pretty good you're gonna have to update sooner rather than later for some bug fix or other.
It's just not really reasonable to expect all distros to enable all kernel flags all the time, a lot of them are not really proven safe or secure. Especially when they're new.
It's an observation. I didn't intend for it to be alarmist, if you interpreted it that way then perhaps I could have worded it differently.
> If you're trying to use bleeding edge kernel features, which this basically is
It's been in the kernel since 2017 (the article noted kernel 4.13 which was released 2017 when I looked it up). That doesn't seem very bleeding edge to me.
> It's just not really reasonable to expect all distros to enable all kernel flags all the time
Of course. And I've already been considering moving away from Alpine for at least some use cases, and this can lead me to use move away for more use cases.
Alarmist is perhaps the wrong word. What I mean is that this is a very strange and high bar for choosing a distro. You aren't "suffering a penalty by using alpine," you're being a beta tester by using a non-LTS ubuntu with a bunch of random flags on or whatever. You can also just.. use a different kernel version with alpine (or whatever distro), no one's stopping you.
> It's been in the kernel since 2017 (the article noted kernel 4.13 which was released 2017 when I looked it up). That doesn't seem very bleeding edge to me.
You can't actually use the version in 4.13 though, you need at least 4.17, because apparently that's what openssl 3.0.0 requires.
Now 4.17 has been around for a while too! But you also need openssl 3.0.0 to make practical use of it, and that's only been out since sept 2021. And also had a massive number of breaking changes.
And then you have to be using a newer kernel version than that to get tlsv3 ciphers apparently. Looks like somewhere around 5.10, though it doesn't explicitly say in the article afaict. If you don't use that then maybe you're gaining some speed but you're also downgrading your security.
And then you need to use bleeding edge nginx and manually compile that against openssl3.
So yeah. It's technically been there for years. But in practical terms no one (or very few people at least) have been using it in anger until the last few months. A new syscall in linux is "bleeding edge" for a while.
Honestly the kernel is the least of your concerns here.