The other way to eliminate protection boundary crossings is to push everything to userspace, a la Snabb, DPDK etc.
TLS session management is rather hairy. Judging by their Linux numbers, I'd take the performance hit over pushing something that complicated into the kernel.
In general you have a point but it’s a judgement call like many things in engineering. Even in the absence of specialized TLS hardware, TLS operations are so common, there is a strong case for pushing it in the kernel if that improves efficiency by a double digit percentage.
SSL_sendfile in particular is an efficiency boon for large static site hosts, it could result in significantly less hardware waste and/or reduced power consumption.
Pushing complexity into the kernel only makes sense when it's not immediately thereafter offloaded to easily-isolated hardware.
Can you be more specific? Is this something that could be done with non-root privileges and without explicit coordination? Like normal TCP sockets?
Modern hardware interacts with software ("drivers") by ringbuffers & data areas.
To support this use case, the hardware generally provides multiple "logical devices". For example, https://en.wikipedia.org/wiki/Single-root_input/output_virtu.... They're defined so that handing an untrusted party control of a logical device limits what they can do, e.g. what VLANs or other network overlays they can interact with. Each one gets its own ringbuffers etc.
Another use case for the same hardware idea is virtual machines. Here, you can think of a userspace process as a virtual machine, minus all the overheads and pretense of being a whole computer.
A userspace process is given access to the memory areas containing ringbuffers & data areas for a logical NIC. A library acts as a driver, and controls the NIC. All interaction is just reads & writes to memory, after setup the kernel is not involved at all.
Bulk ciphering isn't hairy, and it's the same approach as IPSEC; userland negotiates sessions, kernel does the bulk ciphers.
Also, you're thinking of CPU-accelerated crypto, but you're missing two other use cases: here, coupled with sendfile(2), you can reduce the number of back-and-forth between kernel and userspace when you already know what will be written on the socket. The other use case is the (few, for now) network cards that do TLS in hardware, meaning going at line-rate, whatever your CPU speed is (not sure you can do single-thread 200Gbit/s crypto on x86_64).