It's a bit hard for me to reconcile "<1% of the CPU load" (Google) with "up to 50% capacity hit" (Netflix).
It's a bit hard for me to reconcile "<1% of the CPU load" (Google) with "up to 50% capacity hit" (Netflix).
https://secure.freshbsd.org/search?project=freebsd&q=netflix+sendfile
TLS kicks all that in the teeth - each encrypted stream needs its own private buffer which has to go in and out of the CPU before it gets anywhere near the NIC. That costs memory bandwidth (a mere 400Gbps with quad-channel DDR3), memory you could have otherwise used as cache (keeping in mind you need to hold onto it until the client ACKs it), and CPU cycles (probably still a decent chunk of overhead even with AES-NI).So, in their case, they had already optimized things to such a degree that they were having a lot of throughput and very little CPU usage per byte. In fact, it's possible that on their old system the data streams wouldn't have to be touched a byte at a time, as they could be loaded into main memory via DMA from the disk and once that finished they could be transferred to the network card, again via DMA. Once sendfile() is called, no further context switches are necessary for that request.
Testing with `openssl speed -evp AES128` on an i5-4670K I'm able to get 6.7 gigabits of encryption out of AES-NI. I tried running two concurrent copies and they each got that speed, which tells me that it's per-core. That's easily enough speed to saturate a 10G connection.
It's quite possible that they have much older or slower hardware involved, since when you're building a server like this normally the network or disks are the limiting factor.
Bit of a crappy test, though - tight loop over a single 8k buffer (used in-place). It probably executes entirely in L1/L2.
They provide details of the hardware they're using: https://openconnect.itp.netflix.com/hardware/
The fanciest in their "IO-optimized" SSD-driven beast is a 2.7GHz 12 core Ivy Bridge. You've got a 20% clock advantage, and Haswell reportedly reduced AES-NI clock latency from 8 to 7 cycles for a ~14% improvement. Doing the math I get 55Gbps for the lot.
That's without doing any packet processing, no waiting on memory accesses, no interrupts, no context switching, just perfect scaling with a tight loop over the same 8k buffer - and they're driving 4 10Gbps NICs. Not much left over.
L1!
Gmail: Lots of small requests which each last a short time.
Netflix: Small number of large requests with connections that stay open a long time (possibly hours).