Is the RAM mostly used by page content read by the NICs due to kTLS?
If there was better DMA/Offload could this be done with a fraction of the RAM? (NVME->NIC)
If there was no need to TLS, would the RAM usage drop dramatically?
Yes, the RAM is mostly used by content sitting in the VM page cache.
Yes, you could go NVME->NIC with P2P DMA. The problem is that NICs want to read data at once TCP mss (~1448b) and NVME really wants to speak in 4K sized chunks. So there needs to be some buffers somewhere. It might eventually be CXL based memory, but for now it is host memory.
EDIT: missed the last question. No, with NIC kTLS, the host RAM usage is about the same as it would be without TLS at all. Eg, connection data sitting in the socket buffers refers to pages in the host vm page cache which can be shared among multiple connections. With software kTLS, data in the socket buffers must refer to private, per-connection encrypted data which increases RAM requirements.
Back when I was in NetApp, folks had researched on splitting 4k chunks as 3 ethernet packets (NetCache) line so that they'd happily fit and issue 3 I/Os on non 4k aligned boundaries. There was also a similar issue to reassemble smaller I/Os into a bigger packet, because some disks were 512b blocks back then. The idea was to give multiple gather/scatter and the engine would take care of reassembly.
Really looking forward to what interesting things happen in this space :)
B. What’s the current biggest bottleneck preventing higher throughout?
C. Has everything been up streamed? Meaning, if I were to theoretically purchase the exact same hardware - would I be able to achieve similar throughout?
(Amazing work by the way in these continued accomplishments. These posts over thr years are always my favorite HN stories.)
b) Memory bandwidth and PCIe bandwidth. I'm eagerly awaiting Gen5 PCIe NICs and Gen5 PCIe / DDR5 based servers :)
c) Yes, everything in the kernel has been upstreamed. I think there may be some patches to nginx that we have not upstreamed (SO_REUSEPORT_LB patches, TCP_REUSPORT_LB_NUMA patches).
That's pretty impressive if it's literally zero.
How many machines are deployed with NICs?
2. On amd, did you play around with BIOS settings? Like turbo, sub-numa clustering or cTDP?
Yes, I've spent lots of time in the AMD BIOS over the years, and lots of time with our AMD FAE (who is fantastic, BTW) poking at things.
How did you deal with the hardware/firmware limitations on the number of offloadable TLS sessions?
Is there a plan to move to the Connect X-7 eventually?
Depending on the bandwidth available, that'd be either 2x to get the same 800Gb/s as here (or perhaps eventually with 4x to get 1600Gb/s).
Since the alternative for an ISP is to be carrying the bits for Netflix further, the likelihood is they’ll devote whatever space is required because that’s much cheaper than backhauling the traffic and ingressing over either a settlement-free PNI or IXP link to a Netflix-operated cache site, or worse, ingressing the traffic over a paid transit link.
Meanwhile, on the flipside, since Netflix funds the OCA deployments they have a strong interest in not “oversizing” the sites. That said I’m sure there is an element of growth forecasting involved once a site has been operational for a period of time.
And If ZFS, what options are you using?
I found this extremely interesting. ZFS is almost a cure-all for what ails you WRT storage, but there is always something that even Superman can't do. Sometimes old-school is best-school.
Thanks for the presentation and QA!
The most interesting use of ZFS for us would be on servers with traditional hard drives, where ZFS is supposedly more efficient than UFS at keeping data contiguous on disk, thus resulting in fewer seeks and increased read bandwidth.
Also, what percentage of CDN traffic that reaches the user is served directly from your co-located appliances?
It seems like the bottlenecks are different, so if you serve X% of traffic on the NIC node and Y% on the disk node, you might be able to squeeze a bit more traffic?
Also, how amenable to real time analysis is this? Could you look at the request where the request comes in on node A and the disk is on node B, and tell which NIC is less loaded and send the output through that NIC. (Some selection algorithm based on relative loading anyway)
It'd be neat if you could teach the page cache to do replication for you... Then you might use SF_NOCACHE for not very popular content, no option medium content, and SF_NUMACACHE (or whatever) for content you wanted cached on the local NUMA node. I'm sure there's lots of dragons in there though ;)
My motivation for asking comes from these findings in the pdf,
Did the graph show the bottleneck contention on aio queue? Did the graph show that "a lot of time was spent accessing memory"?
What made freebsd a better platform compared to Linux to begin tackling this problem?
Thanks! Super interesting. Both a freebsd fan and I have workloads that I'd love to explore benchmarking to squeeze more performance.
We have an internal shell script that takes hwpmc output and generates flamegraphs from the stacks. It also works with dtrace. I'm a huge fan of dtrace. I also make heavy use of lockstat, AMD uProf, and Intel Vtune.
> Did the graph show the bottleneck contention on aio queue? Did the graph show that "a lot of time was spent accessing memory"?
See the graph on page 32 or so of the presentation. It shows huge plateaus in lock_delay called out of the aio code. Its also obvious from lockstat stacks (run as lockstat -x aggsize=4m -s 10 sleep 10 > results.txt)
See the graph on page 38 or so. The plateaus are mostly memory copy functions (memcpy, copyin, copyout).
We already use FreeBSD on our CDN, so it just made sense to do the work in FreeBSD.
The talk is on Youtube https://youtu.be/36qZYL5RlgY
How does Linux compare currently? I know in the past FreeBSD was faster, but are there any current comparisons?
I have to imagine considerably less (e.g. 100 Gb/s instead of 800).
Looks like Intel's release is coming January 10. https://www.tomshardware.com/news/intel-sapphire-rapids-laun...
https://www.youtube.com/watch?v=36qZYL5RlgY
Some really great talks this year from all the *BSDs, highly recommend checking a look: https://www.youtube.com/playlist?list=PLskKNopggjc6_N7kpccFZ...