As much as I enjoy the results of the work, I'm always a bit curious how the sausage is made. Is pushing the hardware limits your primary job or something you do periodically? How do you go about selecting the gear you use? How much do you work with the vendors? (etc etc) I'd really enjoy a behind the scenes blog post or something wrt this serving absurd amounts of traffic from a single box.
But I do plenty of other things as well, including fixing random kernel bugs. You can read the git log of the FreeBSD main branch to see some of the things I've been working on..
The bottleneck is at your NIC anyways, so seems like there would be a market for NIC that can directly read from disk into NIC's working memory
In 2021 somebody submitted a patch for io_uring support in nginx:
https://mailman.nginx.org/pipermail/nginx-devel/2021-Februar...
I'm not sure if there has been further progress on it so far. In one comment feedback is "it doesn't seem to make the typical nginx use case much faster" [at that time].
But I find this interesting, because io_uring can make almost all things async that can't be used async so far in Linux (open(), stat(), etc) and thus in nginx.
Would io_uring integration in nginx be relevant for you?
Future technology advances increasingly looks like this complex work integrating hardware, OS fixes, team collaboration. People and teams and companies working together, and contributing to shared resources like FreeBSD. Tolerating mistakes at scale, giving credit where credit is due, and all the other things that make respect real, which creates a space to get things done.
Most of us will never get close to these opportunities or contexts, but still it helps us advance our own technique/culture to observe and model your story. And perhaps you'll help new collaborators find you. All the best.
I totally get the cost, convenience, and supply chain risk-value in commodity stuff that you can just go out and buy, but once you're bound to a single network card, this advantage starts to go away, and it seems like you're fighting with the entire system topology when it comes to NUMA, no? Why not a "TCP file send accelerator" instead of a whole computer?
Ah, so this is why everything stutters / falls apart when you switch subtitles on or off -- it has to access a whole different file and resume at the same place in that file I assume? I would think you would want the (verbal) audio separated out in a different file so it can be swapped out on the fly without re-initializing the video stream, and same thing with subtitle files? I'm just making some assumptions based on the behavior I've seen but would be cool to know how this works.
I've never seen this bad behavior myself. Do you mind sharing the client you're using?
Have you examined other NIC vendors? (Chelsio?)
We looked at Chelsio (as T6 was available well before CX6-DX). However, the CX6-DX offers a killer feature not available on T6. The CX6-DX can remember the crypto state of any in-order stream, while the T6 cannot. That means that the TCP stack can send, say, 4K of a TLS record, wait for acks, and come back 40ms later and send the next 4K and DMA just the requested 4K from the host. The T6 cannot remember the state, and would need to DMA the first 4K (which was already sent) in order to re-establish the crypto state, and then DMA the requested 4K. This could run the PCIe bus out of bandwidth. The alternative is to make TCP always chunk sends at the TLS record size, but this was horrible for streaming quality.
This part I don't get. How about DRM? Unless Netflix pre-DRM all contents for all user?
Netflix's DRM is sufficiently good that the Reddit Piracy subreddit has spent the last three months moaning that they have no access to 4K Netflix rips, at least for weeks or months after the content comes out.
Netflix's DRM and key management systems do what they care about pretty well at this point, which is protect the initial airing of popular shows.
It only really matters that this key is unique per package, not per user, because once even a single user can compromise the trusted execution environment and extract either the key or the plain video stream, that piece of content is now pirated anyway. So, key reuse against the same content probably isn't really a major part of the threat model - this attacker could share the key with others, but they might as well share the decrypted content instead.