FreeBSD optimizations used by Netflix to serve video at 800Gb/s [pdf]
people.freebsd.org
people.freebsd.org
Are you complaining that Netflix doesn't want people to pirate content, content they might have licensed from 3rd parties which contractually bind them to not let being pirated?
This + is the development/resources/cost of serving such few people on FreeBSD even worth it.
Note: I'm a huge FreeBSD fan. But consider this totally understandable on Netflix part.
It just makes normal people jump through hoops to watch the things they are trying to pay for. That’s a DRM issue in general though, I acknowledge this isn’t just a Netflix thing.
Huh, that's an interesting take. I feel like something similar might end up being what you need to do with certain video games as well.
For example, I bought Grand Theft Auto IV as a boxed copy back when it came out (though most of my games are digital now). The problem is that the game expects Games For Windows Live to be present, which is now deprecated and some folks out there can't even launch the game anymore. It's pretty obvious what one of the solutions here is.
> When a device uses DRM for the first time, a device provisioning occurs, which means that the device will obtain a unique certificate and it will be stored in the DRM service of the device ... This provisioning profile has a unique ID, and you can obtain it with a simple call. This ID is not only the same on all apps, but also it is the same for all users of the device. So a guest account, for example, will also obtain the same ID, as opposed to the ANDROID_ID.
Source: https://beltran.work/blog/2018-03-27-device-unique-id-androi...
For example on Windows with Chrome you only get 720p playback for Netflix, complete nonsense.
I don't like it, but there is some logic to it. For business types, it isn't merely the existence of ripped copies, but the ease of creating and spreading them.
I don’t think you realize how impractical that is. Take a look at the credits at the end of a movie some time. Or look up the list of people who worked on a particular episode of a show (yes, it can vary throughout a season).
There could be a the address of a smart contact at the end of the credits. Every time more than, say $1000, piles up in that address, whatever is there gets dispensed to the contributors at the end of that month.
Plex could aggregate those addresses and tell you how to allocate your payment based on how you allocated your attention. Yes I know that's what Netflix does, but I control my Plex server. Nobody is then going to find additional ways to monetize that data.
I know it's unconventional, but I really don't think it's crazy to want to reward the creators of content that you consume while simultaneously not wanting to contribute towards the development of ecosystems that prevent people from being in control of their tech.
Studios already plan for this.
For a short time in the 80's, one of my mother's job responsibilities was making sure every single person involved in the production of a movie in the 1940's got their revenue check each quarter, whether it was for $50.00, or 12¢. Hundreds of people. Hundreds of checks.
Perhaps in the 80's it would've been impractical to pay her to multiplex hundreds of $1 input checks into the appropriate set of $50 or $0.12 output checks, but that's now a job that's early done by a computer.
Or perhaps IMBD will include a wallet addresses on cast & crew pages.
Whoever does this is essentially part of the crew now and probably deserves to get paid too, but it would be an easy scam for them to just set up each contributor with a wallet that secretly they control. I'm not sure how to prevent that.
If only we had some sort of distributed ledger that can programmaticaly send payments to anyone from anyone on the network in almost any quantity large or small!
Haven't tried using netflix on it though.
>it says "use my work and don't give anything back".
That's not written.
But compared to GPL you just could read the whole BSD-2 license in 1 minute.
Looking at the 800Gbps Config, Dell R7525 with Dual 64C / 128T and 4x Connect-DX 800Gbps in 2U.
With Zen 4C, 128C and PCI-E 5.0, Connect-7, two node could fit into 2U. i.e doubling to 1.6Tbps per 2U.
That is going from 16Tbps to 32Tbps per Rack. ( Using 40U only )
To things in perspective, if every user were to use 20Mbps Stream at the same time, ( not going to happen due to time zone difference ), the 250M Netflix subscribers worldwide would need 5000M Mbps or 5000 Tbps. That is less than 200 Racks to serve every single of their customer on planet earth. ( Ignoring Storage. ) You could ship a Rack to every Region, State, Nation, Jurisdiction or Local ISP and Exchange and be done with it.
I hope Lisa Su sent drewg123 and his team at Netflix with Zen 4C ASAP to play, cough, I mean help them test it.
Note: We have PCI-E 6.0 ( and 7.0 ), DDR6 on Roadmap. The 200 Racks could be down to 50 Racks by the end of this decade. Assuming Netflix is still streaming at the same bitrate.
I doubt ISP's give an entire rack to Netflix. I wouldn't be surprised if they only get like 4U total (hence why throughput per server is so important to Netflix).
https://openconnect.zendesk.com/hc/en-us/articles/3600345383...
I think it depends on the size of isp, probably a rack would be too much even for the biggest isps, but one 4u too less.
But a lot of the racks sit at internet exchange points, where Netflix rents one or more racks at a time.
If you're going to ignore storage, Netflix could just ship a low-end video server to every one of its customers and be done with it.
Every problem is an easy problem if your pretend the hard parts don't exist.
It's got about 17,000 titles globally [1]. If they have copies in SD, 720p, HD and 4k that would be 68,000 versions (plus some extra audio tracks for stuff dubbed in multiple languages, but I suspect this is fairly minimal in terms of storage though)
Let's assume that the resolutions have the bitrates at 5, 10, 15 and 20 mbps.
The average length of a Netflix original movie is ~90mins [2]
So that would require about 575TB in storage if I have done my maths correctly.
You would need about 20x30TB Kioxia CD6 SSDs for all that. Very expensive but definitely technically possible.
I could totally see it being possible to fit those drives in a single node to push the 800gbps required, not increasing the over rack requirement at all. Not sure if the bandwidth from that many drives is enough, might have to cache some of the most watched stuff to ram)
Not gonna see any in home boxes with all the titles pre loaded any time soon though. As a hard drive array that's still 30x20TB drives.
[1] https://www.comparitech.com/blog/vpn-privacy/netflix-statist...
[2] https://stephenfollows.com/netflix-original-movies-shows/#:~...)
https://openconnect.zendesk.com/hc/en-us/articles/3600356180...
It's probably larger than we'd guess (there's likely a lot of device-specific stupid codec profile crap that may cause extra copies too).
And - even more interestingly to me - it's not a stationary target. Gotta take mobile consumption into account too!
I wouldn't be surprised if there were a few bored and cashed up engineers with private plex servers this big out there. 575TB is way smaller than I would have guessed, although 575tb of SSD is still much more expensive than 575TB of HDD.
I guess it makes sense - Netflix' library isn't that big.
> top of the line commodity hardware
Yeah -- cost in commodity hardware scales super-linearly with performance.
Considering they're really eeking every last bit of juice out of the box, I'm doubtful that distributing it would be cheaper.
Also they're not maxxing the CPU out, they just need memory bandwidth and the Mellanox nics. Storage would be more expensive on more boxes since they can't distribute the storage, they have to use local storage to reach the performance we're after.
Not necessarily. To a first approximation, if they can get 800Gbps out of the disks in a single box, they could split those disks over say 8 100 Gbps boxes and get the same performance out of the disks for the same price. Once you split it into 8 boxes, maybe instead of each box getting 2x 16 TB pci-e 4.0x4 drives, you get them 4x 8 TB pci-e 3.0 x 4 drives. Half size and older generation drives are likely to be less than half the cost. Netflix does have ways to segment cache among their appliances, so they wouldn't need to have the same storage capacity on each box as they do on the combined box.
It's certainly a procurement analysis to see if those savings will add up to overall savings, and there's a good chance it won't; especially if you need to add a 800 Gbps network switch. You do often get a pretty good cost savings by having two single socket servers vs one dual socket server though; again though, probably not if you have to add a switch.
IX.br peak traffic is 20Tb/s, DE-CIX peak traffic is 14Tb/s, AMS-IX is around 11Tb/s.
The 800Gpbs machine is probably enough for a country.
Netflix traffic stats at PIT Chile, this is their only peering connection in Chile: https://www.pitchile.cl/wp/graficos-con-indicadores/streamin...
I’d be vey surprised if the average bitrate was anywhere near the appropriation.
However that wasn’t the point of the calculation, it was looking for a maximum.
Marvel Octeon 10 DPU (with an integrated 1 Terabit switch): https://www.marvell.com/content/dam/marvell/en/company/media...
Probably pretty soon you'll be able to chuck in a few hot swappable 100 TB Nimbus exadrives (https://nimbusdata.com/products/exadrive/) in there and call it a day. 1T in 1U. :)
If only it could start playing faster.
Also are you taking into account encryption for those specs?
And when a single rack is down the whole region might be impacted negatively. The blast radius of such a setup would be huge!
As a CDN you really want to avoid this. Even for regular operations you need to be able to afford some servers going out of service for software update and other maintenance - without impacting availability, latency for users (the latter would happen if you route them to a complete different region), or infrastructure costs (which increase if you can't serve data from as close to the user as possible anymore, and have to pay for additional networking fees).
All major CDN providers have hundreds of regions (all with multiple hosts), and you can't really avoid the former. You could run less hosts per region, but whether 1 is the feasible will depend a lot on other parts of your system.
Looks like Intel's release is coming January 10. https://www.tomshardware.com/news/intel-sapphire-rapids-laun...
How did you deal with the hardware/firmware limitations on the number of offloadable TLS sessions?
Is there a plan to move to the Connect X-7 eventually?
Depending on the bandwidth available, that'd be either 2x to get the same 800Gb/s as here (or perhaps eventually with 4x to get 1600Gb/s).
B. What’s the current biggest bottleneck preventing higher throughout?
C. Has everything been up streamed? Meaning, if I were to theoretically purchase the exact same hardware - would I be able to achieve similar throughout?
(Amazing work by the way in these continued accomplishments. These posts over thr years are always my favorite HN stories.)
b) Memory bandwidth and PCIe bandwidth. I'm eagerly awaiting Gen5 PCIe NICs and Gen5 PCIe / DDR5 based servers :)
c) Yes, everything in the kernel has been upstreamed. I think there may be some patches to nginx that we have not upstreamed (SO_REUSEPORT_LB patches, TCP_REUSPORT_LB_NUMA patches).
That's pretty impressive if it's literally zero.
How many machines are deployed with NICs?
My motivation for asking comes from these findings in the pdf,
Did the graph show the bottleneck contention on aio queue? Did the graph show that "a lot of time was spent accessing memory"?
What made freebsd a better platform compared to Linux to begin tackling this problem?
Thanks! Super interesting. Both a freebsd fan and I have workloads that I'd love to explore benchmarking to squeeze more performance.
We have an internal shell script that takes hwpmc output and generates flamegraphs from the stacks. It also works with dtrace. I'm a huge fan of dtrace. I also make heavy use of lockstat, AMD uProf, and Intel Vtune.
> Did the graph show the bottleneck contention on aio queue? Did the graph show that "a lot of time was spent accessing memory"?
See the graph on page 32 or so of the presentation. It shows huge plateaus in lock_delay called out of the aio code. Its also obvious from lockstat stacks (run as lockstat -x aggsize=4m -s 10 sleep 10 > results.txt)
See the graph on page 38 or so. The plateaus are mostly memory copy functions (memcpy, copyin, copyout).
We already use FreeBSD on our CDN, so it just made sense to do the work in FreeBSD.
The talk is on Youtube https://youtu.be/36qZYL5RlgY
Also, what percentage of CDN traffic that reaches the user is served directly from your co-located appliances?
How does Linux compare currently? I know in the past FreeBSD was faster, but are there any current comparisons?
Since the alternative for an ISP is to be carrying the bits for Netflix further, the likelihood is they’ll devote whatever space is required because that’s much cheaper than backhauling the traffic and ingressing over either a settlement-free PNI or IXP link to a Netflix-operated cache site, or worse, ingressing the traffic over a paid transit link.
Meanwhile, on the flipside, since Netflix funds the OCA deployments they have a strong interest in not “oversizing” the sites. That said I’m sure there is an element of growth forecasting involved once a site has been operational for a period of time.
And If ZFS, what options are you using?
I found this extremely interesting. ZFS is almost a cure-all for what ails you WRT storage, but there is always something that even Superman can't do. Sometimes old-school is best-school.
Thanks for the presentation and QA!
The most interesting use of ZFS for us would be on servers with traditional hard drives, where ZFS is supposedly more efficient than UFS at keeping data contiguous on disk, thus resulting in fewer seeks and increased read bandwidth.
Is the RAM mostly used by page content read by the NICs due to kTLS?
If there was better DMA/Offload could this be done with a fraction of the RAM? (NVME->NIC)
If there was no need to TLS, would the RAM usage drop dramatically?
Yes, the RAM is mostly used by content sitting in the VM page cache.
Yes, you could go NVME->NIC with P2P DMA. The problem is that NICs want to read data at once TCP mss (~1448b) and NVME really wants to speak in 4K sized chunks. So there needs to be some buffers somewhere. It might eventually be CXL based memory, but for now it is host memory.
EDIT: missed the last question. No, with NIC kTLS, the host RAM usage is about the same as it would be without TLS at all. Eg, connection data sitting in the socket buffers refers to pages in the host vm page cache which can be shared among multiple connections. With software kTLS, data in the socket buffers must refer to private, per-connection encrypted data which increases RAM requirements.
Back when I was in NetApp, folks had researched on splitting 4k chunks as 3 ethernet packets (NetCache) line so that they'd happily fit and issue 3 I/Os on non 4k aligned boundaries. There was also a similar issue to reassemble smaller I/Os into a bigger packet, because some disks were 512b blocks back then. The idea was to give multiple gather/scatter and the engine would take care of reassembly.
Really looking forward to what interesting things happen in this space :)
https://www.youtube.com/watch?v=36qZYL5RlgY
Some really great talks this year from all the *BSDs, highly recommend checking a look: https://www.youtube.com/playlist?list=PLskKNopggjc6_N7kpccFZ...
2. On amd, did you play around with BIOS settings? Like turbo, sub-numa clustering or cTDP?
Yes, I've spent lots of time in the AMD BIOS over the years, and lots of time with our AMD FAE (who is fantastic, BTW) poking at things.
I have to imagine considerably less (e.g. 100 Gb/s instead of 800).
It seems like the bottlenecks are different, so if you serve X% of traffic on the NIC node and Y% on the disk node, you might be able to squeeze a bit more traffic?
Also, how amenable to real time analysis is this? Could you look at the request where the request comes in on node A and the disk is on node B, and tell which NIC is less loaded and send the output through that NIC. (Some selection algorithm based on relative loading anyway)
It'd be neat if you could teach the page cache to do replication for you... Then you might use SF_NOCACHE for not very popular content, no option medium content, and SF_NUMACACHE (or whatever) for content you wanted cached on the local NUMA node. I'm sure there's lots of dragons in there though ;)
Pretty cool stuff
I was always blown away by how much more efficient FreeBSD's network stack was compared to Linux at the time. It convinced me to go FreeBSD-only for a few years.
Do you consider that not to still be the case?
After that, Intel and AMD have introduced cheap multi-threaded and multi-core CPUs. Linux was adapted very quickly to work well on such CPUs, but FreeBSD has struggled for many years until reaching an acceptable performance on multi-threaded or multi-core CPUs, so it became much slower than Linux.
Later, the performance gap between Linux and FreeBSD has diminished continuously, so now there is no longer any large difference between them.
Depending on the hardware and on the application, either Linux or FreeBSD can be faster, but in the majority of the cases the winner is Linux.
Despite that, for certain applications there may be good reasons to choose FreeBSD, even where it happens to be slower than Linux.
I'm not denying this, but do you have a source? I've been trying to find modern "Linux vs FreeBSD" performance tests but haven't been super successful. Mostly I find things from the early 2000s when FreeBSD had a clear lead.
Do you have any data to back that up? Everything I've seen recently and my own experience tells me this isn't the case but I also don't have any data to back up my position. Would love to find some good data on this either way.
In the early years, I have run many benchmarks between them, in order to choose the one that was the best suited for certain applications.
However, during the last decade, I did not bother to compare them any more, because now the main reasons why I choose one or the other do not include the speed.
Even if I have right now, besides me, several computers with FreeBSD and several with Linux, it would not be easy for me to run any benchmark, because they have very different hardware, which would influence the results much more than the OS.
For all the applications where I use FreeBSD (for various networking and storage services), its performance is adequate, and I use it instead of Linux for other reasons, not depending on whether it might be faster or slower.
In the applications where computational performance is important, I use Linux, but that is not due to some benchmark results, but because some commercial software is available only for Linux, e.g. CUDA libraries or FPGA design programs.
Many benchmark results comparing FreeBSD and Linux may be influenced more by the file systems used than by the OS kernel.
I have seen recently some benchmark comparing FreeBSD and Linux for a database application dominated by SSD I/O, but I cannot remember a link to it.
The only file system shared by Linux and FreeBSD is ZFS. With ZFS, the benchmark results were similar for Linux and FreeBSD. However, FreeBSD was faster when using UFS and Linux was much faster, when using either XFS or EXT4 (BTRFS was much slower than ZFS). Such a benchmark was much more influenced by the file system than by the operating system.
In conclusion, it is very hard to make a good comparison between FreeBSD and Linux, because you need identical hardware, which must be restricted to the shorter list that is well supported by FreeBSD, and you need to run some micro-benchmark testing some kernel system calls.
Otherwise, the result may depend more on the supported software, hardware or file systems, than on the OS kernel.
But you're right, in the end you just have to set up both for your particular use case with the best optimizations each has to offer and see which performs better.
At the scale of Google / MS / Amazon / Apple if servers would run faster of BSD* they would use it. We're talking about 10's millions of servers here.
https://www.phoronix.com/review/bsd-linux-eo2021/7
It gives you a pretty clear picture.
There are a lot more factors involved in OS choice that could drive popularity other than the speed of the network stack. And BTW, Hotmail runs on BSD. MacOS is a fork of BSD. And Yahoo ran on BSD (and may still).
Hotmail running BSD is dead, I doubt MS is using BSD at all, they're using the same stuff as outlook.com which is not BSD.
I'm not saying Linux is better, just it is overall faster because an order of magnitude more $$$ and contribution.
I suspect most system administrators and engineers use Linux because that's what everyone else is using.
I remember noticing Yahoo properties being almost unusable in GPRS because they did packet loss detective and recovery in such basic ways e.g. no SACK.
[0] https://twitter.com/brendangregg/status/1412201241472471048
The AWS egress savings from this setup must be immense.
And you pay either by 95th percentile (basically "peak usage") or by whole link, not per megabyte sent
But yeah it would be nice if there was someone who could work on it full time
[0]: https://www.daemonology.net/blog/2022-03-29-FreeBSD-EC2-repo...
But my point was that for my requirements, 100MBit are actually sufficient and FreeBSD still is a good choice for me, I was just being snarky about it. (I do find it aesthetically displeasing, though, that my wifi is now faster than my wired network, but I can live with that.)
That's what i was thinking and brought up multipath and not LACP ;)
It is better for me to pirate their content, play it with Plex and be happy. I pay for Netflix, but still have to download it, to see it an acceptable quality. Absurd. The support couldn‘t help. It doesn’t affect, because I have my Torrent/Plex Setup, but for 99.9% of people it is a subpar experience.
I think the best years are over for Netflix. The hard awakening is here to make content that the users want and they are a movie/tv content company, not primarily a „tech company“.
Well 4K vs HD, you are right, but 480p on a Retina display right in front of me. Really obvious.
Yeah this has been the case since forever. It prioritizes instant playback vs forcing 1080p or similar.
Can't speak for iPhone, but on iPad, I've moved to using the website which goes goes to 1080p immediately.
> still have to download it, to see it an acceptable quality
Downloaded content do contain a whole lot more compression than streaming at max phone supported quality, so just a tiny FYI.
But yeah, once your hot data size exceeds cache byebye efficiency
In comparison, we pre-transcode everything to exacting standards, so all our CDN has to do is serve is static files.
But it should be noted that the FreeBSD Openconnect boxes are highly optimized to Netflix's use case. Which is serving a predefined set of content that has been pre-rendered. Youtube and its ilk are a completely different use case.
The Netflix cache is so optimized for serving Netflix movies that for many years we still used Akamai for all of our other CDN needs, but it looks like they may have finally moved that to Netflix's own CDN now.
While standing in a state of mild awe at 800Gb/s I read reviews and consider upgrading my house to 2.5Gb/s equipment... Should I just wait for 10Gbit to get a bit cheaper? Should I ditch copper and go fiber like that guy who was on the front page here recently (probably not, but that was cool)? Maybe raw single core CPU performance is starting to level off a bit, but it seems that networking technologies are still advancing a rapid clip!
Following these reports since 2015, when I compared estimated cost of your 9Gb/s server to F5 load balancer :)
how much does netflix donate to the FreeBSD foundation relative to their profits?
Took about five seconds to Google, it's the first result for "netflix donations to freebsd".
NFLX Q3 2019 earnings were about $5.2B.
So about 0.001%, I guess.
Linux is just the automatic go to. It’s great the big tech companies are rethinking these basics.