How Netflix works: the stuff that happens when you hit Play
medium.com
medium.com
This article, in my opinion, is way off. I base that on the fact that I've been dancing with Netflix for a month, I might end up working with them to try and make NUMA machines serve up content faster. As in I'd be working on exactly what happens when you hit play.
My take on how Netflix serves you movies is nothing like what this article says.
They have servers in every ISP, the servers send a heartbeat to a conductor in AWS, the heartbeat says "I've got this content and I am this over worked", when you hit play the app reaches out the conductor and says "I want this", the conductor looks around and finds a server close to you that is not overloaded and off you go.
That might look easy. It's not. Take a look at this post about how they fill a 100Gbit pipe: https://news.ycombinator.com/item?id=15367421
I'm a kernel guy, I'm old school, I get what they are doing there, that is impressive.
I wish hacker news got excited about the filling the pipe post and less excited about this thread.
There's this very widespread misconception about Netflix being a monolithic service running on Amazons' cloud infrastructure, even though the truth is that just the "rather boring" routine stuff like billing, view history tracking, suggestions, everything necessary to show the UI is running there. Netflix does a good job at this, but after all it's not that impressive when you've architected some distributed systems yourself. Not even their extreme take on the microservices concept - that is after all just a nice way of letting their devs do "their thing" in the way they want with as little restrictions from the environment as possible (which basically only works because they clearly have above-average-competence devs who can deal with these degrees of freedom).
What's really crazy is the way they squeeze unimaginable amounts of bytes per second out of modern hardware and into the internet infrastructure. They saturate 100Gbit links that usually serve an aggregation of many boxes in a datacenter with just ONE box! This is way beyond what even most above-average devs are capable of doing - you NEED old-school guys which still managed to stay on top of the crazy stack created by the evolution of hardware and low-level software in the last decade. There's not many of those out there, and Netflix apparently managed to catch a good bunch of them. These guys do the magic, and the magic they do never touches that damn Amazon Cloud. It just floats way above it.
They do one thing themselves in their stack: their distribution network for content, and they do an incredible job of it. Every post I see about Netflix's CDN for video is insightful and a learning experience.
Then they throw a CRUD app in the cloud on top of it and call it a day. Okay, a little simplified -- there's still some neat tech in the DRM, in load-balancing the CDN, and in keeping all of their tech highly available. But conceivably, Netflix could retain much of their value by simply offering all the content to other websites that displayed it to customers -- the hard part of what they do is the CDN (and contracts with content owners), and opening their platform to other interfaces doesn't change that. (Heck, Netflix might be worth more if they opened their content to other interfaces, since they're not actually very good at the front end experience.)
But when I was doing consulting (about cloud stuff), that was my advice: do the core of your business yourself (eg, CDN) then offload as much of the rest as you can.
I'm busily bringing lmbench up to date and seeing if I can get my mojo back enough to hang with these guys and do some work. It's both exciting to think about kernel work and depressing to realize so few people understand it any more.
And I concur as well on the BSD crew responsible for engineering it. The grey beards are getting harder to find as they retire and I weep each time one does.
I remember OpenConnect being mentioned back when the ISP I use came on board. The service has been really good. I haven't gotten a single buffer in like 3 years.
If anyone is curious, glance through this. It's a pretty cool initiative:
https://openconnect.netflix.com/en/
If you get crappy Netflix service you should call your ISP and point them to this page.
That's disingenuous considering the number of votes and comments each thread have received. HN IS more excited by the pipe post as reflected by user participation.
Would love to hear more detail insofar as this article is concerned.
I'm just depressed that real low level systems stuff seems to be not sexy any more. When I was at Sun the kernel group was the unchallenged top of the heap, it was the place to be. When I was fixing source management at Sun (as a hobby, it wasn't my official job) they asked me to go work in the tools group. Are you frigging kidding me? Noone in their right mind would leave the kernel group for the tools group. Sorry to be snooty but that just wasn't a thing.
I just got back from a FreeBSD conference at Netflix and was asking about the state of the world and it's depressing. People don't seem to write solid papers like they used to. Sun produced papers on vnodes, the VM system architecture and implementation, Sparc, I wrote one on making UFS perform like an extent based file system. Who is doing that now? I looked in the usual conference proceedings and it was full of academic stuff, not very interesting stuff, but no industry stuff.
Once you replaced the userspace with GNU + X11 the kernel was solid, although SunOS and early solaris didn't multitask well. For some workloads I'd end up disabling all but one CPU.
But on SunOS 4.x, you had userspace problems? That was a pretty stock BSD userspace with a lot of bugfixes. Every open source makefile in the world just worked when you typed make on SunOS. Seemed pretty solid to me, I think the only thing I would do is add perl, it took Sun a long time to decide to include perl (if they ever did).
http://mcvoy.com/lm/papers/splice.pdf
and linux took some of that. As I recall, they don't use it much because the user land apps (like web servers) end up touching the data for TLS (https).
Netflix/FreeBSD are working with the NIC folks to get the byte by byte TLS pushed down in the card along with the TCP offload. I think they'll get there first because BSD is so heavily used by CDN people (not just Netflix, Limelight is very active in BSD as well).
I believe this is actually the CERN computer center itself and totally unrelated with Amazon.
Nothing happened.
I saw The spinning loading wheel and the firefox/netflix header saying "blah blah audio video software being installed try again blah restart blah"...
I waited, I reloaded the page, I googled, I checked the DRM settings, I did blah blah blah.
Wasted time and frustrated I then opened microsoft edge and everything worked.
You know that firefox share is at 8% according to W3Counter? You know this kind of crap will only reduce this?
And then we'll only have the big daddy corporate sponsored browsers.
And all because users want to watch netflix.
What a bag of mediocre horse shit.
If there is any netflixer around here: please, please push for better testing to prevent that from happening again (i.e: firefox should be one of the test platforms!)
Tried to get support on the firefox subreddit and they told me I shouldn't be using Firefox Beta if I'm not "technically inclined". Okay, so I switched to the normal release and it was too slow for me to even bother. Back to Google...
I'm wondering if their new quantum engine fixes issues. Are you on that newest version?
Using TLS is a lot more expensive so that costs them money, so I have to respect them for that.
Grandma's TV is still unencrypted, but anyone who updates their client is protected.
They do have to deal with a whole ton of legacy clients.
i think it’s just another instance of humans making things more complicated than they need be. Same line of reasoning Linus went with a monolith vs a micro kernel.
IMO by using HTTP to communicate, they ended up being significantly closer to the original concepts behind OOP and message passing.
http://lists.squeakfoundation.org/pipermail/squeak-dev/1998-...
Now, whether microservices are successful in dealing with the problems they set out to solve, and are worth the tradeoffs they entail is still up for debate.
"... on a huge service like Netflix the entire application going down because a change was made to one part of it..."
To extend your OS example, just because there was a kernel panic in the Bluetooth stack is no reason to stop servicing requests over the Ethernet NIC which is already bound by the web server. There's also no need to edit or recompile the Ethernet driver because of a Bluetooth problem; see others posts about encapsulation.
They should probably update this to say "tries to prevent", as NF DRM has been long cracked.
Circumvented sure, but the actual cryptography components haven't been cracked as far as I am aware.
https://www.reddit.com/r/Piracy/comments/6pkypj/direct_strea...
> An Amazon Web Services data center in Frankfurt, Germany, specially dedicated to CERN.
That is because in my case, the data path portion basically talked directly to a couple PCIe boards. It bypassed the entirety of the kernel outside of some setup API's to claim memory/interrupts/etc. That meant the transfer limits generally came down to lack of PCIe or memory bandwidth (depending on which generation of machine/configuration we were using). The CPU's in the machines spent 99.99% of their time running code we wrote. Despite the talents of most OS developers, generic OS/driver code is not optimized for absolute performance in one case, rather it tends to be tuned to perform well over a wide range of situations. The general goal is to be a fair arbitrator of system resources to multiple competing processes. Further, most general purpose OS's are under the assumption that I/O is slow or low bandwidth. Take the entirety of the linux filesystem/block layer/scsi layer, which is written under the assumption that the system is attached to a high latency low bandwidth spinning disk, so burning a few cycles coalescing requests, or handling the page cache isn't a big deal. That code doesn't scale when you plug it into a NVMe disk with 2GB/sec of bandwidth, much less a storage network with 100GB/sec of IO bandwidth.
Anyway, if you throw all these assumptions away and ignore modern "best practices" development models of assembling piles of unrelated libraries to solve a task, you end up with really lean (probably fits in the L1i cache) software that can perform two or three orders of magnitude faster than similar code written using modern methods.
Netflix connections are typically about 1mbit/sec each (older apps open up ~4 connections per video for reasons that are no longer valid but the apps aren't all updated).
So to fill a 100Gbit pipe they have 100,000 connections running at the same time. Which makes filling that pipe super super impressive.
But, there are a bunch of different ways to solve the problems. I guess how impressive it is depends on they have gone about solving their particular cases. There is a fair number of network accelerators that offload individual stream level management to little cores running on the network adapter itself. Cavium, EzChip and now even companies like mellanox are playing in this space https://www.enterprisetech.com/2017/10/04/mellanox-etherneta....
So, i'm not sure the impressive parts are necessarily in the stream counts but what they must be doing to "align" (for lack of a better term) them. AKA the trade offs between keeping a few seconds of a video stream in RAM vs sourcing it from disk/wherever so that multiple users streams are aligned to avoid having to hit a secondary storage medium. In netflix's case I suspect that requiring fairly large buffers on the endpoint allow them to get away with a much lower QoS metric on any given stream.
Put another way, at least the few times I've watched netflix's bandwidth usage, it seems to be bursty. It blasts a few 10's of MB/s of data and then sits idle for a few seconds while the stream plays and then you get another chunk.
They are using either Chelsio or Mellanox cards and they use the offload but they are doing TLS with the Xeon cpus. So they are getting 100Gbit while touching every byte.
And don't under estimate how hard it is to do 100,000 TCP connections. When I was at SGI we had a bunch of big SMP machines (I think they were 12 cpu Challenge) that someone was using to serve up web pages (AOL? It was someone big). Modems brought that machine to its knees. You would think that would be easy but it was not. A single (or small number of) fast streams is easy, a boat load of slow streams is hard. Think about it, if you have a TCP stack that gets a request and then nothing, you have all the overhead of finding that socket, doing that work, then nothing. It's way easier to have a stream of packets all for one socket.
It's that sort of stuff that they worked on so far as I can tell. Your caching idea is nice but the cache hit rate is very very low. They did way more work in the sendfile area, managing the page cache. Did you read Drew's post? It's worth a read for sure.
sendfile() is good, but the general concept tends to waste far to much time doing filesystem traversals, buffer management, dma scatter gather lists, and a bunch of other crap that gets in the way of getting a blob of data from the disk, encrypting it, and passing it off to a send offload to handle breaking up and apply the TCP/IP headers/checksums. Frankly the minimum MSS size is something that ipv6 should have fixed, given that no one is on 9600bps modems, but didn't.
Good for them for realizing that modern machines have a little less than a GB of bandwidth per pcie lane per direction, and memory bandwidth to match. If you don't mess up the CPU side of things you can even touch all that data once or twice and still maintain pretty amazing I/O numbers.
EDIT: Also in the case of x86 NUMA, you _REALLY_ want to make sure that the nvme/source disk, the memory buffer your writing to and the network adapter are on the same node with the core doing the encryption/etc. That is pretty easy if the "application" controls buffer allocation/pooling, but much harder with a general purpose OS which will fragment the memory pools.
It's not fair to Vimeo, because they actually had a copy and YouTube didn't, but because of multiple experiences like that I was surprised to see a positive remark about their quality.
It's the youtube content creators like PewDiPie, Casey Niestat, Kurzgesagt etc that makes Youtube special just like that community that makes Hacker news special.