Rebuilding Netflix's video processing pipeline with microservices
netflixtechblog.com
netflixtechblog.com
I worked with this team (and its predecessors) during my time at Netflix. They achieved several "holy grails" of video encoding: a perceptual quality metric (VMAF), optimal bitrate selection per 2 second video chunk, and then optimal video chunking to be scene based rather than a fixed 2 seconds. Doing any of that in a research lab would be a challenge, but pulling it off at Netflix's scale is epic.
You might need some background on how adaptive video streaming works to fully grok this article.
But this is also just a story about a massive refactoring of a large, critical system. How many companies have you worked for that aggressively pursued refactoring/re-engineering their central systems? At most other places, I've seen risk aversion, fear, and mismanagement conspire to kill innovation. Not so at Netflix.
Andrey (who works for Netflix) drove the effort. Chatted with him about it.
I've seen streaming services serving half a million HD video users a day running on 7 servers, 2 of them being cheap vps.
They're often first to come to market, aren't intimidated to try new things, resist being bound by incumbents or gagged by censors and are generally wide open to grasping the full thro ... that's enough now.
Paramount looks like Youtube and uses twice as much bandwidth, that is almost impressively bad.
not all "HD" is made equal. bitrate is a major component of the perceived video quality. some experienced it first hand when youtube premium started pushing higher bitrate version for 1080p (https://news.ycombinator.com/item?id=34918698).
(disney+) hotstar is an indian streaming service which recently handled more than 50M+ concurrent viewers (https://news.ycombinator.com/item?id=38344265) for a live sports event. but if you actually watch the stream, it is of a very poor quality.
long story short: hub pushes for bitrate/quality much lower than netflix needs to in order to satisfy its paying customers, especially when they discriminate based on how much one pays.
but i wonder how much it would help with non-cpu scenarios when we need to do transcoding for upcoming codecs which might benefit from using hardware accelerators.
Both being extremely hard problem but I wouldn't say they are "holy grails" level. The latter part is currently being offered by other encoding services. While the first one, VMAF, is much better than all other metric we were previously using, still has many short comings. I was actually hoping all the AI / LLM model we could have something better than VMAF by now.
I guess I'm just old, but I prefer the delay with a couple of weeks of testing versus pushing to prod and having the customer test the code.
It's not necessarily anything Earth shattering, but it may be an issue at some smaller places with fewer resources.
And remember, this is the backend for video encoding. Issues in this aren't necessarily user visible.
The big benefit to their velocity is responsiveness. Apparently the backend team understands their customer and the timelines that the customer wants, and adjusted appropriately.
Just dealing with ads would have been problematic, because those tend to be straight 1080 or 4k with stereo. Nothing fancy, but I'll bet they didn't fit inside the chunk size they were expecting since ads are usually 30 seconds or less. And they don't need the dynamic encoding etc that normal titles do.
I wonder how much benefit dynamic encoding brings in space reduction?
It implies everyone is hopped up on drugs?
More stories lately are about “why we went pack to monoliths and building with borland C++”
Not long ago it was more likely “how microservices solved everything at our company, and why only morons disagree.”
So are we moving towards or away from microservices? Both. We’re maturing to use the right tools for the system.
Surely not! I want my dogmatic clickbait and LinkedIn-style grandstanding thank you very much!
How else will I stay on the hedonic treadmill of staying up-to-date with a new framework or architecture every 3 months?!
If you want to encode video, use ffmpeg. Netflix serves static movies, so encoding is going to be relatively rare and can probably be done on whatever computing resources are already available. Quality-wise, the ffmpeg/x264/x265 people probably are doing a good job already.
If you want to serve video, serve it with HLS or similar from static files stored on a CDN with a bunch of bitrate profiles. Here the problem is more creating or finding the CDN that anything to do with video.
Can't quite figure out what the purpose of all the stuff in the article could be (maybe to justify the jobs of the people doing the work?)
Ultimately it will still likely to be ffmpeg or their internal fork of it? But they are talking about per-title/per-shot optimisation here and not apply a blanket quality profile for every single video.
Netflix also produce their own series and allow them to optimise their encoding further.
They also have a Video Validation Services and Video Quality Services that do automatic quality check of the encoded video as the post indicated.
Is it complex, yes it is. But is it overly complex, maybe not?
Using per-title-encoding, allows you to have less renditions for a given video, having better cache hits.
Another example is pre-fetching data to the CDN nodes(when a new episode or season of a popular show comes out).
Its just extra optimizations that come up and are worth it at scale.
Netflix encoding is actually much more complex. Both per Scene and per TiTle encoding, optimal bitrate selection etc are a lot more than what ffmpeg offers if you do it manually.
Arguably you are right though Netflix could ignore all of that and just brute force the problem with much higher cost in terms of encoding, storage and bandwidth. But i guess at their scale it make sense to do all these optimisation.
After every ad break, the video rewinds to the start of the block before the ad break. So I have to fast forward to the ad break marker.
And people wonder why we pirate :-)
I'm presuming this is because it fires a beacon on pause, but only every X (10?, 30?) seconds on playback.
Not had many issues with it completely forgetting where I am though (it does sometimes get confused as to which episode I watched last though).
Disney is already finding out that a streaming platform isn't cheap to run and hard to do efficiently.
https://thenewstack.io/return-of-the-monolith-amazon-dumps-m...
Are there any good resources from trillion dollar globocorps that get down into the weeds?
What Amazon did, according to this needlessly snarky article that is not Amazon's tech blog, does not conflict with this. It's all theory. In reality, you should not be dogmatic and religious about your architecture choices, but empirical wherever possible. They measured utilization and cost and found they could do better in some cases with monolithic sub-systems. This doesn't mean all of Prime Video abandoned SOA.
"efficient utility of compute resources", etc. Just shorthand everything.
Does it take a famous developer to do it first for everyone to feel comfortable doing it?
No it doesn't. The rest of the monolith is just a chunk in your compiled binary sitting on disk, which is trivial in terms of resource cost. If that code is not running, it is not using any runtime resources.
Microservices will, however, greatly increase resource requirements if they lead to additional serialization/deserialization, which is relatively expensive. If you're doing video encoding, this isn't such a big deal. For web services, it is likely to be the bulk of the resource cost. This is only exacerbated in modern infrastructures where services are more and more expected to use TLS to talk to each other.
As an aside the Prime video article as a bit funny, at one point they have the line (which I hope is sarcastic but I fear that it isn't) "We experimented and took a bold decision: we decided to rearchitect our infrastructure" when their original design just obviously chose tools that didn't fit their workflow.
no no no, that takes time, the hype train doesn't wait for anyone. the sacred monolith it is. all hail the Monolith! crush the microgerms, destroy the filthy tiny services.
Needless to say, microservices are not serverless/step functions.
None. They have a different set of problems than you.
My impression seems to be that he doesn't like the container infrastructure, reflecting my own opinion, though he never calls out explicitly the infrastructure at Netflix as something bad. But every time he talks about work at Netflix, sounds about as complex as I'd image if I'd give the job to a CV driven engineer.
Jesus the quality of conversation here is not good today.
That feels fairly precise. But maybe some folks would disagree with this definition of a monolith.
FB very much does not use microservices. The closest is in infra but the www layer is very much a massive monolith, probably too massive but that's another story. They've done some excellent engineering to make the developer experience pretty good, like you can commit to www and have it just push to prod within a few hours automatically (unless someone breaks trunk, which happens).
Anyway, this person tried to reinvent everything as microservices and it pretty much just confirmed every preconceived notion (and hatred) or microservices that I already had.
You create a whole bunch of issues with orchestration, versioning and deployment that you otherwise don't have. That's fine if you gain a huge benefit but often you just don't get any benefit at all. You simply get way more headaches in trying to debug why things aren't working.
One of the key assumptions built into FB code that was broken is RYW (read your write). FB uses an in-memory write-through grraph database. On any given www request any writes you make will be consistent when you read them within that request. Every part of FB assumes this is true.
This isn't true as soon as you cross an RPC boundary... much like you will with any microservices. So this caused no end of problems and the person just wouldn't hear this when it was identified as an issue before anything was done. So th enet effect was 2 years spent on a migration that ultimately was cancelled.
Don't be that guy. When you go into a code base, realize that things are the way they are for a reason. It might not be a good reason. But there'll be a reason. Breaking things for the sake of reinventing the world how you think it should've been done were you starting from zero is just going to be a giant waste of everybody's time.
As for Netflix video processing, they're basically encoding several thousand videos and deploying those segments to a VPN. This is nothing compared to, say, the video encoding needed for FB (let alone Youtube). Also, Netflix video processing is offline. This... isn't a hard problem. Netflix does do some cool stuff like AI scene detection to optimize encoding. But microservices feels like complete overkill.
You have to have a very bad case of god complex if you look at a codebase that serves >3B users and experiences very little downtime thinking "oh yeah I could completely rearchitect that thing to be better"...
> but that's also the mindset of serious innovation
And I'd completely agree, if the project was to build FB from scratch. However in this case a software engineer that shows up in a mature codebase and wants to redo-it in a different architecture is simply immature, reckless and ignorant.
If anyone at Netflix would like some assistance, I've previously consulted in the areas of large-scale compression optimisation, and I'm sure we can get those 100KB text files down to under 20KB!
I'll help build distributed Kubernetes buzzword-compliant architectures, if that helps anyone get internal promotions as a part of this pan-cultural effort of inclusivity.
My thinking goes like this, with some simplifying assumptions. Let's say you have a monolith with 99% uptime that you rearchitect into 5 microservices, each with 99% uptime, and if any one of those services goes down your whole system is down. Let's also assume for the sake of simplicity that these microservices are completely independent, although they are almost assuredly not.
From basic probability, 99% uptime means there is some chunk of time t for which P(monolith goes down) = 1%. But
P(microservice system goes down) = P(service A down or service B down or...) = P(service A down) + P(service B down) + ... = 5%
In reality P(microservice system goes down) < 5% because they aren't independent and the chunk of time in which service A can go down will overlap that of service B. But still, that means the upper bound of the whole system going down is higher than for a monolith.
But microservices are pretty popular, and I'm sure someone has thought along these lines before. One potential rebuttal is that each microservice is in fact more reliable than the monolith, although from what I've seen in my career I am skeptical that's truly the case.
Where's the hole in my reasoning? (Or maybe I'm right. That would be fine too.)
In practice, this can make the math quite a bit messier, but I don't think it necessarily has been worse overall from my perspective.
So instead of having your system be up or down 99% of the time in a monolith, you'll have it fully up 95% of the time (using your numbers), but of that 5% of downtime, 20% of the time one of your products will be running slowly, or 10% of the time some new feature you launched won't work for specific customers in some specific region, etc.
At my company it makes things like SLA/SLO guarantees for "our services" pretty complicated in that it's hard to define what uptime truly means, but overall I think the five microservice approach, when done well, should have less than 1% of complete downtime, at the cost of more partial downtime
This is an excellent point, but what brought this to my mind was that the microservices in the Netflix article I don't think have this property. It looks to me if any of the VIS, CAS, LGS, or VES go down, then the whole service is effectively down.
Indeed, in my own career what I've seen is that if one microservice goes down the user won't be seeing 500 errors or friends, but the service will be completely useless to the user. You've just gone from a hard error to a spinning load icon, which might in fact be an even worse user experience.
It could be argued that this is just "you're doing microservices wrong", but then we start getting into no true Scotsman territory.
Exactly what it does is that first few hours of triage call goes with people claiming "well my service is up and issue is somewhere else". So find which service failed itself take crucial hours instead of fixing the failing service.
But in a world where Micro Service Incident Commanders can pinpoint failing a service among 1000 micro service within seconds on their vast 80 inch monitoring consoles and direct resolution admirals to fix in next 15 mins. It might just all work fine.
But the whole point is that by splitting it into micro-services you can efficiently and optimally scale each component individually. So it's extremely rare that VIS for example would entirely go down. And because Netflix has tools like Hystrix if one instance is unavailable it will seamlessly route to another one.
And Even if you push bad code there are techniques like blue/green and canary releases which can be used.
This suggests an addition to your model which is that not all outages are equally costly.
I would expect that over the course of years, the uptime and feature velocity of the monolith will decline at a rate faster than that of microservices, if a monolith has 99% reliability, then that inflection point might take a while to occur due to the calculation that you provided.
However, as a company you might be willing to go from 99% to 95% reliability (or 3 9's to 2 9's) to double the speed of feature development.
To my understanding, microservices are primarily implemented as an architecture for collaboration due to the inherent inefficiencies and communication difficulties of large teams. This is why they are often not recommended unless you could be categorized as "Big Tech".
When I was in web dev, my experience was that there was actually no good separation, and collaboration with microservices was in fact more difficult. Every single feature I worked on required changes across several microservices, and having Team A run service A and Team B run service B etc. just meant that I had to get buy in from every team. My team was easy enough to work with, but then for the work needed on service B I would have to learn and start using their processes, attend their standups and meetings in addition to mine, and so on.
Frankly it was a nightmare. But in a monolith, the same work is usually just a few quick arguments during code reviews. Maybe inviting a few extra people to design reviews.
In my experience, microservices just make teams more territorial.
- micro services don’t always block the pipeline; often the failing one can catch up later
- scaling can happen for each micro service
- removing faulty components from the main path means the key services are less likely to crash
- you haven’t explained why feature X is more likely to crash in a micro service than a monolith, eg, you’re assuming components A and B have 0.5% crash in a monolith but 1% when run independently
Your model ignores that most of your crashes come from the same code paths between the two models; only a small contribution is to the crash is from hosting.
Further at least where I work it is clearly that failure rate is higher than 5%. But with cottage industry of observability tools, cloud native solutions ..blah..blah.. telling basic maths to people in responsible positions is sure shot way to get fired. I am already being marked as someone opposed to progress so I can basically take my statistics and shove up mine. There is million times more data about reliability of micro services and they all can't be wrong.
I've worked on Microservices at HBO. IIRC over 2 years my team's multiple services only had 1 complete outage, and 2 or 3 impactful incidents.
Also a nice benefit of Microservices is that you can shove queues and retry logic between services, and replay messages later on when a service is back up. Obviously not appropriate for anything that needs to give real time results, but there are a surprising number of features that don't mind a 30 second or even 5 minute delay so long as success is guaranteed eventually.
Therefore I would question the assumption that things are simpler, code certainly, infrastructure, debugging and deployment certainly not.
I would say breaking things down into services, slowly, as it makes sense is enough. They don’t need to be micro.
Much simpler in what way?
You mean less close = less crash/segfault potential?
If yes, then oh c'mon, modern stacks are incredibly reliable that they almost never crash.
More microservices = more infra level stuff needed = WAY more potential problems
1) Retries. When one replica of a microservice is down, the calling service can retry, get service from an up replica, and the outage is routed around
2) Queues. Microservices lend themselves to queue and worker patterns where downtime on individual services has less effect on overall service availability
4) Outages have narrower impact. One microservice losing access to its database breaks the functionality that relies on that microservice; other functionality runs fine.
5) Changes have smaller blast radius. Most outages are caused by changes; changes in monoliths that cause outages are more likely to take the whole system offline (eg stack overflows and infinite loops crash processes). Changes that cause outages in microservices can’t knock other services offline.
Nice to see them rearchitecture their service around enshittification.
Make my nextflix better? How about cheaper? Did it deliver better content? Is this the work product of 2000 engineers focused on delivering me the worst content in the best way possible? What exactly am I getting for my 12, 20 ... wait what the hell is netflix charging now for their garbage content...
1000? 2000?d engineers at netflix and this is the article we get, this is their flex?
I am underwhelmed.
But I do consider some of what Netflix does to be a platform. Most of it is platform in the sense that some of their open source offerings are commonly adopted such as Spinnaker. But if you look at the adoption of microservices at Netflix a part of that includes Conductor, an open-source microservice orchestration engine.
The Netflix developers that created Conductor left Netflix and formed a new company named Orkes to offer Conductor as a platform. So while its not operated by Netflix, the microservice efforts they've made have been turned into a service offering.
Net negative improvement for users whilst there was presumably some net savings or gains for Netflix.
They've from offering a competitive offering to offering a compelling investment, shedding any guise of caring about their users along the way.
I think this is the year where I go back to stealing content, none of these services are worth it.
They not only have to store the movies, but access and simultaneously stream them thousands of times from anywhere on the globe.
Imagine how many people are watching things like Stranger things series premiere at the same time.
This is basically a scaling issue streaming platforms and publishers created themselves because of copyright.
It makes their product expensive, bloated, clunky and make pirated content much more convenient.
There are no excuses. Games are exactly like that. They just have to do better.
those different versions are transcoded from a mezzanine source, by a massive system, which is the subject of the OP. you can't just write off the main task from the discussion
Do I download everything to my computer, iPhone, iPad, and multiple AppleTVs in case I want to start watching a TV series one place and finish watching some place else?
BTW, of course you can download video to mobile devices
Just make a paid local tier, cheaper, but you need to load it and transcode it locally. Give the user a choice.
Are you really suggesting that it would be a good mainstream product to offer video downloads - in 2024 - where people used a computer to download movies to a computer and then upload them to all of their devices? How then do they watch on their TV with the built in apps?
They setup their own Plex server? When I want to binge 30 seasons of South Park - some on my TV, some on my phone, some on my iPad, do I copy it to all three places?
What keeps track of what I watch and don’t watch?
> Are you really suggesting that it would be a good mainstream product to offer video downloads - in 2024 - where people used a computer to download movies to a computer and then upload them to all of their devices?
If it was the sole offer, I think it would be too restrictive, but as a lower or cheaper/advanced tier, why not?
It was a pain to maintain and had symmetrical gigabit internet. Most people have cable internet with very low upstream bandwidth.
But now you’re asking them to buy another device and what’s the benefit for them?
In today’s world, I would buy an Nvidia Shield that can do hardware transcoding.
But you can already download video to mobile devices ahead of time.
Venting? Local copy would not give you anywhere (mobile/pc/tablet etc) access to their content. Imagine having to carry 20 discs (just 20, not saying 100+ yet) with you everywhere. In case, if you plan to say - 'I don't need mobile access or 20+ discs' - they aren't going to customize their offering just for you or a few users. A USB will not work with mobile devices.
Streaming service just need a relay bastion to punch through NAT for the initial handshake. You just need a server running some kind of client in your own infrastructure.
I am not sure i understand your point...plus netflix has the HUGE problem of optimizing the stream as much as possible to give people fluidity in their experience (exactly for the - low bandwidth - connections)
Giving the user a choice to pay less and run their own binary of netflix would solve most of those issues
1. Shitty connection: no problem, just wait for the movie to load.
2. Preemptive download: on demand pre-load
3. Stream optimization: would be solved by local caching and P2P offload
4. Mobile devices: either caching there (already a feature) or NAT punchthrough, accessing movies you already preloaded in your home infrastructure
5. Naive view of the world: I think many things we do nowadays would be considered naive. What, a computer the size of a chocolate bar in the hand of everyone? This just shuts the discussion off of new ways of thinking. My idea could be technically bad, but "naive" is just a way of saying "out of the common discourse".
Serving a shitton of files is sort of a solved problem. For huge bursts of a single piece of content you just need request coalescing and a few layers of fanout. If you know what content you can even pre-warm the top layer of the cache.
Sure, you need a lot of infrastructure to serve a lot of traffic. But it isn't complex infrastructure.
The hardest part of Netflix's setup is probably the player. Making it request the right quality for the network and device conditions. And IDK much about DRM but I'm sure that decryption keys add some complexity to it. Serving quick recommendations and other things are probably also much more complex than serving a small amount of fairly large files.
1. Encode the video in multiple different formats and resolutions for different devices 2. Encode the sound track in multiple different formats for different devices, and package those up alongside the video file 3. Encode the subtitles in various formats and languages
The number of combinations of the above is, by itself, super complicated, and if you pay close enough attention to the different streaming platforms you can see that they all get it wrong sometimes.
And remember that content is being ingested from multiple different sources, from internal studios to purchase agreements with small international indie studios.
Alright, so you got that taken care of, now you need to get the files out to CDNs. You have your ISP based CDNs, e.g. Comcast really wants to cut costs, you may possibly be running your own backhaul between your own CDNs, and then there are the large CDNs everyone knows of as well.
And video playback isn't just a static thing. People want to be able to pause a video on their TV and resume it on their phone, so every few seconds you are sending completion info on where the video is at, except some playback platforms are so locked down that they allow you to initiate sending data back over the network (!!!) so you have to find a way to estimate how much of the video the user has played back so far. Spend some time thinking how to do that, and you can imagine that it gets horribly ugly.
> Now me and like 100 people share a ceph storage that we all pay $100/mo for. I think the current size is like 1.5PB
https://www.comparitech.com/blog/vpn-privacy/netflix-statist...
It also has 1800 TV series and how many episodes does the average TV series have?
It also has TV shows, I wonder how shows and movies compare regarding bandwidth, storage and offers.
so it will make things better, pinky-promise!
I'm happy to pay them for the occasional good content, that I'll then torrent (because fuck smart TVs), but ... their app/client/website and their system seems to just work. I'm sure there are many things to optimize, etc, etc.. but probably they could reduce their development (and ops) budget by 70-80% if they would stop fucking with the system.
though, of course, that'd require a drastically different mindset, different people, etc.
> While it is still early days, we have already seen the benefits of the new platform, specifically the ease of feature delivery.
1) Implement Apple TV menu integration, like several other services do.
2) Bring back manual rating. My suggestions were way better then.
3) Like literally every other service that has parental controls: I wish you’d just give me a goddamn allow-list. That would be more useful than almost all other effort that goes into this stuff, and relatively easy. But almost nobody does it. It’s very frustrating.
Ways to save a lot of bandwidth:
1) Stop being the only major service that auto-plays video with audio behind the menu when folks are just trying to browse, or are just idling on the menu and talking, and in either case actively do not want that (I know there’s a setting now, finally, that sometimes kinda works for a while—how about flipping that default around?)
It’s available in most of the countries in the world, in a lot of varying devices, requiring a ton of different video processing pipelines, content delivery networks, infrastructure and etc. Even very “straightforward things” like downloading for offline viewing can be a significant effort to implement. Now think of audio sync, post processing, sub delivery, localization, partnerships and etc., you can see how you would need a ton of engineering effort to achieve it. Just the scale makes it much nuanced perspective during implementation.
You and me can dislike whatever content they’re delivering, but it’s very obvious how there are millions and millions of people who still enjoy it.
You get that your post right here is a better pitch for Netflix engineering than their own engineering. Blog about some top those problems, the things that your doing that make your domain hard and interesting...
But you can also just enjoy the story of developer achievement!
I’m not being wide.
/s if anyone took this seriously