How Facebook encodes videos
engineering.fb.com
engineering.fb.com
Most big companies with millions of hours of video uploaded each day have realised it's cheaper to stick a bunch of hardware video encoding chips onto an accelerator board and be able to transcode 100 HD streams simultaneously into all the formats and resolutions you need to host.
The power savings on CPU's pay for the custom hardware in a matter of months.
It does reduce flexibility when new video formats get released though.
Based on the speed I see from Telestream and Elemental. I think they’re software based as well
YT & FB have a few popular videos, and a ginormous long-tail. That long-tail needs to be encoded, and cheaply as it might not get many views.
The TV/Movie VOD over IP industry stands to benefit optimizing the encoding to have the smallest filesize with highest quality. Spending considerable CPU cycles to find that best quality path.
For those like Netflix, why would you invest with ASICs. You likely don't have a huge infra bill (relatively) when it comes to encoding. And CPU vs ASICs, the ASICs mean you loose flexibility to really fine-tune the quality and use the latest codecs.
For those like YT/FB. You just gotta encode that long tail. Getting something cheap that's 80% the quality of CPU-based encoding is good enough for 99% of the content.
Also, good video codec engineers can’t be hired for any predictable amount of money, and ones who speak Verilog barely exist.
NVENC is good at cleaning the noise and fast encoding. ( Or basically Game Streaming ). Which isn't something you want to do if you want to do movies encoding.
Sure, and I think hardware encoders are great when you need speed, certainly for real-time video. But in other cases, well, the medium preset sucks. I always encode videos at `veryslow`, and there's just no way to get close to that with a gpu.
There's a lot of other free video encoding tools, like avisynth plugins, that are just better than all professional tools. I'm not sure why this is, maybe customers aren't sophisticated enough.
You can mass parallelize it by encoding different clips on different CPUs, this is more optimal because it has less communication overhead.
Even with power savings it was usually more economically efficient to run encodes on a large (12+ core) machine than to deal with limited amount of nvenc slots on GPUs.
I've been to a project where, because the development and "hardening" budgets were separate, we knowingly pushed out buggy code so that we could take advantage of the latter.
Had I known that half of the job would be to game the system, I wouldn't have joined.
Bureaucracy is inevitable once a certain number of people join a group.
You should be happy it's only half!
[1] https://arstechnica.com/gadgets/2021/04/youtube-is-now-build...
For example, if you want to only annotate 1 in 100 frames and interpolate with optical flow, you can get the flow in the process of doing video compression.
I don't have inside info but there are various other things like this that I can think of for wanting to do it in software.
Also, hardware encoders suck if the company making the hardware decides to suddenly EOL the product. With software encoders you can easily scale the backend whenever you want, at any time in the future, and without succumbing to supply chain issues.
https://news.ycombinator.com/item?id=26925307
>It does reduce flexibility when new video formats get released though.
A new video codec is released and adopted once every ten years. So I dont think it would be a problem. Hardware encoding also trade compression quality for speed. And they are compensating it with slightly higher bitrate. It will also reduce their incentives for improving open source encoder.
Although I think even Netflix switched to BEAMR ( Cant really blame them though )
I think the next frontier, for both Audio and Video will be Codec designed with LiveStreaming / Low Latency in mind.
( Or pretty much everything computing, I wish we could focus on latency, from Hardware Input, Display, Network, Disk, etc. Apple is certainly moving in that direction without talking about it. )
Convinced?
If we unpack your calculation to slightly more typical human behaviour: Watching 1 stream × 10 hours/day × 365 days non-stop, one person would need 10 years to watch the whole catalog.
Doesn't seem unreasonable to me.
sources:
[1] https://www.whats-on-netflix.com/news/how-long-would-it-take...
[2] https://www.comparitech.com/blog/vpn-privacy/netflix-statist...
Also, IIUC, they re-encoded about once a month, but the re-encoding didn't necessarily take one month of compute time.
But also this might be more about protocols, but AFAIK streaming protocols have mostly been abandoned to the favor of chunk downloads over HTTPS ?
Facebook, which videos will we transcode into what versions.
Google, with this list of videos to transcode how can we do it cheaply on a massive scale.
I didn't spot anything that pitted approaches against each other.
The advantages are:
- You don't waste CPU encoding video into formats that won't be used.
- You can use a standard caching solution to reuse those chunks.
- Everyone gets the perfect encoding, always.
- If most people watch the first 2 minutes and then give up on a 20 minute video, you don't waste CPU encoding the other 18 minutes.
- You can introduce new encodings instantly for all videos, without going back and re-encoding historical videos.
- You don't waste storage on video chunks in formats that will never be used.
- It's really simple.
The disadvantages would be:
- You have more unpredictable load if a lot of people start watching (different) videos at once (although the common case of a lot of people watching the same video is still fine) and you could "cap" the load by switching to a fallback format to avoid becoming overloaded.
- There might be an initial delay when playing or sweeping a video whilst the first chunk is encoded. On the other hand, it can't get much worse than it is already, and you could make sure these initial chunks are prioritised, or else serve the initial chunks via a fallback format.
Most of codecs targeted at low bandwidth for mobile streaming track scene changes and if you make chunks in a naive way (split via I-frame borders and encode 'em independent of each other), when reassembled final video will look choppy due to broken scene change relations.
So after encoding each chunk you will have to carefully save relevant parts of encoder state and reuse it for the next chunk. Seems doable, but tricky to get right?
Have you ever written a stateless transcoder like this? Of course it can be done, but saying “you could simply encode those chunks” and “It’s really simple” is pretty misleading especially if you are changing frame rates or sample rates or audio codecs during the encoding process.
That said, if there is someone that could do this at scale it would be Facebook.
Also this would mess up ABR streaming at least for the first people to watch the video which would not really guarantee “the perfect encoding always”.
If the input video source has been prepared properly (i.e., constant framerate, truly compliant VBR/ABR, fixed-GOP), or if your input is a raw/Y4M, then segmenting each GOP into its own x264 bytestream is rather trivial.
If the input is not prepared for immediate segmentation, it is also somewhat easy now to fix this before segmenting for processing. Using hardware acceleration a transcoder could decode the non-conforming input to Y4M (yuv4mpegpipe) or FFV1, which can then have a proper GOP set.
That's not always reasonable, though. If I upload a video at 4K, there needs to be some "baseline" encoding so that when the video is published, there's something playable without streaming 4K to, say, cell phones with a resolution smaller than 4K.
Even then, chewing on that 4K file into a 1080P resolution video "on-demand" for a desktop user on high-speed internet is no small task. First and foremost, you need to assume concurrency: if two people request the video at the same time, there's a complex problem of coordinating the encoding in a large distributed system so that the video is encoded once (or a very small number of times). You also need to do the encoding _faster than the video can be played_, or at least faster than the baseline/fallback version can be retrieved and sent. You also need to queue the next chunk(s) of video up, so you're not watching a chunk, buffering, watching a chunk, etc.
In a system at the scale of FB, it's not smooth sailing for compute jobs like this: you're subject to network latency, noisy neighbors, failures (disk/network/software/power/etc.). The case where you're able to stream the original file from storage, start encoding it and streaming the output back to storage and to everyone around the world who is requesting it _at that moment_, and coordinating the encoding of the next chunk is actually not very likely.
Want to talk about weird failure modes?
- I start watching your video and click all over the seek bar. Am I DOSing your compute cluster?
- Two users on opposite sides of the world request the same location of the same file with the same fidelity. Does one of those users get a dirt-slow experience, or do I double my compute costs?
- A thousand users start watching the same video at roughly the same time. A software bug causes the encoder(s) to crash. Do 1000 users suddenly have a broken experience, or does the video pause while your coordination software realizes there was a failure, releases the lock, and restarts that encoding job from the top while your users all get in line for the new job?
I'd argue that this is the _least simple_ approach. In the happy path, you get a nice outcome while reducing compute cost, and users get high-quality video. In the unhappy path, users get slow loading from slow encoding instead of reduced resolution, or you start to need to trade performance for compute (do you encode twice in two datacenters, or move compute further from the viewer?).
As other people have pointed out, video encoding isn't stateless. You can chunk and encode in parallel, but there are tradeoffs.
The biggest problem is that you've taken what is effectively a CDN problem (serve the first chunk of a video) into a CDN and a CPU scheduling problem
Serving video is cheap, apart from the bandwidth. So anything that reduces the number of bytes transferred yields savings. Real time encoders are not as efficient as "slow" encoders.
For low volume videos (ie 99.5% of all video) the biggest cost is storage. So storing things in a high quality, or worse still, original codec makes storage expensive. Not only that you still have to transcode on the way in, or support all codecs ever made, in real time.
In short, yes, for some applications this approach might work, but for facebook or youtube, it wont.
Anyway, of course videos are already chopped into chunks that are stored separately. It’s much easier to distribute and cache these independent chunks. On demand encoding doesn’t change that.
I don't see why this constraint is in place, you can absolutely serve video for certain-res users only with certain codecs (youtube certainly does this).
Context on that wonder: I've always noticed that news outlets seem to carry downright horrible quality user-generated video clips of rallies, protests, and the like. Where everybody's carrying around stellar-quality video gear in their pockets these days, I've never figured out why that is.
1) Livestreaming can have very poor quality when bandwidth constrained (e.g. at an overloaded cell site at a rally) 2) Viral videos get reencoded many times in many formats. The cumulative encoding errors are not only limited by the lowest quality reencode, but also by the defects in all previous encodes with various codecs.
I recently tried to implement video uploading for an open source project, but naively choosing ffmpeg parameters can often result in noticeable quality loss / large output file size / long encoding time. And easily all three of those at the same time.
The quality from going from source -> 15mbit/sec h265 -> 2mbit/sec h264 (like on a classic social media or whatever site) is absolutely terrible compared to going from source -> 2mbit/sec h264.
The problem is when you go from 15 to 2 in my example is that the encoder spends all the time trying to basically encode the artefacts. It really doesn't work well at all.
Welcome to the world of video compression. If you're not able to take the time to learn the ins/outs of how to use a codec as well take each incoming video's specifics into consideration, then you'll be needing to borrow someone's middle of the road presets. Dedicated settings make decisions based on the frame size, bitrate restrictions, things like HLS vs download/play, 1pass/multipass etc. All of that determines GOP size, reference frames, etc.
(Which could include max size over time constraints like VBV.)
It has the ability to do trial compression (w/ scene splitting) and evaluate quality loss up to a desired factor.
[1] https://en.wikipedia.org/wiki/Video_Multimethod_Assessment_F...
[2] https://netflixtechblog.com/toward-a-practical-perceptual-vi...
Surprising to me that it’s so large, I’d expect something like 5% or less account for half of watch time. I wonder how it compares across platforms (YouTube for instance would almost certainly be in the <5% bucket, I’d imagine)
A. what % of all videos generated the majority of overall watch time over the last week?
B. what % of all videos generated the majority of overall watch time over the last year?
If the watch time is growing by a high percentage month-over-month, perhaps the numbers would be similar. But, if overall watch time is flat, then you'd expect A to be much lower than B.
Both when played on my phone or via the web page.
Videos sent by the same people on other services like Signal, Slack or Messages are fine.
In fact, prioritizing latency might be the reason for the garbage quality you're seeing: reduce the bitrate, and people will receive the video more quickly and more reliably. Whether that's a smart decision is another question.
* ISP has poor peering to facebook servers
* ISP has cache/CDN node installed, but that's overloaded
* the content that OP viewed isn't popular so it has to be pulled from origin, which adds another layer of complications
Moral of the story: Video is hard, apart from Youtube, Twitch, Netflix and Amazon Prime I don’t know any service that plays video flowlessly. O.K. maybe some ad networks too.
" The number of Haystack photos written is 12 times the number of photos uploaded since the application scales each image to 4 sizes and saves each size in 3 different locations. "
[1] https://bigdata.devcodenote.com/2015/04/haystack-facebooks-p...Just imagine if all that ingenuity was focused on solving humanity’s problems, instead of sharing conspiracy theories and advertising.
My engineering colleagues that voraciously use and consume Messenger and Instagram don't see a problem either.
There a lot of "bad" companies. Certain advertising, defense, pharma, law, etc. firms are shady or morally bankrupt. It doesn't stop them from finding people to do the work.
This is the new one I found:
removebackground.app
To Collect feedback metrics to judge usage and throw resources to optimize rather than blind optimizing every single video. These techniques are very simple but makes sense to use when you have such a high demand. But I would expect more from facebook.