Video streaming at scale with Kubernetes and RabbitMQ
alexandreolive.medium.com
alexandreolive.medium.com
Granted, I knew that my service needs to be able to scale at different bottlenecks, but a lot of "build your own video service!" tutorials start with:
- Build a backend, return a video file
- Build a frontend, embed the video
And that leaves a lot to be desired in terms of performance. I think the actual steps should be:
- Build a backend that consists of:
- Video Ingestion service
- Video Upload / Processing Service that saves the video into chunks
- Build a streaming service that returns video chunks
- Build a frontend that consists of: - Build or use a video streaming library that can play video chunks as a stream
Edit: From the author's links, I found this website which is very informative: https://howvideo.works/What does it do?
Mostly just building for fun though.
This is both a semi-shameless plug and probably a few levels deeper than what you're looking for, but I organize a conference for video developers called Demuxed. The YouTube channel[1] has 8 years worth of conference videos about streaming video (and the 9th year is happening in a couple of weeks). The bullet points you mentioned are definitely covered across a few talks, but it's certainly not in any kind of "how to" format.
At this point you also need to chose what streaming protocol you want to use. You have mostly two choices, HLS if you want to get things done quick, or MPEG-DASH if you want more control (but you'd need a separate HLS pipeline for iOS anyway…)
> Build or use a video streaming library that can play video chunks as a stream
As someone who's worked on a web streaming player, I'd strongly recommend not to build one but to use an existing one (or, in short: use HLS.js)
Can you explain exactly what you mean?
What exactly does this mean?
Throw subtitles in multiple languages, and different audio tracks, into the mix, and all of a sudden streaming video becomes a nightmare.
Finally, if you are dealing with copyrighted materials, you have to be aware as to what country your user is physically residing in while accessing the videos, as you likely don't have a license to stream all your videos in every country all at once.
Throw this all into a blender and what is needed is a very fancy asset catalog management system, and that part right there ends up being annoyingly complicated.
My very first job was writing streaming media services for radio. Integrating the ad services over events embedded into the stream identifying content was a pretty nifty solution. You send the metadata to your ad server, it selects the ads based on your user and media metadata, you swap out the player's media url to play the ad(s), and keep the stream going in the background. When the commercial break is up, you kill the ad stream and swap the media stream player back in. Presto changeo, seamless ads in 2007.
You also track media playback through the stream embedded events.
The hard part back then was live video/audio encoding. These cards lived at the stations, and streamed to our cdn. Geocoding for American service bases in foreign countries was interesting too, the streaming rights for stuff limits the countries you can stream in.
Monitoring the live encoders was an interesting problem - I setup an event-based solution so I could capture the audio waveform and analyze the frequency buckets on dozens of streams at once. Could tell if a stream was silent pretty quickly, and alert the station engineer via email.
Back then, a lot of this stuff wasn't off the shelf like it was today. Even on the frontend,working with Flash, window's media player sdk, and the QuickTime sdk and integrating it with Javascript was challenging. And the backends were all in php and Java 1.4, and mysql, and a whole lot of memcache.
Was a really great experience for a new professional programmer to write all that.
I was working on streaming in the late 90s - that was real wild west, had to write solutions myself, and it felt very duck taped together. Luckily it was an offshoot of a largish ISP so we had bandwidth and a lot of servers. incoming streams often used bonded ISDN lines.
But the devil is in the details and most people that attempt it don’t even get the database models right on the first or second time, they underestimate the complexity particularly when it comes to video and struggle when the use cases widen. Document asset management systems are way easier but usually don’t understand media as in depth as a media asset management system.
The system differ in that it was not user generated video content. It was coming from the cameras in our fitness studio.
Here is the article if anyone intereste to read about: https://dev.to/dvliman/building-a-live-streaming-app-in-cloj...
I'm not actually sure on balance how much transcode gets done in hardware vs software, since it's also very amenable to using batch compute that's otherwise idle. I'll guess that most or all live transcoding - streams, on-the-fly transcode into formats not pregenerated - are done in hardware, and transcoding new formats for the back catalog are probably done on a mixture of mechanisms where and when capacity is available. (Source: Googler, not on YouTube though.)
AWS is easier, but you can do it with anything. The basic steps are:
1. Upload the file somewhere 2. Transcode it 3. Put the parts somewhere 4. Serve the parts
You should really transcode everything into HLS. It's 2023, and everything that matters supports it. If you want 4k you can use HLS or the other thing (which I keep forgetting the acronym for).
If you want to get fancy you can do rendition audio, which not everything supports. Rendition audio means sharing one audio stream amongst N number of video streams.
You can use FFMPEG to transcode, but I'd suggest using AWS MediaConvert. It's cheap, fast, and probably does everything you want. Using FFmpeg directly works, but why bother. You will get an option wrong and screw everything up. You don't want your video to not work on some random device that 50k people are using in some country you didn't think about.
He's using RabbitMQ but you should use SQS, because SQS can trigger lambdas...which means no polling required. But use whatever queue you want.
You can kick the process off by attaching a Lambda to S3, which will start the process when the file is uploaded.
You can kick your "availability activation" off by attaching a Lambda to the S3 output bucket.
Background: I help run a streaming service and built the backend pipeline.
This omits the entire "metadata management and analytics" side as well. That's left as an exercise for the user.
Cloud services like S3 and Azure Storage were invented specifically for hosting images and video. That’s their origin story, their foundation, their very reason for being.
Similarly, cloud functions / lambda were invented for background processing of blobs. The first demos were always of resizing images!
Building out this infrastructure yourself is a little insane. Unless you’re Netflix, don’t bother. Just dump your videos into blobs.
It’s like driving to your cousin’s place, but step one is building your own highway because you couldn’t be bothered to check the map to see if one already existed.
PS: Netflix serves video from BSD running directly on bare metal because at that scale efficiency matters! If efficiency doesn’t matter that much, use blobs. Kubernetes is going to be even worse.
Ultimately you’re competing with the likes of Twitch (IVS), youtube and cloudflare on price and they ALL run their own compute so at certain size you will have to run your own hardware to stay competitive especially now that zirp is in rear view mirror.
See Dropbox as another example of this but in documents/storage space
For reference Google Transcoder is about $0.13 per minute of encoding (at four resolutions). Mux is at $0.032 / min, AWS Media Convert $0.0188.
I should note I know Mux's pricing well, use them a lot, happily. It gets a bit confusing with Google and Media Convert because I'm not sure how these costs map to the resulting bitrate renditions that get created and I've not got the time for a deeper dive to get a more straight apples to apples comparison (ignore scale discounts)
[1] https://cloud.google.com/transcoder/pricing
You do that by using CloudFlare, fastly, or another CDN.
You can get bandwidth costs down to < .001/gb by committing to $1500/mo in bandwidth. The CDNs will pull from S3 once, then cache it forever (assuming you do it right).
Do you really want to spend your time messing around with ffmpeg and setting the correct GOP values? Or trying to create your own b-frame track?
At some point you actually need to do what all these settings are for. Until that time you should use other services.
If you’re interested in working on transcoding I’d highly recommend taking a look at Shaka-packager/streamer.
Also looks pretty complex.
The stabilization step presumably does a video encode …. that’s extremely expensive in terms of time, compute and money I wonder why it’s necessary.
Hetzner or other bare metal providers would probably be a better idea.
AMD Alveo MA35D Media Accelerator
https://www.xilinx.com/applications/data-center/video-imagin...
It uses un-utilized infrastructure around the world and incentivizes independent network operators to join the network (kind of like a 2-sided marketplace for video-specific compute).
Please sign up for a free account and check it out! We'd love to get your feedback
Also personally I wouldn’t use rabbitmq … it’s pretty heavyweight… there’s lots of lightweight queues out there. Overall this architecture looks like it could be simplified.
Also, the post doesn’t mention if the video encoding uses GPU hardware acceleration. Makes a big difference especially if using spot instances …. ffmpeg in CPU is extremely computationally expensive.
Presumably all input videos need reencoding to convert them to HLS.
I run https://atomictessellator.com solo, using kubernetes, and my database, Minio object store, application servers, quantum workers, everything is all on kubernetes, it’s self healing and much simpler to run all the infrastructure the same.
Recently I had a node failure while I was sleeping and the whole system healed itself while I slept, the monitoring system didn’t even alarm me because the small blip of increased latency while the pods rebalanced wasn’t above the alert threshold so it didn’t even wake me up.
What happens in the article infra when the rabbitmq or database nodes fail? The whole system goes offline, which seems very silly setup when you have kubernetes sitting right there, who’s primary function is to handle all of this.
I’m not claiming that it’s totally bullet proof, I never said that - I’m saying that if you had a kubernetes cluster anyway why not benefit from its abilities? Especially when the alternative is single node, single points of failure, which is clearly inferior.
The "what if the storage detaches" argument could easily apply to the single node VMs too, in which case the outcome would be a total system failure.
We are discussing the contrast between the articles architecture and running everything on K8s ... and I'm saying that running everything on K8s is clearly better
> What happens in the article infra when the rabbitmq or database nodes fail?
I agree; I was replying to this invented problem:
> What happens in the article infra when the rabbitmq or database nodes fail?
It makes sense if you read the reply as a reply.
We run all of our stateful and stateless workloads on 10+ kubernetes clusters at work in multiple datacenters in multiple continents, and we serve 500 million users a month with it.
I wrote the first BORG version of DFP backend systems at Google, where we served billions of users billions of ads a day, and we used stateful infrastructure management on some of the first container runtime systems that inspired k8s during it's development.
Using rabbit and "most databases" native fallover strategy is fine for toy projects, but when you're operating at this scale, you need automated infrastructure provisioning and all of the automated tooling around it.
Then you get a bit worried that the Postgres Helm chart, while good, doesn't do what you want. So you update to use a dedicated clustered Postgres, using some Postgres clustering tech.
Finally, you're at so much scale you can throw giant wads of advertising cash at the problem, and you can use anything you like and it'll work. You just need to choose the best thing for your particular problem.
Is that still true?
I wouldn't call the parent comment charitable enough, because there definitely can be some reasons for running stateful workloads even outside of containers altogether (familiarity included), but at the same time it feels like a lot of effort has been invested into making that a non-issue.
For example, how many database Operators are now available for Kubernetes: https://operatorhub.io/?category=Database&capabilityLevel=%5...
Honestly, as long as you have storage and config setup correctly, it's not like you even need an Operator, that's for more advanced setups. I've been running databases in containers (even without Kubernetes) for years, haven't had that many issues at small/medium scale.
We are using RabbitMQ because it's my company target solution. There might better so lighter solution that would fit us but having just one for every solution is easier to maintain.
Great comment about GPU hardware acceleration for encoding, I'm going to look this up.
That's pretty important context.
That said, I wouldn't call RabbitMQ that heavyweight myself, at least when compared to something like Apache Kafka.
Gumlet (https://www.gumlet.com): Per-title encoding (Netflix's approach) to optimize and transcode your videos to boost engagement rates. Moreover, securing your videos is easy with digital rights management solutions paired with Widevine and Fairplay. Made for developers, by developers.
Mux: Developer-friendly video infrastructure for your on-demand & live video needs.
I love Gumlet because of their pricing and support.