OBS Studio Lands AV1 and HEVC RTMP Streaming Support
phoronix.com
phoronix.com
I am sad that we are extending RTMP instead of adopting something better. I want things like
* Codec negotiation happening at connect time, so we don't hardcode things
* Multi-track. So users can upload multiple audio tracks, or do Simulcast. I would love a world where users upload the multiple quality levels. Then anyone could run a stream service, not just those that can afford transcodes.
* P2P. Stream your OBS output directly to another user. It is frustrating to see users struggle with setting up RTMP servers/HLS when they want to share their video with just one or two other people. The other use case I really care about is co-streaming.
* Standardized. I would like to see whatever protocol we use be discussed in the IETF. It frustrates me that it is just companies behind closed doors deciding things. If an individual/open source/startup has an idea or use case it is never going to make it into RTMP.
I have been trying to add WebRTC support https://github.com/obsproject/obs-studio/pull/7926 so I am hopeful for the future.
MoQ (Media over QUIC) development is underway at the IETF - with involvement from Twitch/YouTube/Facebook/Cisco/etc. - but will probably still take quite some time before its finalised and can be adopted.
I have a PR open to OBS for it right now https://github.com/obsproject/obs-studio/pull/7926
IPv4 users seem to manage a connection ~70% of the time in my experience. Remember only one peer needs to have a router with a suitable config for everything to work.
* PR https://github.com/obsproject/obs-studio/pull/7192
* RFC https://github.com/obsproject/rfcs/pull/43
I tried to keep the size as small as I could. It is just basic audio+video (no Simulcast/ICE Renegotiation etc...) It has been a tough process, but I am going to keep working on it until I can get consensus.
2. "Unlike some other protocols that only support specific video and audio formats, SRT does not limit you to a specific container or codec, since it is media or content agnostic. SRT operates at the network transport level, acting as a wrapper around your content. This means it can transport any type of codec, resolution or frame rate. This is important because it can future proof workflows by working transparently with MPEG-2, H.264, and HEVC for example."
2. SRT, in practice, is only ever used with MPEG-TS as the container for audio/video data, which does not support AV1.
Which I'm sure is still significant work, but is it much harder to do with SRT?
The "just" is doing a lot of work here. Extending existing RTMP infrastructure is obviously much easier than spinning up an entirely different one. And you'd still have specify a way to use anything but MPEG-TS to send audio/video over SRT, which de facto does not exist (at least not that I can find).
I agree that extending RTMP isn't ideal, but other solutions have their own set of limitations, problems, and challenges. Whether that's SRT, RIST, or WebRTC.
For better and worse, RTMP and dealing with it at scale is well understood.
WebRTC has so much potential, I have big ideas for using it in a number of projects, if I can ever just clone myself a couple of times. Thanks for your work and enthusiasm!
I constantly dream of a much more simple, but also powerful OBS replacement, with separated components for sending inputs to a "compositing" server, that can then rebroadcast via WebRTC. Maybe with customizable templates that could be more powerful than OBS's scenes. With the ability to let viewers customize/change their views further... It's a bit much to type out all here, but the WebRTC-powered future is very exciting.
I believe both AV1 and HEVC/h.265 specify optional Scalable Video Coding, although I don't know if that's a common feature of encoders or if this integration supports it. If the stars align, it's straightforward to choose parts of the stream to throw away to fit into available bandwidth or decode capacity without transcoding. Of course, decoder support is also needed.
> * P2P. Stream your OBS output directly to another user. It is frustrating to see users struggle with setting up RTMP servers/HLS when they want to share their video with just one or two other people. The other use case I really care about is co-streaming.
The nginx RTMP plugin supports converting an RTMP stream into an HLS stream. It's pretty easy to set up. Even if OBS used WebRTC or some other p2p solution, you would still need some type of signaling / connection negotiation setup.
On startup you could be connecting two clients that have completely different sets of codecs. What if you have a low powered device that only has a H264 hardware decoder? What if you have two desktops that prefer AV1. WebRTC lets each side publish what codecs they support, and their priority
> It's pretty easy to set up
Yea it is not bad if you are comfortable with those things. I want to empower users who have no idea what Linux, codecs or systemd is. I want to create something like 12 year old me would have used. I could have had a lot of fun with my friends if we could have done P2P.
> you would still need some type of signaling / connection negotiation setup.
I use QR codes for this.
For public signaling is also cheap enough that it could be provided for free. I think the project name was simplewebrtc? Hosting RTMP/HLS for free is prohibitively expensive though.
You recode on the server.
* I don't want to upload my video from my security camera to a remote host. It has security and privacy implications
* I don't have the bandwidth available
* I don't want to pay the bandwidth and compute costs.
* I don't want to incur the extra latency
I can see your point but it has nothing to do with what platforms like Twitch are for.
I don’t think P2P is the answer to everything, in the same manner servers aren’t the answer to everything.
Even consumer nVidia cards support at least three simultaneous encode streams, which would be perfectly cromulent for a relay-only streaming service (high / medium / phone).
Of course at some number of viewers in a centralized setup it probably does make sense to send just upload the top quality and make separate lower quality streams for small efficiency gains that add up.
Wouldn't it make sense to have separate channels for distinct content streams (e.g game capture, face camera, overlays..) so that you could independently choose/optimize each channel's compression and thus make the whole stream more efficient?
I don't know anything about compression algorithms, but I hypothesize that maybe now the algos would also be able to do a better job since the frames have less abrupt discontinuities (in the viewport perspective).
This would also enable stream viewers's to activate/deactivate channels when needed. For example, when the overlay/webcam/etc occludes the view on layer (channel) below.
In general the compression algorithm will make an intelligent decision on where to spend bits. However you may gain a bit more accuracy here. The better win may be to set priorities (maybe make the scrolling text crisp right away because it doesn't move often anyways) but I am still skeptical that you would see much benefit.
> enable stream viewers's to activate/deactivate channels when needed
On the other hand this would be very cool. I can also see moving away from fixed aspect ratios to moving the bits around the screen. Maybe the webcam can be above the (non-video) chat rather than over top of the video. Or even just mapping the different streams (that I have selected) into my aspect ratio better (like Google Meet and other video conferencing tools do).
If RTMP was in the IETF my comment wouldn't apply anymore.
I'd be happy to improve the one I submitted (more bands, steeper rolloff,etc) but the project does not want such a feature.
I'm not butthurt, I just find it odd but half expected the rejection going in based on what I read.
The Cynic in me wonders if the project is funded by plug-in vendors and this might be a conflict! I lean towards the rabbit hole theory though.
I don't know how complex the code is though, of course.
E: Their response seems to be that users who would find this useful already use a plugin.
But nee users will have to figure that out. That is one of the worst excuses they could have offered.
BTW the code is completely self contained. Nothing to maintain. That's due to a clean API on their part.
Never say never...
Audio is 80% of video.
Most people will watch to a potato-quality video if the person talking can be understood fairly well, but will probably not bother with something where the audio is echo-y, with wind, and sounds like it was recorded over a wet string, even if the image is 4K/HDR.
How advanced the included audio controls is certainly up for debate, but it's not possible to focus on 'just' video.
Every additional feature is another thing to maintain and test. Saying no leaves more time for everything else so there needs to be an exceptional reason to say yes. This is true for all projects, not just open source.
Yes, it's a bit more convenient to do it all in the same program, but the complexity and maintainance expansion has its own costs.
But also, what's the problem with having this as a plugin? Isn't it better to keep the core OBS small, and provide extra functionality via plugins?
https://github.com/obsproject/obs-studio/pull/7961
I guess most frustrating to me was the request to format the code and fix the commit message as if it was on track, only to be rejected for "we don't want this in the first place".
Plugins suck. There are a lot of people who will not bother to go that route. It immediately raises questions about what OS do they work on, do they cost money, how do I install them?
They include compression and noise gate among others. Enough to clean up my microphone input. But not much in the way of tone control.
Serious people WILL explore the plugins route, but one has to go from "What is this thing I just downloaded and how do I use it?" to "Okay, I've got the hang of this, I'm pretty good at it, but the stock tools are holing me back from making something better." I would venture to say that the majority using OBS will never make it to that second phase. For those who do, have a well SEOed blog post and YouTube/TikTok video ready to show them how to take it to 11.
Usually, those things are easy to check, almost a part of just triage where you try to cover lots of ground but just the surface. Takes small amount of time for the maintainer unless there is a lot of ground to cover obviously.
But unless they explicitly signal the change is wanted, I wouldn't (in the future) take that as a signal that the change in general is wanted.
As a tip, before trying to contribute any larger patches to FOSS, create a issue to ask if this is something they'd actually be willing to merge, if you chose to build it. Then you know for sure up front if it's worth spending time on.
> Plugins suck. There are a lot of people who will not bother to go that route. It immediately raises questions about what OS do they work on, do they cost money, how do I install them?
Personally, I appreciate the architecture of "small core - allow extensions" so I'm happy they don't expand OBS in that way.
As well, I'd appreciate it a lot if you made what you did into a plugin, better audio control is wanted by a large set of OBS and I know many do the piping outside of OBS to get a proper equalizer. So if you spend time on making it into an extension, I'm sure you'd find lots of people grateful for that.
Oh I completely agree. This was scratching an itch on my part, plus I had been under the impression they didn't want it. But it wasn't clear, did they not want to develop it themselves? Why the 3-band? So I finally sat down and did something and fit the submission guidelines to find out.
Now I kinda have the bug. It might be best to make some kind of pipewire effects that can be used in anything (on linux).
Pipewire effects for equalizer and so much more already exists (personally I use Easy Effects for that), and the people I know using OBS are kind of stuck with Windows because of other reasons. Also, it's nice to have everything you need in one package.
But, you're the author after all, you decide what you want to do :)
The majority of the time a DAW ships with things like EQs, compressors, etc, they are actually some kind of plugin.
Plugins don't suck, your opinion is not everyone else's fact.
Plugins do suck. But the entire recording industry uses plugins.
Both are true. Loading arbitrary blobs of 3rd party code into your plugin host is a nightmare; adding arbitrary blobs of amazing new functionality into your plugin host is amazing.
That does suck, from multiple people too. Code review processes that put off developers from contributing and don't value their time, rather than helping them to contribute are a pet hate of mine.
Reading that PR discussion, I don't see any sort of 'gross'-ness.
The submitter submits a (presumably unrequested) PR. One project maintainer adds comments to bring it up to project standards. Later, another dev clarifies that after discussion, they agree that it doesn't belong in the core project.
For anyone else out there who doesn't like 'their time wasted', discuss your idea with project maintainers before submitting a PR. Just because you think it is a good idea doesn't mean it fits within their vision for the project.
> For anyone else out there who doesn't like 'their time wasted', discuss your idea with project maintainers before submitting a PR. Just because you think it is a good idea doesn't mean it fits within their vision for the project.
Of course. I give the same advice. But that's not a valid excuse for treating would-be contributors like this. You can reject PRs but in a more respectful and polite manner, without the misleading additional time-wasting requests. Rejection and professionalism/politeness are not mutually exclusive.
clarification request: with scare qutoes around "their time wasted," does that imply that you do like your time wasted? or that you don't think all the effort spent making the PR was wasted?
You know the answer, so forgive me for not responding to your not-in-good-faith question.
To clarify my point: If you think about this from the project maintainer's standpoint, the submitter is the one doing something wrong if they are getting angry at this outcome. They are speculatively doing work assuming that it is going to eventually get integrated, without seeking any information about its value beforehand.
No I don't, otherwise I wouldn't have asked. And if you're going to assume bad faith on my end then there's probably not much more to discuss.
Curious, what if the maintainer had just replied with a "finger" emoji instead? Would that change your opinion on whether this was gross, or is the manner of the reply completely irrelevant to you? (Since you seem to be presuming bad faith on my part, I feel I have to add this disclaimer to my thought experiment. Don't assume I'm implying there is some sort of equivalency here with the response that was given and a hypothetical "finger" response. I am not. This is a thought experiment, not an implication that the maintainer acted in such a way).
If so, then I would suggest that our disagreement here is just about what is polite and what's not, and what a person's social responsibility is regarding being polite. In my opinion there's a responsiblity on the maintainer to be polite in this case. If you disagree with that, that's just an "agree to disagree" sort of thing.
If your opinion wouldn't change if the response had been a "finger" emoji, then I would suggest that our disagreement here is about whether politeness is something that matters.
The person who submitted this code went out on a limb and implemented something without talking to the maintainers first. That's completely fine. I'm sure they learned quite a bit by implementing it. That's a win by itself.
It's the getting upset that it didn't get integrated that I think is unreasonable. The maintainers donate their time to write and curate code to make progress on their vision of the software.
Submitting code like this (not talking beforehand) is like giving a friend who likes photography a camera without asking what they want. There's a chance they love it, but there's a good chance they wont. If you care about getting one outcome over another, you need to communicate beforehand.
The hardline opposition to adding hevc etc support to rtmp came from ffmpeg+gstreamer afaik, and from the document it looks like at least ffmpeg is on board.
...Might have to brush off our broadcast software and add this :)