Ultra Low Latency WebRTC Streaming – Open-Source Media Server
antmedia.io
antmedia.io
In answer to the questions about TURN, this approach won't be lower latency than TURN but will scale better. TURN servers are just (nearly) passive relays, so the sending client needs to set up as many outbound streams as there are receiving clients in the session.
The advantage to the TURN approach is that you can do end-to-end encryption.
The server this post is about "forwards" streams. It's in a class of infrastructure traditionally called a "Selective Forwarding Unit." So the sending client can send one stream, and the SFU copies the stream and sends it to the receiving clients. The server needs more CPU and outgoing bandwidth as the receiver count grows, but the sending client doesn't need much of either.
Once you have streams piping through a server, you can do lots of other things, of course. This server can transcode to several formats, including HLS and RTMP, and can copy streams internally across a cluster to scale more than a single machine can.
Media servers are really fun to work on: lots of small(-ish) hard problems that touch low-level network protocols, memory and cpu optimization, architecture (because eventually you want a clean plugins interface, etc.).
If you're interested in this stuff, check out Mediasoup, an open source WebRTC SFU with a very nice design in which all the low-level stuff is c++ and all the high-level interfaces are exposed as nodejs objects.[0]
Therefore i am waiting for the PERC standard. It will allow to send the audio/video to server only once and the server will retransmit to all peers - everything still end-to-end encrpted.
"Easy to use" is in the eye of the beholder. Mediasoup isn't a packaged solution; you need to write some javascript to get a server up and running. From this perspective it's a framework that sits at the same level as, say, Express. But the API is flexible and powerful. It's exactly what you want if you're building an SFU cluster for a specific use case and want to decide on the architecture yourself. Here's excellent "1:N broadcast" sample code: [0]
With that sample code, you can broadcast a single WebRTC stream to ~500 clients on one vCPU of an AWS c5.large instance. To scale to more clients you need to use more than one vCPU. The new mediasoup v3 APIs make this pretty easy. There are now "pipe transports" and "plain rtp transports" for moving the low-level rtp bits around between processes and between machines.[1]
[0]-https://github.com/michaelfig/mediasoup-broadcast-example [1]-https://mediasoup.org/documentation/v3/differences-between-v...
At the moment, everyone in the industry is working to figure out how to provide a viewer experience that is similar, or better than the current expectations of live streaming.
Live streaming is hard. Ultra-low-latency is just as hard.
However, I'm still not sure why we need ultra-low-latency. Will it improve the live experience? Maybe. Aside from live-action sports or significant events, do we as content consumers really need ultra-low-latency?
Twitch chat is fun.
Simply for passive viewing probably not. For live streaming in which the viewer is also using actuators, e.g., gamepads, wheel, etc., for controlling the stream capture device, then yes.
1. http://www.wired.co.uk/article/england-vs-croatia-live-strea...
I subscribe to both Formula 1 and MotoGP online streaming services and couldn't care less if the broadcasts were even multiple minutes delayed. They may very well be already and I wouldn't know. As far as I can tell they both just use standard CDNs delivering DASH or similar (basically just short video files played in sequence).
That seems more than enough for sports. Multiple minutes may result in getting real time information from elsewhere (e.g., Twitter, local radio) but anywhere from 10 to 30 seconds seems fine.
We call it the "Twitter effect". Essentially people either see things on social media or app notifications before it happens. Kinda ruins it for some people.
We are working on sub-30 second latency in 4K w/ server side ad insertion, but it's hard work. Especially with encoding with DRM baked in.
For the FA Cup final I had two feeds, one the 4K BBC version, and one on SDI from Wembley. The goal went in, and I'd completely forgotten it by the time the 4K goal went in.
The trouble is, things like twitter, whatsapp etc, will work to the lowest latency, and for a football game that's about 50-600 nanoseconds to the earliest viewer. Twitter can push the message within a few seconds (or the VAR result or whatever)
Latency needs to be in the sub-10 second goal-glass range.
If you delay all your friends by 30s +-2s then that would probably be fine...
That changes the problem somewhat.
If you're not watching with someone else, it doesn't make that much difference; but 30 seconds is plenty for someone to cheer for a goal and the notification to distract you from seeing it.
So given that what you are competing against (cable TV) already has a significant delay I doubt a normal web stream can't reach the same or better latency with just standard HTTP download of video files.
In Santiago, during national soccer matches everyone would watch, goals would reach us through the sound of neighbors rejoicing sooner than they would from actual television coverage because the satellite dish had a latency of say 3 seconds over cable.
And if the problem was actually that serious it would be easier to fix it by just adding additional delay to the faster technologies instead of having to reduce latency in internet streaming at huge cost.
Results in sports can switch from one side to the other end within split second, so whoever has the lowest latency will gain massive advantages. It's very similar to stock trading.
It’s the entire reason that I won’t cut the cord.
For 1->n streaming you just want very high quality broadcast->server and then presumably you want the downstream clients to not have the option to send packets back to the broadcaster. Transcoding is hard though, so consumers might be stuck with just the source video quality. I assume this lib just passes the frames on from the source (twitch did just that through almost all of its pre-acquisition growth).
Some of my details might be a little off, but I did spend a few months trying to build something very similar to this with gstreamer. WebRTC was hard to grok.
I'd say ultra low latency would be sub-frame, so <40ms. Low latency on the 1 frame to half a second, normal upto say 1.5s, and high latency beyond that.
But that's for point to point or multicast circuits, and ISPs don't like multicast as they can't charge for it.
>Latency To Broadcaster: 7.71 sec.
>Latency Mode: Low Latency
Edit: It's the same author
A few seconds would also such for live streaming where the streamer interacts in chat.
in fact, for scaling a solution that uses "near realtime broadcast", you probably don't want encryption at all. ( afaik you can't disable webrtc encryption. )
in my tests i can get 3x more performance just using rtmp/rtsp. ( ~8000 video streaming connections on a aws c5.xlarge at 800kbit/s ).
WebRTC can spawn large number of p2p connections.