I immediately jumped ship to WebTransport when Chrome added support. But I suppose there's no other option if you need P2P support in the browser.
392 karma · joined February 15, 2023
I immediately jumped ship to WebTransport when Chrome added support. But I suppose there's no other option if you need P2P support in the browser.
Cloud services are pretty TCP/HTTP centric which can be annoying. Any provider that gives you UDP support can be used with QUIC, but you're in charge of certificates and load balancing.
QUIC is client->server so NATs are not a problem; 1 RTT to establish a connection. Iroh is an attempt at P2P QUIC using similar techniques to WebRTC but I don't think browser support will be a thing.
As for the media pipeline, there's no latency on the transmission side and the receiver can choose the latency. You literally have to build your own jitter buffer and choose when to render individual frames.
And for context, Cloudflare is using a fork of my open source library (kixelated/moq) that I've been working on for a few years, plus I authored the original MoQ drafts. I know my post might look derivative at first glance... but it's the other way around.
Logs should be bursty, because they're most useful when debugging rare issues. If you have identical log lines, then that should have been a metric instead.
Metrics should be sampled based on frequency, because they deduplicate. I'm a huge fan of logarithmically sampling metrics.
AT&T customers are less likely to complain if AT&T throttles a HLS broadcast, as the quality will just be lower.
> A client can theoretically detect a bandwidth fall (or even guess it) while loading a segment, abort its request (which may close the TCP socket, event that then may be processed server-side, or not), and directly switch to a 360p segment instead (or even a lower quality). In any case, you don't "need to" wait for a request to finish before starting another.
HESP works like that as far as I understand. The problem is that dialing a new TCP/TLS connection is expensive and has an initial congestion control window (slow-start). You would need to have a second connection warmed and ready to go, which is something you can do in the browser as HTTP abstracts away connections.
HTTP/3 gives you the ability to cancel requests without this penalty though, so you could utilize it if you can detect the HTTP version. Canceling HTTP/1 requests especially during congestion will never work through.
Oh and predicting congestion is virtually impossible, ESPECIALLY on the receiver and in application space. The server also has incentive to keep the TCP socket full to maximize throughput and minimize context switching.
> From this, I'm under the impression that this article only represents the point of view of applications where latency is the most important aspect by far, like twitch I suppose, but I found that this is not a generality for companies relying on live media.
Yeah, I probably should have went into more detail but MoQ also uses a configurable buffer size. Basically media is delivered based on importance, and if a frame is not delivered in X seconds then the player skips over it. You can make X quite large or quite small depending on your preferences, without altering the server behavior.
> But perhaps another solution here may be to update DASH/HLS or exploit some of its features in some ways to reduce that issue. As you wrote about giving more control to the server, both standards do not seem totally against making the server-side more in-control in some specific cases, especially lately with features like content-steering.
A server side bandwidth estimate absolutely helps. My implementation at Twitch went a step further and used server-side ABR to great effect.
Ultimately, the sender sets the maximum number of bytes allowed in flight (ex. BBR). By also making the receiver independently determine that limit, you can only end up with a sub-optimal split brain decision. The tricky part is finding the right balance between smart client and smart server.
As for maintaining existing standards, it depends on your goals. If you're not trying to push boundaries then absolutely, HLS/DASH is great for quickly getting off the ground. But if you're looking for something in-between HLS and WebRTC then MoQ is compelling.
It's really difficult to compare the latency of different protocols because it depends on the network conditions.
If you assume flawless connectivity, then real-time latency is trivial to achieve. Pipe frames over TCP like RTMP and bam, you've done it. It's almost meaningless to compare the best-case latency.
The important part is determining how a protocol behaves during congestion. LL-HLS doesn't do great in that regard; frankly it will perform worse than RTMP if that's our yardstick because of head-of-line blocking, large fragments, and the playlist in the hot path. Twitch uses a fork of HLS called LHLS which should have lower latency, but we were still seeing 3-5s in some parts of the world.
But yeah, P90 matters more than P10 when it comes to latency. One late frame ruins the broth. A real-time protocol needs a plan to avoid queues at all costs and that's just difficult with TCP.
I'm doing that in my implementation: the main thread immediately transfers each incoming QUIC stream to a WebWorker, which then reads/decodes the container/codec and renders via OffscreenCanvas.
I didn't realize that DataChannels were main thread only. That's good to know!
It's just a lot of work to get everything right. It's kind of working, but I removed synchronization because the signaling between the WebWorker and AudioWorklet got too convoluted. It all makes sense; I just wish there was an easier way to emit audio.
While you're here, how difficult would it be to implement echo cancellation? The current demo is uni-directional but we'll need to make it bi-directional for conferencing.
I'm still just trying to get A/V sync working properly because WebAudio makes things annoying. WebCodecs itself is great; I love the simplicity.
The section is about data channels, which uses SCTP and is ACK-based. Yes, you can use RTP with NACK and/or FEC with the media stack, but not with the data stack.