End-to-end congestion control cannot avoid latency spikes (2022)
blog.apnic.net
blog.apnic.net
I'd bet you could do something similar with Meet and Zoom–my understanding is video bitrates for those services are lower than for e.g. Netflix which we showed are much lower than network capacities. But it might be tricky because of the latency-sensitivity angle, and we did not look into it in our paper.
[1] https://www.sandvine.com/hubfs/Sandvine_Redesign_2019/Downlo...
For a given quality, bitrate will generally be higher in RTC apps (though quality may be lower in general depending on the context and network conditions obviously) because of tradeoffs between encoding latency and efficiency. However, RTC apps generally already try to underutilize links because queuing is bad for latency and latency matters a lot for the RTC case.
op used the term presumably to describe "live content" eg. the source material is not available as a whole (because the recording is not finished); which can be considered a subset of "streaming video"
the sensitivity in regard to transport characteristics stems from the fact that "live content" places an upper bound for the time required for processing and transferring the content-bits to the clients (for it to be considered "live").
It's common knowledge to run your API nodes at 20%-60% CPU load, no more, exactly to curb tail latency.
You can deploy a 24-fiber optical cable and allow many thousand virtual circuits to run on it in parallel using packet switching. Usually orders of magnitude more when they share bandwidth opportunistically, because the streams of packets are not constant intensity.
Running thousands of separate fibers / wires would be much more expensive, and having thousands of narrow-band splitters / transcievers, also massively expensive. Phone networks have tried that all, and gladly jumped off the physical circuits ship as soon as they could.
Both are technologies that one rarely has any contact with as an end user.
Sources: Things I've read a long time ago + Wikipedia for ATM, Wikipedia for MPLS.
This was undoubtedly true (and not even close) 20 years ago. As technology changes, it can be worth revisiting some of these axioms to see if they still hold. Since virtual circuits require smart switches for the entire shared path, there are literal network effects making it hard to adopt.
was slightly at a loss in what exactly needed to be shown here until i clicked the link and came to the conclusion that you re-invented(?) pacing.
that is certainly an interesting economical optimization problem to reason about, though imho somewhat beyond technical merit, as simply letting the client choose quality and sending the data full speed works well enough.
addition:
i totally agree that things have to look economical in order to work and that there are technical edge-cases that need to be handled for good ux, but i dont't quite see how client-side buffer occupancy in the seconds range is in the users interest.
Bufferbloat > Solutions and mitigations: https://en.wikipedia.org/wiki/Bufferbloat#Solutions_and_miti...
[0] - https://www.cs.princeton.edu/courses/archive/fall17/cos561/p...
1. Seeing the future 2. Building a ten times higher capacity network 3. Breaking Net neutrality by deprioritizing traffic that someone deems “not latency sensitive” 4. Flood the network with more redundant packets
Or is the whole text a joke going over my head?
The article ends with “ I will leave you with a question: Are we trying to make end-to-end congestion control work for cases where it can’t possibly work?”
So, it seems to me that there may not be any good solutions to latency spikes. The article is basically pointing out that you either pursue one of the unfortunate solutions mentioned, or be resigned to accept that no congestion control mechanism will ever be sufficient to eliminate the spikes. This seems a valuable message to the people who might be involved in developing the sort congestion control mechanisms they’re talking about.
Seeing the future obviously can't happen. Building a higher capacity network is just wasted money. Breaking NN is going to be unpopular, not to mention determination of "not latency sensitive" is going to be difficult to impossible unless there's a "not latency sensitive" flag on TCP packets that people actually use in good faith. And flooding the network with more redundant packets is just going to be a colossal waste of bandwidth and could easily make congestion issues worse.
It is theoretically plausible that end-users could mark their packets 'latency/jitter sensitive' or 'high throughput'. If there was an Internet-wide consensus that those two options made useful trade-offs then there would be an incentive to use it correctly. That would be as NN-safe as any other scheme.
For example, maybe 'latency/jitter sensitive' buffers less, but at a 25+% throughput penalty. 'High throughput' would then be faster, but have much higher latency/jitter.
The trick would be ensuring fairness without requiring too much hardware at scale.
but treating the "importance" of flow inversely to their bandwidth usage works astonishingly well (flow-queuing/fair-queuing)
both sides claim fairness/net-neutrality and performance on their side but one is deployed and working while the other one is standardized and looking for an application.
presumably there is room for a compromise but defining "queue building flow" w/o regard to link capacity seems like a step in the wrong direction
Is this problem almost exclusively to do with the "last-mile" leg of the connection to the user? (Or the two legs, in the case of peer-to-peer video chats.) I would expect any content provider or Internet backbone connection to be much more stable (and generally over-provisioned too). In particular, there may be occasional routing changes but a given link should be a fiber line with fixed total capacity. Or is having changes in what other users are contending for that capacity effectively the same problem?
* We absolutely would spend the money to avoid globally hitting that number.
* We'd use ToS so that user-facing TCP traffic wouldn't get dropped, but instead some less latency-sensitive and more loss-resistant transfer protocol traffic would instead.
* We'd have several levels/types of load balancing from DNS-based as they come into the network (typically we'd direct to the closest relevant datacenter but less so as it gets overloaded) to route advertising to Maglev<->GFE balancing to GFE<->application balancing and so on.
* etc.
I would expect that'd be true to some extent for any content provider. There are surely some problematic hops in the network (I've seen alleged leaked graphs of Comcast backbone traffic flattening out at 100% at the same time every day) but the entire network is oversubscribed...and running out of capacity regularly in practice? No way.
that's a definition of oversubscription.
> around 70-95% utilization depending on link size.
2/3's with dumb queues, <100% with the computational tradeoff of sqm
For their upstream links to actually be maxed out (thus "experiencing congestion" as Hikikomori put it) with any regularity is more remarkable—suggests they screwed up their capacity planning or just don't care. I kind of expect that from Comcast but not ISPs in general.
For those links to be of varying capacity (like the 5G/Wifi networks the article mentions) would be truly surprising to me.
I said this:
> the entire network is oversubscribed...and running out of capacity regularly in practice? No way.
and you took that to mean "the network is not oversubscribed"? and this is what you're focusing on? No, the "...and" was the important part. Forget the word oversubscribed. It's a word you introduced to the conversation, and it's a distraction. I don't care about the theoretical potential for congestion; I care about where congestion mostly happens in practice.
https://www.theverge.com/23655762/l4s-internet-apple-comcast...
> Congestion signalling methods cannot work around this problem either, so our analysis is also valid for Explicit Congestion Notification methods such as Low Latency Low Loss Scalable Throughput (L4S).