Unbloating the buffers
dgroshev.com
dgroshev.com
Anecdotally, switching from DOCSIS 3.0 to 3.1 seems to have improved my worst-case latency, because DOCSIS 3.1 includes a CAKE-like algorithm (PIE) in the modem itself:
https://www.cablelabs.com/blog/how-docsis-3-1-reduces-latenc...
The initial bandwidth setting is really to get it started at a good value for your usual bottleneck device, such as your local cable modem.
Always starting off as if you had fast ethernet to the ISP when you actually have a super-el-cheapo link is a waste of time and effort. Even if you have a good algorithm like CAKE.
--dave
Now, to get around that some devices perform regular speed tests and dynamically adjust the high water mark. This said there are some limits at which you should perform these tests as too often and you may affect the actual applications you're trying to run.
With the limit set to 1 Mbps (1000/1000), my upload latency dropped from 80ms to 25ms, but speed was hard-limited to 1000/1000. With the limit rasied to 1G/1G, cake stopped working and my upload latency returned to 80ms.
So I stand by my original comment. You still have to configure the speed limits manually.
But I think davecb may have been confusing traffic shaper limits with TCP congestion control behaviors, and maybe the impact of a large TCP initial window releasing a sudden burst of packets that may be large enough to build up a bit of a queue on particularly slow links. (It was a serious problem in the ADSL era; now, only wireless gets that slow, and large bursts of packets are as likely to help as not when frame aggregation enters the picture.)
Right. This is a scheme for trying to manage the flows from the end you control so that the clueless FIFO nodes out there don't overload.
Flow management should be at the point where multiple flows feed into a bottleneck. Usually at the point where the LAN connects to the outside world, the telco modem or cable box. And, in the other direction, where bulk bandwidth goes into limited bandwidth to the consumer at the upstream end. I'm encouraged to hear that this finally got into DOCSIS. Now if only we could get it into AT&T modems.
It's even worse than that, a lot of consumer Internet connections are bottlenecked by some shared link, which means your max bandwidth literally depends on what the other customers are doing at the moment. There is no "correct limit".
Three homes ago, my Comcast Internet only supported 135mbps before latency spikes while using active shaping, so I downgraded from the 250mbps plan to the 125mbps plan and no longer had to use the shaping device at all. Another home could only get 75mbps, so I switched to a 25mbps plan just to prove that low bandwidth excelled at low latency; it did.
This does mean that console video games download slower, but I grew up on modems, so I can keep myself occupied while my computer is downloading — and instead of needing a shaping router, I saved half my internet bill per month.
In general fq_codel is superior to pie in just about every way. There is a specific thing that pie does that makes it appealing to hardware designers is that it is O(1) on egress and fq_codel is not (but there are ways around that). fq_codel achieves less latency than pie does, and faster than it does. codel takes a bit longer than we would like.
fq-pie still has worse tcp latency (using the same fq algorithm as in fq_codel and cake), and cake solves all kinds of edge cases that no other fq+aqm+shaper can - per host FQ (solving torrent issues), ack-filtering, nat transparency, a better codel model, easy to use link layer compensation, diffserv - and my favorite feature is actually that it runs more or less the same as a default qdisc, as a shaped one. I had hoped it would replace fq_codel at both line rate and at the shaped cases, but it did prove a bit too cpu intensive on cheesy hardware and the edge cases not noticeable enough on simple benchmarks.
I would like to multi-thread it, and add some new features, discussion here: https://docs.google.com/document/d/1tTYBPeaRdCO9AGTGQCpoiuLO...
So do I set my max to 900Mbps all the time? I'd really rather not.
And this is worse for the upload bandwidth, as there's so precious little to spare of it.
Delighted to see vyos pick it up!
Next is the shaping vs variable rate issue, on outbound and inbound. On outbound - Ideally the fq_codel or cake algorithms are located right on the bottleneck link, and respond to changes in the available bandwidth from the upstream dynamically (asserted by ethernet pause frames for example), and managed underneath by BQL (ethernet) or AQL (some wifi(ath10k,ath11k, mt76, mt79)) , though true native fq_codel support exists for the ath9k) which store up enough data only to service one interrupt. These algorithms are much (20x!) lighter weight than shaping to a fixed rate is, and like I said
Codel was designed to deal with variable rate links. A fixed shaping rate was not what we intended, but has become one of the most common and finicky use cases, dang it. It is the FQ that matters most in most cases, however.
On inbound shaping, it's a SWAG, where we recommend 85% or so of the provided rate as a starting point. You might need more, you might need less, due to the behaviors of "slow start" in particular, it is never going to be accurate enough, and we keep encouraging the ISPs merely to manage their own egress well enough (ideally dynamically and without a shaper) so inbound shaping is not needed. Libreqos/preseem/bequant/paraqum all make middleboxes that do ISP level shaping. Libreqos (for whom I work these days) has pushed CAKE to where an ISP can do 10k subscribers at 25Gbit at 50% of cpu on a cheap ($2k) Xeon or Ryzen box.
So you get something that works well on the rrul test... in the morning but not the evening. Folk then complain that their bandwidth from their provider is variable rate and shaping via their other box does not work all the time, and I go back to hoping more ISPs get told how fix it at, egress, truly right at their end.
There is another project which perhaps can be ported to vyos, called cake-autorate, which attempts to use active measurements to cope with variable rates.
Kathie nichols has also spoken about Codel multiple times.
OpenWRT, for example, can be flashed on many older routers and will be able to solve bufferbloat without replacing the hardware.
Software such as network schedulers in Linux can work around that problem. But ultimately, it should be fixed in the various networking devices that create the problem in the first place.
"bottleneck link" is, of the path from the packet source to the target, whichever link is trying to exceed the available bandwidth. The problematic router (prior to that link) will receive packets faster than it can send them, so it needs to either use the queue, drop the packet, or use the ECN bits in the IP header to tell the target to tell the source to slow down. A router with bad queue management will enqueue packets until the queue is full, before picking one of the other options. This adds latency depending on the queue size; big queues on slow links cause significant latency.
With bad internet connections, the bottleneck link is often the connection between the home router/modem and the ISP's infrastructure. If upload bandwidth is fully used, bad queue management on the home router will cause latency to spike. If download bandwidth is fully used, bad queue management on the ISP's router will cause latency to spike.
With fast internet connections, the bottleneck link might be somewhere else, e.g. at the interface between two ISPs.
Because the problem only happens at the bottleneck link, it is possible to work around a bufferbloat problem in the ISP's infrastructure by introducing another artificial bottleneck in your own home network. So by using a bit of Linux software that artificially limits bandwidth, that software creates a virtual "bottleneck link" where it can control both ends, and thus "fix" the problem (as the link affected by bufferbloat will no longer be overloaded, and thus no longer use its queue). A real fix would require updating the firmware/configuration on the ISP's routers.
If you have an older, slower router in your network that is not acting as your gateway router and does not need to compensate for bufferbloat in somebody else's device (eg. you're using it as a secondary WiFi access point), then you should almost always use fq_codel on that device's network interfaces.
Some devices like some cable modems have even more drastically anemic CPUs because the CPU was never intended to be part of the data plane and most traffic is supposed to be offloaded to special-purpose hardware. Those may be genuinely unable to even handle CoDel, and are a big part of the reason why DOCSIS 3.1 invented PIE AQM rather than adopting CoDel as the standard AQM.
What am I missing? Are things like cake/codel "easier" to deploy than creating a single paired qdisc and class?
Most people do not measure in-stream latency, rather multiple streams, so, assuming you did or did not pick up fq_codel, a packet capture will show you the tcp RTTs being managed or not. Note that on a system where the tcp is originating and then transitioning htb on the same box, another subsystem (TSQ) managed the rates fairly well before hitting HTB.
Prior HN discussion: https://news.ycombinator.com/item?id=38597744
My biggest gripe about that project is that everyone focuses on the ECN change, and nobody looks at the effects on normal everyday traffic, where fq_codel or cake utterly smoke the dual-pi AQM. There actually is L4S support in fq_codel now, but they are not testing it. :(