What Can I Do About Bufferbloat?
bufferbloat.net
bufferbloat.net
The cake traffic shaper in OpenWRT is amazing for fighting bufferbloat in your home network and it can also do almost perfect fairness in dividing the available bandwidth per LAN host with very little configuration. Just get it as part of the SQM tools in OpenWRT and enable it. For the per-host-fairness take a look at the "Make cake sing and dance" from this link: https://openwrt.org/docs/guide-user/network/traffic-shaping/...
If you use an Edgerouter, you can get the cake traffic shaper but you'll have to do without the easy web interface OpenWRT has: https://community.ubnt.com/t5/EdgeRouter/Cake-and-FQ-PIE-com...
As Arie mentions, this is a little more involved on the EdgeOS stuff, but doesn't look too complex for those that are used to a CLI or two.
Edit: Entirely off-topic, but an equally interesting find – There is a WireGuard client for the EdgeRouter line: https://community.ubnt.com/t5/EdgeRouter/Release-WireGuard-f...
If you were happy with your setup before hearing about the EdgeRouter, I hope you'll still be happy with it now :).
I'd hope I'd have upgraded my LAN to at least 2.5 gigabit by then, anyway.
I purchased them a Netgear R7800 and installed hnyman's LEDE build [1] to enable SQM. Night and day difference in latency response. No more staring at a white screen for 3 seconds per URL click.
The build has been stable for several months. I wouldn't recommend this for non-technical users or anyone not willing to spend time troubleshooting, but it has been a great improvement. I couldn't find any other device capable of doing this without running x86 hardware or something else silly.
A few other people mention it, but yes, this is only going to work on slower connections on current SOHO hardware. I think the R7800 can do software SQM at up to 150mbps or so. Plus, if you have a gigabit symmetric connection, hopefully you aren't having bufferbloat issues.
Just wish a popular manufacturer would release an easy-to-use router with SQM so I could install it for non-technical users and forget. Ubiquiti is somewhat close to that, but I believe their prosumer hardware (USG) is running a slow processor at the moment and doesn't even support SQM without installing custom kernels.
[1] https://forum.lede-project.org/t/build-for-netgear-r7800/316
Was the internet speed so high that you couldn't use a normal supported router like a TP-Link Archer C7 at half the cost? I need to do more testing but it seems my C7 can handle my 100/100Mbps fiber connection doing SQM without too much issue.
Edit: Maybe you mean Wifi to wired may have a disabled offload? That path does go through the router and not directly through the switch. For bigger installations I end up having one of these with wifi disabled as the router (firewall, dhcp, etc) and individual ones connected through ethernet as dumb access points (same SSID on all and straight bridge from Wifi to Ethernet). That should also avoid any issues and is a good setup to get more wifi coverage with a simple config.
These days I just run OpenWRT on x86, no more will my router sit in a broken state that I can't fix by logging in over the LAN or WAN (via OpenVPN ofc). Wish PFSense would get sane defaults in this regard!
[2] https://gist.github.com/gonzopancho/760ab9ecee9dfbc1b6033e48...
To answer your question: I have no idea. Would be neat if a much cheaper model had the horsepower though.
I hear hardware offload is possible, but I have yet to try a build that has the patches for it.
One random question: Does it ever make sense to add SQM to a tap_soft interface? I have two locations and both have SQM set up to minimize bufferbloat, but when I VPN from one to the other there is some bufferbloat on the VPN connection.
http://blog.cerowrt.org/post/net_neutrality_customers/
sch_cake (derived from fq_codel with an integral shaper) is about to enter the Linux mainline after years of baking in openwrt.
https://lwn.net/SubscriberLink/758353/05c20f25115c852d/
We can end bufferbloat, now, on everything.
I imagine the main thing that's holding back ISPs from using it on their consumer routers is the (minimal) configuration the user has to do to set to upload bandwidth limit. If ISP routers were also using decent queuing algorithms and not buffering like crazy then that wouldn't be necessary of course, but it doesn't seem like that's going to change any time soon.
Funnily enough, I've just got a firmware applied to my cable connection with Virgin Media in the UK that mitigates the Puma 6 issue on their provided DOCSIS 3 modem - it's eliminated buffer bloat as well, which is a nice side effect.
People are listening, Docsis 3.1 includes active queue management for instance, which goes a long way to preventing the issue.
This thread covers the issue pretty well: https://www.dslreports.com/forum/r31614833-Equip-Intel-Puma-...
I've run a lot of tests since the update, and it's definitely on a par now with my old FTTC connection, jiter especially is much improved.
We have been able to for years. htb+sfq has been around forever.
For directly managing bufferbloat, I never understood this to be a requirement. Either technique is meant to create back-pressure on the TCP stack so it can more effectively manage it's window.. and both do that perfectly well.
My understanding is that you would want AQM in order to keep your bandwidth utilization closer to the actual wire speed than a simple priority queue would.
So, we've had the tools to deal with bufferbloat for a long time, it's just that this mode now just provides slightly better peak performance.
If you have a traffic shaper being fed by a deep, dumb queue, you aren't giving useful back-pressure to TCP until it's too late. You need an AQM that gives either ECN marks or packet drops soon enough that you don't build up or sustain deep queues of packets. That's the core of what bufferbloat is. The queue(s) in front of HTB in your router might not be as stupidly oversized as the queue in your cable or DSL modem, but they're still susceptible to the same problems when there's no AQM component. Some TCPs do a decent job of backing off when latency climbs, before the buffers actually fill and start causing packet drops. But AQM in the router can more directly observe and act on congestion, and works with older TCPs and non-TCP traffic.
There's probably over 2b routers without bufferbloat fixes installed, and only if we work together to get them deployed will the internet as a whole get better, more capable of handling web, games, videoconferencing, and other applications that demand consistent low latency.
Thanks.
I was happy to wake up a few weeks back, and find RFC8290 published, and sch_cake being readied for mainline, and multiple commercial products finally shipping what we'd worked on all these years.
Results from the DSLReports speed test [2], using the same router, the only difference being turning smart queue on:
Smart queue off: https://i.imgur.com/zeY4rTd.png
Smart queue on: https://i.imgur.com/jfHpiFb.png
Notice the improvement in bufferbloat score from D to A+.
Subjectively, I have noticed the connection seems more responsive and there no longer seem to be latency spikes when utilising all the upload bandwidth.
It’s certainly worth considering bufferbloat, if you suffer from latency spikes when using all your upload bandwidth (I used to suffer from this a lot more, when I had an ADSL connection, with only 1 megabit upload).
When you need to shape even more bandwidth, routers based on the Marvel XP Armada chipset do very well. I flashed a Linksys WRT1900ACS with OpenWRT and was able to shape about 600-750Mbps before it ran out of horsepower.
I love my ER-X, but once I got my gigabit connection it can't quite keep up compared to plugging directly into the modem. It's close enough that I'm not looking to replace it, but the next time I need a router I may go the NUC build your own route.
I have done this at my parents place with a cheap Intel Celeron based computer with a single network port, but 2 logical networks with vlanning. One is for their personal network and the other is for their guest suite they offer through Airbnb.
It’s only 20MB/s fibre, so not much cpu power needed. The no name computer cost less than $250, came with a 60gb SSD and I put pfsense on it.
I am always happy to hear of a bloat free connection.
Note the comment "the actual rate limits will be set to 95% of the specified value". That explains why I saw a 5-10% dropoff in throughput when I enabled it with honest numbers.
I went from a buffer bloat of 1.2s UP and 0.6s DOWN to effectively 0. It's like I have a whole new internet connection.
OP and everyone else involved gets a million internet points from me!
Measuring RTT(round trip time) or packet loss rate. Unfortunately all the early (and still common) TCP protocols control their send rate by measuring packet drop. Sending faster and faster until upstream routers somewhere start dropping traffic. This has the effect of completely filling the outbound buffers of the slowest link in the chain.
To combat this, most bufferbloat fighting algorithms focus on dropping traffic before buffers fill, namely RED (random early drop) and CODEL.
Newer algorithms like TCP Vegas and BBR measure RTT and lower transmit rate when they detect buffers down the line filling, preventing bloat.
In most cases, you still need a router configured to prevent bufferbloat because even a single naughty protocol on your network can fill outgoing buffers.
The most important thing to know about controlling bufferbloat with QoS is that you MUST have an accurate estimation of your max upload/download rate. This is because you can only control bloat if you are in control of the slowest link where it will build first. All effective bufferbloat solutions rely on artificially making your router the slowest link, usually by limiting bandwitdh through them to ~90% actual.
Once you have control of the slowest link you can pick which packets get dropped as the buffers fill. And using something like FQ_CODEL you can assign equal bandwitdh to all IP's on the network. The nice thing about controlling bandwidth this way vs hard speed limits per user is that it allows users to use as much bandwith as they want until staying below line rate requires sharing
Also notice I keep saying outbound. You actually have far less control of bufferbloat on the inbound end. Best you can hope for is that dropping inbound packets coming in over the configured rate will stop whatever server from sending you more as quickly, but this isn't always the case. Luckily, in my experience, the vast majority of "lag" is due to outbound buffers filling, not inbound. So setting up bufferbloat fighting bandwidth sharing is usually extremely effective
1) BBR is currently not something I'd recommend at home.
2) The hope has always been that the core two bufferbloat-fighting algorithms (BQL, and fq_codel) would end up in the cable, fiber or dsl modem hardware, so that no shaping would be required, as there would be sufficient backpressure from the link itself to regulate the link intelligently. The cpu costs on this are nearly 0! BQL is 6 lines of new code in the device driver. fq_codel has been shipping in linux for 6 years. It's just a matter of turning it on...
But: lacking that support from the ISP-supplied gear, we shape with htb + fq_codel, (as you say), to ~90% of the link rate... with another box - or even in the same box if the device driver can't be fixed and is overbuffered. We are painfully aware of how much cpu shaping costs but modern cpus usually have enough oomph to handle it.
btw: We've come up with a new deficit shaper (sch_cake) that lets us get to ~100% of the isp bandwidth (so long as you get the wire framing exactly right), while providing vastly better queue management in the sqm system.
4) fq_codel is fair to flows, not devices. This works well in the general case, but has edge cases where abusive apps that open a lot of flows gain priority. Adding per host fq (while retaining per flow fq), even through nat, was the number 1 request from the users for sch_cake, and one of the main reasons why cake exists.
$TC filter add dev $IF_WAN parent 11: handle 11 protocol all flow hash keys nfct-src divisor 1024
At https://openwrt.org/docs/guide-user/network/traffic-shaping/...
Unfortunately every modem I've seen, including "business" models, doesn't use CODEL.
Especially regarding SQM... Is SQM just a new name for QOS or am I missing something?
QoE is a better, less overloaded, in the "qualitative, perceived" sort of description.
So (after endless debates) - we defined SQM as as a superset of classic as-defined-for-internet QoS and hoped for the best. (https://www.bufferbloat.net/projects/cerowrt/wiki/Smart_Queu... )
There's been plethora of other trade names for what we do with htb+fq_codel - streamboost, adaptive qos, etc.
I like that eero, edgerouter, and openwrt and derivatives also call what we do sqm. It simplifies the discussion, and the core scripts for linux generically are available as the sqm-scripts on github.
I also hope we see more RFPs specifying RFC8290.
Sorry yeah, I forgot about how the term has been mangled over the years. I was coming from a 3GPP perspective where the term is specifically defined as experience [0]
In particular:
* only the QoS perceived by end-user matter
* QoS definitions have to be future proof;
* QoS has to be provided end-to-end.
* QoS attributes (or mapping of them) should not be restricted to one or few external QoS control mechanisms
Not the first time the ivory tower of Telecomms has been out of touch with the outside world ...
[0] 4.1 (p. 7) https://www.etsi.org/deliver/etsi_ts/123100_123199/123107/05...
“SQM” is shorthand for an integrated network system that performs better per-packet/per flow network scheduling, active queue length management (AQM), traffic shaping/rate limiting, and QoS (prioritization).
“Classic” QoS does prioritization only.
“Classic” AQM manages queue lengths only.
“Classic” packet scheduling does some form of fair queuing only.
“Classic” traffic shaping and policing sets hard limits on queue lengths and transfer rates
“Classic” rate limiting sets hard limits on network speeds.
It has become apparent that in order to ensure a good internet experience all of these techniques need to be combined and used as an integrated whole, and also represented as such to end-users."
Netgear R6300V2
Using fq_codel because the others use more CPU than I’m confident these devices can handle.
I find that most people suffer from bufferbloat unknowingly, having bought fancy routers 5-10 years ago.
BufferBloat: What's Wrong with the Internet? - discussion with Vint Cerf, Van Jacobson, Nick Weaver, and Jim Gettys
https://queue.acm.org/detail.cfm?id=2076798
Bufferbloat: Dark Buffers in the Internet - Jim Gettys & Kathleen Nichols
(Also, it has a broken caching DNS server, and forces that broken server into its DHCP responses; the server it returns is not user-configurable.)
So what's an ordinary consumer to do?
The key issue here is contention, or a busy egress interface - that's when buffering occurs. I do not understand why adding capacity, or reducing the probability of a busy egress interface, "won't help at all".
pfsense has fq_codel.
Anything (1000s of routers) from lede/openwrt has the most advanced bufferbloat-fighting stuff in it, followed by dd-wrt, tomato, etc. If you need high bandwidths the multi-core arms are the best. fq_codel is pre-configured on all links (ethernet/usb/fiber/wifi/whatever) but if you need shaping to the ISP provided rate you need to configure it. All the research that went into fixing bufferbloat queuing problems everywhere landed in openwrt and lede first. Most of the research that improved tcp everywhere came out of google.
Most gaming routers now sold commercially have some variant of fq_codel in them in their trade name ISP "qos" system.
Also fq_codel derived anti-bufferbloat work has landed in many commercial wifi routers on the wifi side (eero, google wifi, some ubnt products, meraki, many others). The paper behind all that was: https://arxiv.org/pdf/1703.00064.pdf - happily that work was "good enough" to enable by default, and boy, does it make a difference if wifi is your bottleneck.
The current premier dsl router with cake is evenroute. I think there are several new models from several manufacturers that are going to get it right, soon.
Not tracking FIOS (gpon fiber) closely at the moment. Yes, fiber networks have bufferbloat, but it's harder to hit, and generally smaller than on dsl and cable technologies. I configured cake on sonic fiber recently and got 60ms back. (going from 60ms latency under load to 3ms )
Regrettably shaper setup is finicky and requires a few minutes of testing with a site like dslreports or a tool like flent.org to get right. If more ISPs published their shapers' bitrate and burst rate settings, life would be easier here... but the hope has always been they'd just ship a router with this stuff on and remotely configured to be "right".
In some ways, what are you doing about bufferbloat is
Depending on the hardware, you can upgrade the link itself with newer, smarter WiFi drivers. After pretty much solving the bufferbloat problem for wired connections, many of the same developers moved on to fixing WiFi, and some fruits of that effort are currently available in OpenWRT and LEDE.
I don't see anything about links anywhere around 10gig or higher, if I wanted this on my 100gig backbone or 40gig metro links.
IF things like fq_codel were deployed on those we'd not see latencies climb that much at all, we'd see bandwidths decline to the actual capacity available - and only the biggest flows would be hit to do so.
FQ_codel is lightweight enough to fit directly in high speed hardware and it does indeed run on 40gigE plus devices on linux, ddpk, BSD, in software. But it takes a long time for new chipsets to adopt new algorithms even if they incorporate support for deeply desirable features like ECN.
That said... 10Gbit to the home.... ooohhhhh. it's really hard to bloat that!
there's an awful lot of lit on FQ, what we do with fq_codel is to not only interleave packets better but apply congestion control signals at the right time so competing tcp flows don't overwhelm the link (with under 10ms of buffering (v seconds common on fifo ISP links)).
https://en.wikipedia.org/wiki/Fair_queuing
Of course, being perfectly fair to flows is sometimes undesirable, but making something strictly higher priority[1] is fraught with peril as you end up with a classification nightmare.
Having fq gives you the best shot at smaller flows completing sooner, and of big flows sharing better with each other.
Having vastly reduced buffering improves the responsiveness of competing TCP flows a lot, grabbing bandwidth whenever available, faster.
My take on folk that want "prioritization" is ask them to try some variant of sqm with just fq and codel and get back to us. being fair with well managed buffers works really well.[2]
[1] making something lower priority than best effort is actually a good idea. [2] but if you really want some flows or devices prioritized, see the sch_cake work mentioned on this thread. I still tend to think per host FQ is what many want rather than attempting to raise the priority of certain flows from certain services.
We use diffserv for this, for apps willing to use it. Example: ssh sets the imm diffserv bit for interactive use. cake respects that (I've cited the relevant paper elsewhere, another place is https://www.bufferbloat.net/projects/codel/wiki/CakeTechnica... but after extensive testing we settled on 3, rather than four tiers of priority)
stuff derived from the sqm-scripts use the same method (using htb + fq_codel) but the problem has always been that diffserv is not respected end to end. However, within your network, you can make your intention known and have it work, if you have the bottleneck.
Also, we have always made the latency/bandwidth tradeoff explicit - if you want less latency, you must want less bandwidth. It's the only safe answer to apps gaming the diffserv markings.
But I'm on cabel, not DSL.
In reality I don't see buffer bloat on the internet adding jitter of more than a few milliseconds. I do see loss though or 20, 50, even 100ms. I'd rather have 50ms of jitter than 50ms of loss, but that's just my application.
fq_codel and cake use a tiny bit of packet loss to get a sending host to back off, for example to keep a large download flow within the limits of your home link. Other flows aren't affected.
Bufferbloat regularly adds hundreds of ms on home internet connections, you can get an indication of your bufferbloat on http://www.dslreports.com/speedtest
I don't see any buffer bloat or excessive jitter on my home internet (at least on wired connection) on BT ftth.
But, I too am allergic to loss. fq_codel supports ECN, (explicit congestion notification), which is enabled, now, universally by IOS. As near as I can tell, the ECN usage of 6% (https://www.ietf.org/proceedings/98/slides/slides-98-maprg-t... ) in france is almost entirely from free.fr's deployment of fq_codel which they enabled by default in 2012 (!!!!!!). I had expected all the ISPs to have lept on this by now....
I then get a loss on a 30mbit stream (3000 packets a second) of 150 packets. At the exact same time I get a loss on a 20mbit stream of 100 packets, and a 10mbit stream of 50 packets.
This is an outage for 150ms, probably because of a reroute in an MPLS network somewhere.
My packets have already been emmitted by the time any round trip resend would have come back.