New phenomenon breaks inbound TCP policing
forums.whirlpool.net.au
forums.whirlpool.net.au
I have seen this with my wife's computer. Though I haven't seen the large-window phenomenon (haven't looked), what I do see is dozens of active connections opened to the same destination, which of course defeats both TCP's link sharing and standard QoS algorithms.
I work around the issue by bucketing any Akamai IP ranges I find into a very-low-priority queue, and let those TCP connections fight it out. Seems to have worked well.
For those interested, here are the Akamai IP ranges I use:
23.0.0.0/12
23.32.0.0/11
23.64.0.0/14
23.72.0.0/13
104.64.0.0/10
2001:428:4403::/48
2001:428:4404::/48
2001:428:4405::/48
2001:428:4406::/48
2600:1400::/24Even if it is isolated to the servers distributing MS's updates, I'm having trouble seeing how this could even be MS's fault-- it's the server side (owned and operated by Akamai) that's misbehaving here.
If there are also congestion-algorithm shenanigans like the forum post suggests, then Akamai is definitely at fault. But the issue I see is that of connection hogging.
We don't use TCP because its fast. We don't use it because its reliable (although that's really useful). We use it because _we kept breaking the internet_. Once you get above a certain threshold, the network can't keep up with you and packets start getting dropped. The problem is that backing off just a little doesn't allow the network to recover.
Instead, we need to use exponential backoff in the face of packet loss to ensure that the network as a whole can recover.
But if you're pretty much the only connection misbehaving, and everything else backs off, then you can kinda get away with not using exponential backoff. The problem is that the applications that is was "kinda okay" to do this for was VOIP and friends, where realtime delivery is really important and exponential backoff causes noticeable drops in quality.
For a great read about these kinds of issues, check out the TCP-Friendly rate control RFC: https://tools.ietf.org/html/rfc5348
Another aspect of this problem is that the network is too hesitant to drop packets [1], so by the time you've noticed packet loss things have gotten bad enough that the drastic backoff is needed. Widespread deployment of ECN and AQM would allow for more rapid feedback before any huge backlog develops, and consequently a less extreme response to congestion signals could be used.
[1] Arista would rather their 10GbE switches add up to 100ms of queuing delay per port than drop a packet: https://lists.bufferbloat.net/pipermail/cerowrt-devel/2016-J...
That anyone can think adding 100ms of latency to a 10Gbe switch, even under heavy contention is a good feature is absolutely staggering.
Try this : iptables -A INPUT -m statistic --mode random --probability 0.001 -j DROP
And see how your internet works. TLDR: sometimes loading times go through the roof, some instant messages go through in <0.1s, and on occasion it takes 30+ seconds, on occasion it's a DNS query that gets dropped and a page load suddenly takes 1 minute for no identifiable reason, large downloads always "get fucked" (suddenly lose 90% of their bandwidth and take several minutes to recover). Burstly traffic doesn't work. If you start your firefox with 20+ tabs open 80% of them will never load.
You will not enjoy the experience.
So yes, people think that adding 100ms of latency is better than dropping a packet under contention.
A 10GbE network in a datacenter without bufferbloat would have RTTs orders of magnitude smaller than the 100ms queuing delay Arista considers acceptable; the effects of a congestion event would be ancient history by the time Arista's queues could drain. Even outside the datacenter, 100ms is a pretty long time for most connections in a managed-queue world. A congestion event on a device using fq_codel won't kill your DNS request or TCP handshake; it'll slow down an established flow and if you're using ECN you won't even lose a packet. It's only in a DDoS-like scenario of thousands of unresponsive connections (such as TCPs with a large initial window) beginning to transmit simultaneously that you'd see some flows getting unfairly penalized, but things would equalize within a few RTTs if the traffic was real TCP and not a true DDoS. You only see it take minutes for a download's throughput to recover if you're going over multiple satellite links or through a severely bloated queue.
The real kicker is that the connections are all to servers (Akamai) on port 80, so any serious blocking breaks all web browsing. The cynic in me says the whole Windows 10 update thing has been made to operate in lockdown environments when non-well known ports are blocked. Intentional, or not, the Internet is basically broken while this happens as Windows is ubiquitous and people all over the world who have successfully used inbound rate limiting to create successful shared Internet connections are going to be getting angry support calls. I hope my post goes viral so it starts to get seen by the likes of Microsoft and Akamai engineers. The local ISP I spoke to where I initially noticed this problem pretty much fobbed me off with the old "nobody else has reported the issue."
Daniel
The trace is indeed a total mess, but I'm not convinced it's anything to do with TCP acceleration. There's absolutely massive levels of reordering and packet duplication happening in ways which are not consistent with TCP acceleration at all. It's much more likely that it's some kind of configuration problem elsewhere in the network.
From eyeballing the trace, almost half the payload segments there are duplicates, while a much smaller proportion are retransmits. (You can tell the difference e.g. using IP ids or by TCP timestamp TSvals / TSecrs).
Submitters: the HN guidelines ask you please not to rewrite titles except when they are misleading or linkbait. It seems like in this case the rewrite made it more misleading.
If anyone suggests a better title, we can change it again.
I'll have to see whether this is what causes my 100 megabit downlink to behave as if it's capped at 30 megabits sometimes. The router can barely keep up as it is.
Not all iptables rules are created equal.
State tracking has significantly more overhead than other types of rules/filtering/shaping (even though state tracking is required for certain types of shaping). You may or may not have already done this but if not try this: iptables -t raw -A PREROUTING -i eth0 -p tcp -m tcp --dport 443 -j NOTRACK
iptables -t raw -A PREROUTING -i eth0 -p tcp -m tcp --dport 80 -j NOTRACK
(replace eth0 with your Internet interface)State tracking on port 443 and 80 is mostly useless and it's where you're likely to see the majority of your (high bandwidth) traffic. Setting NOTRACK on those ports can make a HUGE difference while still enabling your squirrel-powered router (let me guess: It requires no active cooling? haha) to shape traffic like "teh big boys."
Thanks for that info. As soon as I can work out why the default congestion control / single connection speed on FreeBSD 10 is so bad, I'll be running that as a router.
I have seen my wife's Windows 10 computer upload crap from time to time and I suspected but could not confirm that it is this. How is this even possible, given that we're behind a NAT? TCP simultaneous open, I suppose?
I've not yet been able to figure out how to block these shenanigans. I really don't appreciate MS/Akamai profiting off of my (rather limited) upload bandwidth.
http://arstechnica.com/business/2012/05/skype-replaces-p2p-s...
https://en.wikipedia.org/wiki/Microsoft_Notification_Protoco...
They also introduced serious security problems when they made that change... So instead of Skype messages going from one client directly to another they go through Microsoft's servers (where they are stored and intercepted by TLAs) unencrypted.
They also introduced a new "feature" whereby their systems read everything you write:
http://www.h-online.com/security/news/item/Skype-with-care-M...
If you have a decent router (I have a Mikrotik; DD-WRT or such should work too) you can bucket all Akamai traffic into a low-priority queue. I posted an IP list elsewhere in this story.
It still has some issues when multiple machines download the same update, maybe this is when you home sysadmins could start running a caching server?
I believe it was really only update related but I saw it happen when actually no updates were available. I just ended up limiting the corresponding process bandwidth whenever it gets annoying.
One into the other, I'm mainly just very surprised this kind of thing can happen. Except that I'm fairly happy with win10 though as opposed to the usual MS bashing we can hear. One thing into
"It was the same range of source addresses and this was with Windows server and then Office updates. What seems to be happening is that instead of the sending server reducing its window size when packets are dropped, it just keeps re-sending large windows, which are obviously being dropped at my end. The queue algorithm has no idea of this and it will be letting packets through at a rate it thinks is correct, so the flow continues even though much of the traffic is dropped. However as the traffic keeps coming, the link is totally saturated."
Translation: someone broke TCP flow control.
There's also the Windows Update Delivery Optimization P2P feature: http://windows.microsoft.com/en-us/windows-10/windows-update...
I basically just configure my network as if everything connected to it – both clients and the ISP – are bad actors trying to DoS me (albeit, a DoS regulated by TCP Vegas). Between my ISP's bufferbloat and these shenanigans from MS/Akamai, it's not a bad approximation.
Keeps the local networks happy and fast. Isn't as expensive as routing everything through a datacenter because only misbehaving IP's get rerouted. Had the added side benefit that I could "protect" offices from the "involuntary" win10 upgrade.