Boosting upload speed and improving Windows' TCP stack
dropbox.tech
dropbox.tech
* What are the reasons for disabling TCP timestamps by default? (If you can answer) will they be eventually enabled by default? (The reason I'm asking is that Linux uses TS field as storage for syncookies, and without it will drop WScale and SACK options greatly degrading Windows TCP perf in case of a synflood.[1])
* I've noticed "Pacing Profile : off" in the `netsh interface tcp show global` output. Is that the same as tcp pacing in fq qdisc[2]? (If you can answer) will it be eventually enabled by default?
[1] https://elixir.bootlin.com/linux/v5.13-rc2/source/net/ipv4/s... [2] https://man7.org/linux/man-pages/man8/tc-fq.8.html
re: pacing: Awesome!! I would guess it is similar to Linux "internal implementation for pacing"[1]. Looking forward to it eventually graduating form being experimental! As a datapoint: enabling pacing on our Edge hosts (circa 2017) resulted in ~17% reduction in packet loss (w/ CUBIC) and even fully eliminated queue drops on our shallow-buffered routers. There were a couple of roadbumps (e.g. "tcp: do not pace pure ack packets"[2]) but Eric Dumazet fixed all of them very quickly.
[1] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin... [2] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
I got hit by the exact same issue which is described in the fermilab paper, namely packet reordering caused by intel drivers. It took me several days to diagnose the problem. Interestingly enough, the problem virtually disappeared when running tcpdump, which, after a lot of reading on the innards of the linux TCP stack, and prodding with ebpf, eventually led me to conjecture that it was a scheduling/core placement issue. Pinning my process clearly made the problem disappear, and then finding the paper nailed it.
Networks are not my specialty (I come from a math background, am self taught, and had always dismissed them as mere plumbing) , but I have to say that I came out of this difficult (for me) investigation with a great appreciation for networking in general, and now enjoy reading anything I can find about them.
It's never too late to learn, and I have yet to find something in software engineering which is not interesting once you take a closer look at it!
> Networks are not my specialty
I wish all network non-specialist were like you!
The original symptom was very low throughput, which is what prompted the investigation . Without tcpdump, low throughput and high reordering, with tcpdump, high throughput ( which is why I couldn't figure out what was going on).
I'd be very interested if someone with kernel experience could tell me what's specific about tcpdump.
This entire thread is interesting because it is highlighting a similar problem with my virtualized router. When I pinned the router vm to specific cpus the problem goes away. I switched to openstack which doesn’t give me the best control over cpu capabilities and the problem has manifested in a worse form.
My uninformed opinion is that there are underlying concurrency problems with multithreaded user land-kernel interaction and some nic drivers (consumer intel and Broadcom hardware)
I ended up reading quite a bit about congestion control while investigating a different issue (sending data from a 25gb box to a 1gb one over a 10gb/100ms latency link didn't work well since the bigger nic would saturate the smaller one which then dropped packets, and this caused the tcp window to shrink significantly ), and it was extremely interesting. The whole problem of having multiple agents competing with different strategies and incomplete information to maximize their network throughput also reminded me of economics.
Tcp pacing essentially solved the problem (and if not available brutally traffic shape).
I don't have the memory or patience to write a long and inspiring blog post, but it comes down to:
Even with IOCP/multiple threads: network traffic is single threaded in the kernel, even worse, there's a mutex there. Putting the effective limit on PPS for windows to something like 1.1M for 3.0GHz.
The task of this machine was /basically/ a connection multiplexer with some TLS offloading; so listen on a socket, get an encrypted connection, check your connection pool and forward where appropriate.
Our machine basically sat waiting (in kernel space) for this lock 99.7% of the time, 0.3% was spent on SSL handshaking..
We solved our "issue" by spreading such load over many more machines and gave them low-core-count high-clock-speed Xeons instead of the normal complement of 20vCPU Xeons.
AFAIK that issue persists, I'd be interested to know if someone else managed to coerce windows to do the right thing here.
Otherwise you can't easily do LACP because it can't be offloaded to the card.
We tried without LACP, but again, only one PCIe card.
Definitely, though enabling conntrack on Linux has similar characteristics (forces single thread with some kind of internal mutex) though it can do 5x the b/w..
We tried having stateful firewalls in front of our windows boxen, that's how I know.
Seems like Cloudflare has an older blog post detailing this too: https://blog.cloudflare.com/conntrack-tales-one-thousand-and...
Anyway, to answer your question: AAA GameDev (and their backends, if highly tailored) are Windows.
The harder part was aligning the outgoing connections; for max performance, you want all of the related connections pinned to the same CPU, so that there's no inter CPU messaging; for me that meant a frontend connection needs to hash to the same NIC queue as the backend connection; for you, that needs to be all of the demultiplexed connections on the same queue as the multiplexed connection. Windows may have an API to make connections that will hash properly, FreeBSD didn't (doesn't?), so my code had to manage the local source ip and port when connecting to remote servers so that the connection would hash as needed. Assuming a lot of connections, you end up needing to self-manage source ip and port anyway, and at least HAProxy has code for that already, but running the rss hash to qualify ports was new development, and a bit tricky because bulk calculating it gets costly.
Once I got everything setup well with respect to CPUs, things got a lot better; still had some kernel bottlenecks though, I wouldn't know how to resolve that for Windows, but there were some easy wins for FreeBSD.
Low core count is the right way to go though; I think the NICs I used could only do 16 way RSS hashing, so my dual 14 core xeon (2690v4) weren't a great fit; 12 cores were 100% idle all the time; something power of two would be best.
Email in profile if you want to continue the discussion off HN (or after it fizzles out here).
[1] Load balancing/proxying, but no TLS and no multiplexing, on FreeBSD.
I had been thinking that RSS/PCBGROUP was totally abandoned and could potentially be removed.
But yes, I think I ended up using both RSS and PCBGROUP. This was on a server running only one application (plus like sshd and crond and whatever), so it was dead simple to line up listen socket RSS and cpu affinity; I had a config generator script that would look at the number of configured queues and tell HAProxy process 0 to bind to cpu 0 and rss queue 0, up until I ran out of RSS queues; we needed a config generator script anyway, because the backend configuration was subject to frequent changes. If it was only listen sockets, RSS would have been sufficient without needing PCBGROUP, but locking around opening new outgoing sockets was a bottleneck and PCBGROUP helped considerably, but it was still a bottleneck. This was on FreeBSD 12.
Edit: I also found some patches[2] I sent to freebsd-transport that I don't know if anyone saw; I don't remember if I updated the patches after this... I know I tried some more stuff that I wasn't able to get working. Don't apply these patches blindly, but these were some of the things I had to fiddle with anyway. I think I saw there was some stuff in 13 that likely made outgoing connections better.
[1] https://www.mail-archive.com/haproxy@formilux.org/msg34548.h...
[2] https://lists.freebsd.org/pipermail/freebsd-transport/2019-J...
With everything tweaked, we got to 2M clients per server, and actually it was hard to find the limit, because I wasn't able to direct enough traffic to the machines under test.
The software and configuration changes weren't big, but it was a huge impact. On the other hand, if RSS and PCBGROUP weren't in the kernel, I don't think I would have been able to add something similar, and we would have had to something wild and crazy (or try Linux and see if it would do the job). Now, I really did want to write a raw packet tcp proxy in userspace, but I knew it would be a lot easier to manage and quicker to get working with something off the shelf.
Of course, maybe there's a better solution to the root bottleneck, which was always opening a new outgoing tcp connection; even with all the tweaks, that was still the bottleneck, but fixing that needs someone more skilled than me, and I guess it's a pretty niche use case to be opening so many outgoing sockets. Accepting tons of sockets is way more common and way more optimized.
It's inaccurate to describe traffic processing as single-threaded in the kernel.
Hmm. It definitely was not years ago. Perhaps something to do with a specific NIC driver? Or perhaps Cutler retired?
[1] https://docs.microsoft.com/en-us/windows-server/networking/t...
I just tested rn with DropBox, GoogleDrive, and OneDrive, all with their native desktop apps. I simply put a 300MB file in the folder and let it sync.
DB: 500 KiB/s
GD: 3 MiB/s
OD: 11 MiB/s (my max bandwidth with 100Mbps)
I don't know what causes the disparity here, but I have been annoyed by this for years, and it's the same across multiple computers I use at different locations.Another funny thing is if you just use the webpage, both GD and DB can reach 100Mbps easily.
Edit: should mention Google's DriveFS can reach max speed too, but it's not available for my personal account (which uses the "Backup and sync for Google" app).
[0]: https://support.google.com/googleone/answer/10309431#zippy=
The workaround was to back everything up to the "Google Drive" folder since this seems to be the only folder that Backup and Sync can actually restore.
At one point I set it up to use my second SSD as the local storage. Then I needed that SSD elsewhere, so I just took it out. It was impossible to restart the damn thing. It kept complaining about missing folders. I even tried uninstalling and reinstalling it, but it kept its settings.
Since I barely used that machine, if ever, and I'm not particularly familiar with Windows, I never really looked into how to completely clean up the configuration. But the point is that there clearly are some pretty stupid decisions about some products.
System-wide settings and state should be stored in C:\ProgramData.
But on a more useful note how I have handled this in the past is to download the complete Google Drive data using Google Takeout. Not the greatest solution but it has worked.
https://abevoelker.github.io/how-long-since-google-said-a-go...
Of course there are alternatives now, but I like to plug this page whenever I can.
That thing is far too aggressive about network bandwidth. It will upload 20 files at the same time and the speed limit setting doesn't work.
I've never seen it transfer more than five files at a time, which sometimes drives me crazy when there are a lot of small files to sync.
I still use DriveFS for everything else, at least for now. Rclone is capable of mounting the drive but it's not really designed for that.
It would be interesting to re-try the experiment on Linux or FreeBSD using BBR as the TCP stack and see if the results are any better for dropbox.
FWIW, my corp openvpn is kinda terrible. My upload speeds via the vpn did not improve at all when I moved and upgraded from 10Mb/s to 1Gb/s upstream speeds. When I switched to BBR, my bandwidth went from ~8Mb/s -> 60Mbs, which I think is the limit of the corp vpn endpoint.
(I work on a QUIC implementation in Rust.)
Dropbox does at least resume fairly reliably though, so I can generally ignore it the whole time... unless I have something I want to sync ASAP. Then I sometimes use the web UI and cross my fingers that I don't get a connection hiccup ಠ_ಠ
PS. One known problem that we have right now is that we use a multiplexed HTTP/2 connection, therefore:
1) We rely on the host's TCP congestion. (We have not yet switched to HTTP/3 w/ BBR.)
2) We currently use a single TCP connection: it is more fair to the other traffic on the link but can become bottleneck on large RTTs.
Ping result:
Pinging nsf-env-1.dropbox-dns.com [162.125.3.12] with 32 bytes of data:
Reply from 162.125.3.12: bytes=32 time=27ms TTL=55
Reply from 162.125.3.12: bytes=32 time=27ms TTL=55
Reply from 162.125.3.12: bytes=32 time=27ms TTL=55
Reply from 162.125.3.12: bytes=32 time=27ms TTL=55
Ping statistics for 162.125.3.12:
Packets: Sent = 4, Received = 4, Lost = 0 (0% loss),
Approximate round trip times in milli-seconds:
Minimum = 27ms, Maximum = 27ms, Average = 27ms
App Ver. 122.4.4867Is the OS being Win7 a factor? (Work computer, can't update [yet]).
Download speed is normal (100Mbps).
netsh interface tcp set heuristics disabled
netsh int tcp set global autotuninglevel=normal
netsh int tcp set global congestionprovider=ctcpI tried, and the first two lines helped the speed bump to 2.5MB/s! The third one doesn't seem to have any immediate effect.
Still not OneDrive level, but I'm more than happy.
I probably took them from Win7 forums or some stackoverflow spinoff, the TCP problems on Win7 are not entirely unknown.
Maybe they have grand academic visions and papers, but I've been using them for well over a decade and I feel the client quality has gone downhill over the past few years. They keep adding unnecessary stuff like a redundant file browser while the core service suffers.
The scaling model of the hardware is rather simple: hash over packet headers and assign a queue based on this. And each queue should be pinned to a core by pinning the interrupts, so you got easy flow-level scaling. That's called RSS. It's simple and effective. What it means is: the hardware decides which core handles which flow. I wonder why the article doesn't mention RSS at all?
Now the socket API works in a different way: your application decides which core handles which socket and hence which flow. So you get cache misses if you don't tak into account how the hardware is hashing your flows. That's bad. So you can do some work-arounds by using flow director to explicitly redirect flows to cores that handle things but that's just not really an elegant solution (and the flow director lookup tables are small-ish).
I didn't follow kernel development regarding this recently, but there should be some APIs to get a mapping from a connection tuple to the core it gets hashed to on RX (hash function should be standardized to Toeplitz IIRC, the exact details on which fields and how they are put into the function are somewhat hardware- and driver-specific but usually configurable). So you'd need to take this information into account when scheduling your connections to cores. If you do that you don't get any cache misses and don't need to rely on the limited capabilities of explicit per-flow steering.
Note that this problem will mostly go away once TAPS finally replaces BSD sockets :)
Anyways, nice reference for TAPS! Fo those wanting to dig into it a bit more, consider reading an introductory paper (before a myriad of RFC drafts from the "TAPS Working Group"): https://arxiv.org/pdf/2102.11035.pdf
PS. We went through most of our low-level web-server optimization for the Edge Network in an old blogpost: https://dropbox.tech/infrastructure/optimizing-web-servers-f...
In particular, the collaboration with Microsoft was great.I wonder what it took to make that happen.
In the future we are planing on having an HTTP/3 support which will give us pretty much the same benefits as SCTP with a better middlebox compatibility.
Theoretically, UDP would be the best choice if you had the time & money to spend on building a very application-specific layer on top that replicates many of the semantics of TCP. I am not aware of any apps that require 100% of the TCP feature set, so there is always an opportunity to optimize.
You would essentially be saying "I know TCP is great, but we have this one thing we really prefer to do our way so we can justify the cost of developing an in-house mostly-TCP clone and can deal with the caveats of UDP".
If you know your communications channel is very reliable, UDP can be better than TCP.
Now, I am absolutely not advocating that anyone go out and do this. If you are trying to bring a product like Dropbox to market (and you don't have their budget), the last thing you want to do is play games with low-level network abstractions across thousands of potential client device types. TCP is an excellent fit for this use case.
I am not saying that it should be done from scratch. But most recent research done in the recent years about the protocols used on the web tend to be built on top of UDP and not TCP, for many historical reasons.
In theory TCP would be the better choice, but in practice this is more complex than you assume.
I think that many people have a knee-jerk reaction when talking about TCP vs UDP, but they probably don't know as much as they think... (parrots)
[0] https://dropbox.tech/infrastructure/how-we-migrated-dropbox-...
[1] https://dropbox.tech/infrastructure/dropbox-traffic-infrastr...
Honestly I don't understand these orgs that don't go OneDrive/O365 suite. What product value does dropbox have when competing within Microsoft's own ecosystem?
Fixing something like this can help lots of use cases, but may have been difficult to spot, so I'm sure the Windows TCP team was thrilled to get the detailed, reproducible report.
And I found another Dropbox blog post about rewriting their sync engine from Python to Rust: https://dropbox.tech/infrastructure/rewriting-the-heart-of-o...
But it isn't clear whether the outer shell of the app might still be Python.
TL;DR is that they had RACK (RFC draft) implemented as an MVP but w/o the reordering heuristic.
[1] https://techcommunity.microsoft.com/t5/networking-blog/algor...
* 32-bit x86: https://web.archive.org/web/20191104120802/https://download....
* 64-bit x86: https://web.archive.org/web/20190420141924/http://download.m...
(those links via: https://www.reddit.com/r/sysadmin/comments/e4qocq/microsoft_... )
Or use the even older Microsoft utility Network Monitor, which is still available on Microsoft's website: https://www.microsoft.com/en-us/download/details.aspx?id=486...
Supposedly Microsoft is working on adding to the existing Windows Performance Analyzer (great GUI tool for ETW performance tracing) to display ETW packet captures, which will succeed Message Analyzer and Network Monitor: https://techcommunity.microsoft.com/t5/networking-blog/intro...