Optimizing global message transit latency: a journey through TCP configuration
ably.com
ably.com
To make a few examples: on many distributions you get 1024 as the file limits, 4KB of shared memory (shmall) and Nagle's algorithm is enabled by default.
Another thing that we noticed at work (shameless plug to getstream.io) when it comes to tail latency for APIs / HTTP services:
- TLS over HTTP is annoyingly slow (too many roundtrips)
- Having edge nodes / POPs close to end-users greatly improves tail latency (and reduces latency related errors). This works incredibly well for simple relays (the "weak" link has lower latency)
- QUIC is awesome
Is there a reason I'm missing why this wouldn't be worth jumping on?
https://0pointer.net/blog/file-descriptor-limits.html is a good overview of the unfortunate reason why this is and how it should be handled.
It's still a bad idea for performance reasons (though `poll` can actually be worse on dense sets), but it's not actually the open-file limit that's the problem.
It would've been more awesome if it supported BBR for congestion control. QUIC gains in practice can be annihilated just by not having BBR implemented in the protocol, so sometimes QUIC could be even slower than HTTP2 over TCP (if TCP is properly configured).
I think QUIC can be awesome, and I hope it will. But I wouldn’t say we’re there yet. Low level kernel-adjacent things takes time. Networks are extremely heterogeneous and weird. Maybe in 5-10 years.
In the short-medium term, I think we could get much more bang for the buck if there was an easy way to improve the defaults on Linux and/or its distros.
Of course depending on usecase etc the benefit from first-order network behavior improvements is almost certainly more important than the second-order cache pollution effects of replicated/seperate network stacks.
Optimizing a running service is often underrated. Many engineers focus on scaling horizontally or vertically, adding more instances or using more powerful machines to solve problems. But there’s a third option: service optimization, which offers significant benefits.
Whether it's tuning TCP configurations or profiling to identify CPU and memory bottlenecks, optimizing the service itself can lead to better performance and cost savings. It’s a smart approach that shouldn’t be overlooked.
That's just not true.
If the author is reading this - cwnd unit (in linux) is in packets not bytes. I'm not sure where they got the formula from (that's the receive window scaling formula, and has nothing to do with cwnd), but iirc slow start defaults to 10 packets. With their mss being 32K (a loopback connection) the actual cwnd size in bytes is at least 10 * 32K.
Thanks for taking the trouble to comment on my blog post. You are correct that cwnd unit is indeed in packets.
And the mss of a loopback connection is, as you say, 32K. But here we are concerned not with connecting to other processes on the same server, but routing across the WAN to the other side of the globe. And unfortunately, unless we are able to enable jumbo frames, we are stuck with the historical relic of 1500 packet sizes. This means that the initial congestion window size of only 15K.
Thanks for correcting my error about cwnd units. I'll fix the post later today.
Thankfully newer congestion control algorithms like BBR (which you can enable, though linux currently ships an older version) do not suffer from this and do not go to slowstart.