The Latency/Throughput Tradeoff: Why Fast Services Are Slow and Vice Versa
blog.danslimmon.com
blog.danslimmon.com
Practically speaking the only way to really reduce latency to the minimum of the software is idle resources awaiting work.
Indeed. This is similar to the buffer-bloat situation, where limiting your download speed to 90-95% of max capacity will allow for both high throughput and acceptable low latency.
Maybe the system could tune itself in an audio compressor[1] like attack-release style, where it would never go above say 95% utilization (random number), but could dynamically lower it if low-latency requests suffer, then slowly raising it as low-latency requests tapers off.
[1]: https://themixingtips.com/what-is-compressor-attack-and-rele...
This also doesn’t get into the real challenge, which is the trade offs for tail latency for median (or more typically 99%) latency. Most of the time the things you do to flatten tail latency increase latency lower in the curve.
Of particular note: even if you're already operating in a distributed, partitioned etc. environment, improving per-node efficiency often means that you can serve a given workload with fewer nodes and reduce scaling costs. In the extreme case, the less efficient system would require so many nodes even if it scaled linearly that the more-than-linear coordination costs would kill it entirely. In other words, there's always an absolute limit on how "big" your system can grow. To raise that limit you must often address both per-node performance and scaling efficiency.
This is not merely hypothetical. When I was working on distributed filesystems, a competing system that was genuinely more scalable in terms of node count was so ruinously inefficient on each node that they not only couldn't beat us on latency but they basically couldn't beat us on throughput per dollar at any level either. So they resorted to politics instead of engineering, but that's another topic. The key point here is that eliminating waste will typically improve both metrics while adding scalability will only improve one at most and might actually not even raise the ceiling on that.
Once a system is "well implemented" as I've described, the OP's tradeoff becomes relevant. But TBH a lot of systems just never get to that point. They end up being good enough for business purposes even when there's still tons of room for technical improvement. Often, some simple math around utilization levels and controls around batch sizes will suffice.
2. The trade-off is more about computation in general. In typical "service" scenarios you almost always want to prioritize latency over throughput, while in most "batch processing" scenarios which you want to optimize overall processing time you want to prioritize throughput.
Examples:
1. Erlang server: Great preemptive scheduling thus the theoretical overall computation throughput is relatively lower. But in practice, Erlang servers are extremely fast and responsive. (Prioritize latency over throughput, great for "services")
2. Traditional C/Java blocking (non-async) server spawns threads: Theoretical computation throughput is greater but in reality, it's very easy to DDoS it. (Great for batch processing but suck at "services")
The problem reminds me how we made a lousy internet connection usable across a whole dorm using queues in m0n0wall to prioritize low latency traffic
The article assumes "max throughput" means "operate the cluster at 100% utilization". But for a parallelized service that's actually assuming they want "max throughput per $". As described, BigTable can have low latency and high throughput if you are willing to overspend on compute.
I've encountered these situations and solved it using common sense that might vary case by case. In one case we had 3 types of database servers: writer, normal-query, slow-query. The last one was added because certain queries would use up significant cpu and i/o.