0. Another thing people may want to optimise for is performance per watt, but I won’t say much more about it.
1. There are cases where bandwidth optimisations are latency optimisations, eg if you can fit more of your processes onto one box, you are reducing the average distance between the processes and whatever they talk to and hence average latency
2. A very obvious thing to do when optimising for latency is increase bandwidth enough that the bandwidth doesn’t throttle you
3. I feel like mostly if you are aggressively optimising latency, there isn’t much Linux tuning to do. Maybe I’m wrong – I don’t really know much about this – but I think it’s mostly pinning to a core, running tickles, doing user space networking, and then hardwarey things like tuning page size, SMT, power-saving settings, and other things like choice of hardware.
[1] https://pdfs.semanticscholar.org/bce7/5f78d340cac32dccd8631f...
Sometimes you can improve both, but often it's a tradeoff.
Another latency metric that you'll see, often w/respect to web apps and microservices is "P99" and similar. This is the amount of time in which 99% of requests get their response. For a higher percentile, you get a better idea of worst-case performance.