Ask HN: What is an acceptable latency for a web load balancer?
What do you think folks?
What do you think folks?
However, this should only be done at the beginning of the connection. After this the client will have a symmetric encryption key that is much faster to use. Their load balancer should be caching these sessions so that subsequent connections don't need to re-negotiate a session key.
If this 110ms is only on the first request, and a cache miss on the sessions, then I'd say that's probably something you should be expecting. If it's after the TLS session has been set up, or on a cache hit, that sounds bad. It also could be that their session cache isn't large enough and is forgetting sessions too soon, causing more TLS negotiation than may be necessary.
300ms would be rather on the high side. RTTs with todays infrastructure are often 30-50ms if you hit an Edge server (opposed to a server around the world). So you should end up up at around 100ms for the complete TLS connection establishment.
Can you provide the timings this command produces for you?
curl -s -w 'TCP=%{time_connect} TLS=%{time_appconnect} ALL=%{time_total}\n' https://YOUR_REMOTE_HOST_GOES_HERE -o /dev/null
That should provide an easy to use ruler to compare measurements.Here is the output for 100 reqs: https://ybin.me/p/1e68f80f6910ba0d#USqFo3Loksx5rHQIe16NHp299...
$ curl -s -w 'TCP=%{time_connect} TLS=%{time_appconnect} ALL=%{time_total}\n' https://<my-domain>/ -o /dev/null
TCP=0.021106 TLS=0.052928 ALL=0.143908I'm not saying this to minimise like "oh it's only 100ms" I mean there are providers out there that will work with you on this. They're not cheap.
If you pay enough you can reshape the internet - I once had an issue with a backup process and ended up having two ISPs start peering with each other to fix it.
I don't like playing these kind of games but sometimes you have to do it.
If you're talking about "tech", I'd say 100ms doesn't matter for the vast majority of use-cases. Bear in mind that Australia/New Zealand have a ~250ms round trip latency to most web services, 4G can easily have a 300ms latency to Google, and developing countries are often in a much worse position. Depending on <100ms for anything is a non-starter for most of the web.
If you're a service provider several layers removed from a human, then there may be enough between you and the user that it's critical that you respond in <100ms, but this is not most of the industry.
More generally, I find this sort of comment to be similar to the sort of comment that asks why anyone would write a web service in a language as slow as Ruby or Python when we have languages like Go and Rust. They ignore the fact that there are huge, successful companies running on Ruby, Python, and other "slow" languages, all for very good reasons.
> And don't forget about the studies out there that show a relationship between load time and revenue in e-commerce: "A 100ms decrease in checkout page load speed amounted to a 1.55% increase in session-based conversion."
This is true, although I believe this relationship doesn't hold when you're moving from 200ms to 100ms, it's more for when you're moving from 1000ms to 900ms.
I also have huge issues with a lot of those studies; 100ms from 200ms -> 100ms is immensely different than 1500ms -> 1400ms.
The speed of light is quite impactful when you're looking at different countries or continents.
Starlink: hold my beer.
I'm not a HFT, I know I can't increase the speed of light, I know I can't tunnel through the earth to shave a few microseconds off a transmission. I can appreciate the difference between 300ms rtt and 900ms rtt - the later being a practical minimum when you're coping with outages in the 100-200ms range and networks reroute.
Give me jitter over drops any day.
It's not about <100ms latency, it's about a <100ms outage
I know I can't rely on providers to not have outages on 100ms -- a link fails in Sudan and BGP flaps on a peering it cane take seconds, let alone milliseconds, to respond.
I have to be very careful, because if I want a two-way communication at that 250ms round trip time to NZ, that means I have to ensure my packets get there and back in 250ms, and not drop in the middle and require retransmissions, which rapidly escalates latency from 'barely noticable' up to 'painful' levels (1 second or above is meaningless, may as well send a fax[1])
0ms : send packet containing informaiton
125ms : receive packet
135ms: process informaiton, send response
260ms: receive response
270ms: process response
Job done.Trouble is, a 100ms outage from 0ms to 100ms means
0ms : send packet containing informaiton
125ms : fail to receive packet
135ms : notice packet missing, issue request for retransmit
260ms: receive retransmit request
260ms: retransmit
385ms : receive packet
395ms: process informaiton, send response
620ms: receive response
630ms: process response
Given I therefore need (630-270 = 360ms) of buffer, that's more than doubled the response.But it gets worse
125ms : fail to receive packet
135ms : notice packet missing, issue request for retransmit
(request gets lost)
395ms: notice retransmit request not received, issue request for retransmit
405ms: issue retransmit request
530ms : receive retransmit request, retransmit
655ms: receive packet
665ms: process information, send response
790ms: process response
So to cope with a 100ms outage my application needs to insert a 520ms buffer on a 270ms conversation - practically tripling latency, and converting a near- real time convertsation into near-catastrophic talking over each other.Now this is fine, I can have two circuits, send the packets down both, and the chance of both being broken is minimal -- assuming the links are really separate (which is another challange. Hint, both patchs down the same undersea cable on different frequencies is not separation)
However getting network providers to acknowlege, let alone measure, outages of 100ms, is practically impossible on a country basis, let alone an intercontinental basis.
[1] hyperbole. Just.
If the response to that is "there isn't one", well, there's your real problem.
What OP is asking is whether the performance they are seeing is comparable to the industry-leading load balancers that HN readers are familiar with. The answer to that question is independent of any contract.
Software can always be made better and can always be made faster. But you do have to know when to stop, and for me it’s always at “what my users will tolerate.”
Part of the reason there are all these terms and conditions documents that nobody reads is to prevent this ambiguity (to a point). Although legally it's a bit weak because you cannot really expect people to have agreed to all terms (but it does reverse the burden of proof somewhat, so it's still useful).
A quick google search indicates with TLS the whole think takes at least 4 RTT_1, plus the RTT_2 because the LB still has to fetch your data, so your hard limit is 24ms, plus how long it takes to actually transfer the additional data (especially certificates).
This leaves about 86ms.
Now, does your client perform any checks for a revoked certificate? Assuming these checks are also done via TLS, and a "far away" server with a RTT of 15ms, then that's a whooping 60ms your client is taking to ensure that the certificates are valid. In that case your LB would merely be 26ms slower than the theoretical optimum.
Give httping [0] a try for benchmarks.
To compare network latency in terms everyday techies can follow, use video gaming: 15ms latency is 60 fps, and 100ms latency is 10 fps.
Which game would they want to play?
That said, it's not clear if you mean "time to first byte" or "time to transfer 1000 bytes" or both combined. The first is less of an issue; figure out which part of the request and data transfer timing is changing.
(Disclosure: In a past life I built a white label global content delivery network.)
It's fine because the latency is constant. Your brain internalizes that there is a delay between moving buttons and things happening. Your brain internalizes that you need to shoot where things are going not where they are now.
If anything, latency is an example that consistency is more important than speed.
Why do gamers care about not just TV frame rates, but how many frames delay (latency) come from the TV's video scalar / processor?
When dealing with round trip times (RTT) for TCP or video gaming, latency limits the number of round trips or control loops per second and the gap between input and action.
To your consistency point, agree -- jitter matters a lot to human perception and response too, why so many games "lock" to 30 or 60 fps rather than vary at rates above 30 or 60.
Also, it was not an example. It was an analogy to put the differences in milliseconds into amounts of time that matter to a broader set of people than networking times matter to.
You could feed a thousand frames per second to a TV and the TV can have 100 ms latency (delay to display a frame after it received it).
You could feed 30 frames per second to the same TV and it'd still have a 100 ms latency.
The latency is rarely advertised. It's highly variable per TV and per image. There's the time it takes to start displaying an image and the time until the image is fully set (it is much longer to go white to black than white to grey).
That said, anyone who wonders why they can't download a file over TCP/IP at gigabit speeds on their gigabit FIOS is realizing there is a link between latency and throughput (network frames per second).
Put another way, latency is not related to number of frames that can be sent at once (bandwidth) but if RTT factors in flow control, those two work together to limit throughput.
TCP is not exactly a pipe. It's sending a message and waiting to receive a confirmation back before sending more. It's control logic to control the flow in a pair of pipes, that's affected by the latency back and forth.
UDP doesn't do that. The Ethernet link doesn't do that. The HDMI link to a TV doesn't do that.
If, by contrast, you're showing me you understand this, sounds like you do. We had developed an internal tool before Aspera's FASP[1], for our clients to deliver us video, and us to distribute content among hubs at scale, at speeds that saturated links. Understanding the mechanics and limiting factors of information propagation was necessary to be on that team. Turns out most devs don't have a great grip on the physics -- in the most common mental model, everything happens instantly. And so things get slow.
There is one thing I'd like to insist on though, I'm not sure whether people could actually notice 100 ms of latency after user input. You assume that users do? but I don't think they necessarily do. I think the full delay between when you press a button on a gamepad and when the frame changes on the TV is somewhere around that and gaming consoles are generally considered very usable. (Of course latency in computer networks is a different matter).
My daily job among other things is to optimize software and networking infrastructure covering 4 continents and 50 datacenters. I'd say I have a certain grasp of networking. ;)
This is where the Xbox Series X really started to feel even more PC-like to me. I play at 165Hz with frame rates that exceed 200fps in games like Destiny 2, Valorant, Call of Duty: Warzone, and CS:GO on my gaming PC. I do this a lot of the time by dropping a lot of the quality settings lower because I personally value frame rates over visual quality. Running around the versus multiplayer mode in Gears 5 in 120Hz felt like I was playing on my PC. With frame rates hitting 120fps, input lag is reduced, and the experience was suddenly so much smoother than what I’ve ever experienced on an Xbox One X.
That same feeling of PC-like smoothness plays out in Dirt 5 with the 120Hz mode enabled. Sure, the game drops to rendering at 1440p and some of the visual quality is lowered to achieve 120fps, but when I’m sliding around corners and the input latency is reduced, it’s far better than some mud and snow rendering just that little better on my 4K TV to the point I probably wouldn’t notice the difference.
It’s that feeling that’s really important with this new Xbox, and I can’t stress it enough.
-- https://www.theverge.com/2020/10/15/21515790/xbox-series-x-p...
TLS has a non-trivial bootstrap cost, due to the need for extra network round-trips during establishment (mitigated with 0-RTT, if available), the additional byte overhead of certificate exchange (which can be substantial relative to small payloads if you have a long chain or are sending unnecessary certs, and more so if you're using RSA keys), and the cost of performing the crypto operations for the public key part and key exchange.
So if your use-case is "client arrives from the blue with a one-off API request and latency is important" then you are going to suffer with TLS.
If this is a web application (on the other hand), then what's actually going to happen is that the browser is going to establish HTTP/1.1 connections over TLS to the load-balancer, keep them alive, and re-use them. Assuming a well-configured load-balancer, it's also going to have the same from it's back-end to your actual service implementation.
However, once that's done, the only necessary overhead is the symmetric key stuff (microscopic) and maybe rekeying in long sessions.
Using curl to do a one-and-done connection to your service (as suggested by another poster) will give you an estimate which is only relevant to the first use case I described.
However it's possible Citrix is just not as optimized as other providers/services or not support newer protocols like TLS 1.3
However, 100ms doesn't seem critical
Some of the settings there link to tips for more investigation. Look at the session resumption settings in particular.
If possible, Eliptic Key certificates and ECDHE can be a lot faster than RSA certificates and RSA based DHE. Especially if the load balancer is CPU constrained. Make sure TLS sessions or tickets are honored (tickets prefered over sessions so servers don't need to maintain a session database)
In my personal opinion, I'd rather not have the load balancer terminate TLS, but you lose a lot of features that way, and that might not be an option.
Tested with this: https://ybin.me/p/53f0aa4e1204b470#c561Frz6wpKkfvsTklDdZFOer...
and got this: https://ybin.me/p/1e68f80f6910ba0d#USqFo3Loksx5rHQIe16NHp299...
To test session resumption you can do: "curl https://your-url https://your-url https://your-url https://your-url" instead.
I can tell you from my personal experience that +85ms overhead is definitely not what I'd expect for TLS termination.
https://developers.google.com/speed/libraries
d3.js on my network is ~200kb and downloads in 20ms with ssl.