Don't trust default timeouts
robertovitillo.com
robertovitillo.com
Case in point: Google's Waze. If I have a slow mobile connection (e.g. edge or even 3g), Waze will repeatedly fail to load a driving route. It will think for a few seconds at most, then timeout and tell me there was a problem. If it only would wait a few more seconds to load, then the app would be useful. Instead, due to their crappy choice of timeouts, the app becomes useless.
There was no way to set a timeout in Fetch because the browser, acting as the user's agent, has a sane default (~75 seconds on average, but it varies by browser and platform).
Developers often pick TERRIBLE values for timeouts when left to their own devices.
Hell, the author of this exact blog post has picked 10 seconds in all of his examples. That a FUCKING BAD timeout. It's far, FAR too short for many use cases.
It isn't necessarily. It all depends on the use-case. If most of the operations are finishing in 5ms then the probability of something finishing after 10s are rather low - and timing out and retrying early is probably the way to go.
Someone else in this thread recommended setting the timeout to around the P99 time that operations take. I think that's a reasonable starting point, even I might move it towards P99.9.
I worked (and am still working) on adjusting timeouts for systems doing billions of requests/s. One takeaway from that is that the actual value of timeouts is often not too important if you look at one system in isolation. The latency distribution will be rather logarithmic. Most requests might e.g. finish in 20ms. Then you get a P99 at maybe 3 digit ms, and a P99.9 at 10s (example numbers). From there on it will make a minor difference in availability if you now set your actual timeout to 5s or to 120s - it might just be noise along your other error sources.
However it makes sense to align the absolute timeout with timeouts of dependencies. E.g. if you have a chain of
client -> service A -> service B
and service A always times out first, then the client will get an error that service A is broken - but nobody can easily diagnose whether it was service A or service Bs fault. If service B times out first, the client can get an error message that indicates that. Therefore it makes sense if upstream service timeouts are shorter (even if only by a second).
In the same model if the client times out first the services actually do only observe the client dropping the connection. They don't know whether the client timed out or cancelled the operation for other reasons. And therefore they might also not record that something in the service is actually not ideal. For that reason I would recommend setting client timeouts higher than service timeouts (if you are aware of them).
However there is yet another exception to this thing, which are TCP connection timeouts. If you can configure them separately, it makes sense to have those rather low and performing multiple retries. That can improve overall latency, since dropped SYN packets will only be retried by the OS after 1s.
If you're doing anything across datacenters, you have to take a minimum baseline of 10 - 15 seconds to account for extra latency on top.
If you do billions of requests/s I bet you don't care that requests fail? You probably can't even see that requests are failing because you'd have no logging, too expensive at this scale.
I do financial systems, most load typically doesn't go above 1k/s, but every request matters because a dropped request is a dropped payment, possibly tens of millions of dollars lost! There is a ton of issues caused by having too low timeouts set by developers (anything below 30 seconds). I had to reconfigure a ton of systems and libraries to have higher timeouts and ignore configuration passed by developers.
That's an assumption. Our customers care a lot. And we have sufficient monitoring in place. All recommendations I provided above where about maximizing availability, and provided based on the experiencing of improving the experience for lots of users.
> I do financial systems, most load typically doesn't go above 1k/s, but every request matters because a dropped request is a dropped payment, possibly tens of millions of dollars lost!
Without knowing much more about your system: If you are losing that amount of money for failed requests (which can e.g. happen due to random network blips) you are doing something wrong. You should invest into different strategies than increasing timeouts.
If the client decides to drop (timeout) and consider the transaction cancelled, while the server is processing it and will consider it done. That's a catastrophic issue that needs to be addressed. It is one of the most common bugs I've seen in the wild (root cause: too short timeouts).
How to make highly critical systems reliable enough in the face of hardware and software issues is a complex topic. At this level this involves a holistic approach to get every component to cooperate together (timeout is a minor example). A HUGE amount of work is to detect errors, and more importantly to propagate errors across diverse stacks (software should be aware of database errors, services should detect other services failing).
My understanding is that real payment systems solve the problem by just taking a day to finalize transactions…
Two stage commit is important because it has: 1) Predefined transaction id prior to final submission that allows you to validate the status, so if your request to commit gets 503'ed or you get a timeout you can reliably query to know if it was processed or not 2) Unlimited resubmissions of the final commit. It doesn't matter if I perform the final commit api request 1 time or 100 times, it will never cause a duplicated transaction to occur. So if I get a timeout or a 503 I can resubmit knowing that if my original commit request went in my new submit will be a no-op, and if my last commit request didn't get processed then this time it hopefully will be processed.
This pattern isn't just a payments pattern thing either. This is heavily used in distributed systems where failures can occur. UPS' API used to use this as well so you could be sure that you don't pay for duplicate shipping labels or cause duplicate shippments.
Make the client include the id of it's last known transaction and only apply the transaction if it's up to date, otherwise tell the client to refresh and try again.
The practical risk is that this puts a ton of complexity on the client, to keep track of states and perform some follow up actions. The added complexity means more bugs and each additional step can fail hence compounding the problem rather than solving it.
For instance, if I want to generate a shipping label that goes from my house to your house and I do two attempts, how does the receiving service know if I made two distinct attempts (I want to ship 2 similarly sized items) or if a transient error occurred in between making me attempt a re submission?
You solve this by creating an inactive request with the criteria (shipping label from my house to your house). This step is not idempotent but that's OK, because if I resubmit I just create a 2nd inactive request that may never actually be finished.
The second step is to say "this request is good and I want to proceed with it". That step is idempotent and marks the existing request as not just inactive but puts it in an active state.
A shopping cart flow is a user managed 2 stage commit (review your cart, submit the cart order). No matter how many times I submit my order it won't cause duplicate orders because I'm submitting a specific shopping cart.
UPS, Paypal, and others just use a computer/api-managed 2 stage commit
You can't always rely on a client generated ID, because you would have to know that the client id is unique enough. The server is the only one who can really generate a transaction id that it knows is globally unique and efficiently queryable in its backend.
Why would you intentionally drop 1% or even 0.1% of requests?
This also assumes O(1) I assume...
(I'm thinking of how to apply reasonable timeouts to background celery tasks)
I’d think infinity is not a valid state.
For waze’s case, I supposed their priority is not on salvaging the 1% longest request (though critical to you), and instead preserve server resources for the 99% faster clients. That’s not a “wrong” value on their side, and probably have been carefully tailored to get the right tradeoff.
What resources? Buffering a response takes a minuscule amount, and if even a tiny fraction of people try again it will waste far more.
And even if it did take more in total, it would not be by much. This justification for saying it's not a wrong value is very weak.
Let's say 10 seconds, typical intuitive but bad timeout. This will cause requests to fail for no reason other than users are in Asia or Africa, high latency. This will break the application when it's used or deployed across datacenters because high latency. This will cause requests to fail when the server is a bit busy (couple seconds more to process requests). Worse, it will cause chain reactions under load, creating more retries and even more load, causing other services/servers to timeout too.
Better go for a long timeout. A long timeout doesn't break the application.
I'm pretty sure infinite timeout also breaks the application, in a way people rarely realize that it is because of the timeout. People would rather think it "just didn't work, don't know why" instead of being very clever and realized "it must be low timeouts!!!"
To be pedantic though, infinite timeouts don't break applications except some rare cases of resources exhaustion. If an application is completely unresponsive, it is dead for good, not because of the timeout, need to fix the root cause (often resource exhaustion like swapping or it's waiting on another IO or service that's frozen).
You don't need a timeout, you need a "cancel" button.
This happens regularly when I open large files in some app, they take a fair bit of time to load, Windows offers a popup to kill the app after few seconds. Have to carefully wait and not click anything.
When it comes down to UI there are even more options. Since you have a human on the other side, you can transfer to them the responsability of deciding when to timeout. The UI certainly shouldn't become completely unresponsive while a request is being made.
That's an example where a timeout and retry would have fixed the problem. If it had been an API call behind an app, it would have hung indefinitely.
Some libraries sadly have their default timeouts set to infinite.
Systems I work with that have a default timeout are a pita. You end up having to make pointless retries when you'd have been happy to wait.
There are exceptions. If cancelling and retrying is has a decent chance of routing around the original problem. In that case, a timeout makes total sense. The other case is if you have a workload where operations tie up a mixed set of resources (e.g. threads + blocked backend calls) and only some of your incoming ops are dependent on the blocked resource. In that case, timeouts make sense in that they at least allow you to make forward progress on the unblocked requests. Although tbh separate queues and thread pools is the safer way to handle this. Because your caller with the timed out calls is gonna keep retrying and eventually these retries will crowd out the requests that can make progress in your incoming request mix.
Don't use infinite timeouts if you don't have a another way to cancel the operation.
There's nothing complicated about it, so there's no reason your code can't implement timeouts and cancellation the same way: timeouts are a cancellation triggered autonomously after some time passes.
Not retrying+timeouts has similar effects to cancellation. The operation ceases to go forward. But it is not the same. It's a lot more expensive than imperative cancellation (need to rebuild, resend, reparse the request) and it has a lot of production risks that waiting with cancellation doesn't. For example, naive retries can expose backends to thundering herds, and less naive retries can have strange issues caused by exponential backoff where you'll have requests sitting around doing nothing for half their own timeout, before giving up because the next retry did not hit before the end of the parent request's timeout.
And yeah, not having another way of cancelling is not nice, but sadly not entirely uncommon.
There's no way to cancel the operation remotely, because you're not authenticated yet. And you may not have any other access.
Timeouts are also a good defense strategy against bugs.
Enforce short timeouts, preferably less than 3 seconds, and definitively no longer than you expect the user to have patience for, if the operation is part of an interactive workflow. Any task that has a legitimate reason to take longer gets pushed to a background process or cron job. This makes timeouts on the frontend a regular occurrence, so you'll be forced to handle them just like any other error condition. Result: a more robust program.
But this probably depends on what kind of program you're building. Most of the stuff I build and support are consumer-facing, so anything that isn't instantaneous is cause for concern. Other types of applications, though, might have more patient users.
Not every app is forced to support terrible connections. We definitely had issues with hung operations until we had a timeout, though there was some disagreement about failing “fast” (10 seconds is “fast”???) vs hanging inexplicably for much longer, but not forever, periods of time.
Edit: when REST calls occasionally failed after 10 seconds, rather than hanging the UI, the support calls stopped. 2 or 3 people a day had to hit submit a second time. They got over it, as opposed to reloading the app/page. This was the most cost effective way for us to handle this.
Timing out would at least let you, for instance, flip a circuit breaker off or fail fast and have the resulting monitoring very specifically tied to the actual problem in the system, not to mention avoiding resource contention issues like I mentioned.
Go supports timeouts on network & file read/write ops, which can be used to interrupt them.
This is not true of Rust. Some of the convenience wrappers (std::io::Read::read_exact, etc) on top of the basic primitives (std::io::Read::read, etc) do retry for you (and explicitly document it), but not "Rust" as a whole. The primitives map one-to-one to calls of read/write/sendto/recvfrom and bubble up ErrorKind::Interrupted to the caller just fine.
https://github.com/rust-lang/rust/issues/11214
To support user intervention when a task takes longer than expected, all blocking syscalls should be interruptible.
std::fs::File's impl of std::io::Read:
https://github.com/rust-lang/rust/blob/7fc048f0712ba515ca11f...
-> https://github.com/rust-lang/rust/blob/7fc048f0712ba515ca11f...
-> https://github.com/rust-lang/rust/blob/7fc048f0712ba515ca11f...
Again, as I said, it corresponds one-to-one with a call to the underlying read API. The retries for ErrorKind::Interrupted are done by higher abstractions like std::io::Read::read_exact, and they explicitly document that they do this.
https://github.com/rust-lang/rust/blob/7fc048f0712ba515ca11f...
Have they taken out the EINTR retries which were added for the issue I linked?
Pre-1.14 -EINTRs were quite rare in "normal" Go programs so the stdlib basically ignored them, but 1.14 introduced preemption which resulted in many more -EINTRs and quite a few Go programs were broken as a result. So in many ways this behaviour was necessary to un-break backwards compatibility. If Go had made the interruption semantics -- which had existed for at least a decade before Go came about -- clearer from the outset then maybe this whole business could've been avoided.
This is symptomatic of the reasons why container runtimes (at least, those written in Go) have historically been vary wary of Go updates. Several years ago, each Go release would change some minor semantics of the Go runtime and cause breakages...
You need the retry button anyway, in case the server is throwing errors. And there's often no good place for a cancel button, without putting up a big 'Loading' animation.
Maybe increase the timeout on retry?
And that's the problem with timeouts, everyone has different expectations.
For any user-initiated action, automatic retry seems strictly better than failing on a single timeout.
With Postgres you can use roles to set timeouts, maybe you want a longer timeout for crons, shorter for HTTP endpoints.
Sadly we were using mongo which doesn’t have equivalent functionality. Ended up monkey patching the client library to define a reasonable default timeout.
As for "if the requester goes away", remember that the requester might be a few hops away. E.g., the HTTP connection from the mobile client drops; the web server and its connection to the DB is still alive and well. I can forcefully shut that connection, but that's somewhat of a drag (I'd rather keep it open, since it is perfectly good).
Beyond closing the connection, and being able to issue some form of "cancel this query" request: HTTP/1 lacks it entirely, PostgreSQL requires opening a separate connection, Redis lacks it entirely, and I think both Mongo and MySQL lack it entirely.
Even support for "time this request out" is spotty.
MySQL has a kill command, which does need to be done on another connection (might also need more permissions, it's been a while). It's been a while since I used a lot of MySQL, but this was definitely a pain point when things went sideways.
A co-worker of mine actually added support for timeouts to a database we were using. (It is a smaller, less-well known DB.) I added it to the Python side.
Good cancellation support in the language is really critical here, I found. In Python, it was a breeze to add timeouts and get rid of long running requests, if, say, the network connection dropped: you cancel the future, and that cancellation propagates to all the sub-futures. It is even hookable so that one can — if the network protocol supports it — propagate that across the wire to other services.
The DB in our case was written in Go, however, so that was tougher. Golang's best method (that we learned of at the time) is to thread a "Context" object through your code paths. We were working with existing code, of course, and it lacked this, and it's harder to add in hindsight.
Of course, once we got the server to stop hanging on queries of doom and return a more appropriate "that's a query of doom, and would hang the server" error, the complaint was that the server wasn't executing those queries anymore…
async def process():
try:
await db.slow_operation()
except CancelledError:
synchronous_functions_work()
await db.cancel() # This future will not complete on timeout
p = process()
try:
await asyncio.wait_for(p, timeout=1)
except TimeoutError:
await p # Required for db.cancel() to run!!!
Now fire-and-forget on a timeout is perhaps the most reasonable approach, otherwise you'd get timeout on timeouts, so better implementation would be to restart p without awaiting it, or putting it on a background cancel-list. But it can be really confusing when you are not aware of this behavior.Edit: Seems they actually fixed/changed this in 3.7: https://bugs.python.org/issue32751. So instead you have to write robust except-blocks that must never timeout.
Default timeouts in the database layers are hidden time bombs that turn operations that just legitimately take a bit longer than some value the library author set that you didn't even know existed into failures that get retried over and over causing even more load than just doing the thing once. Don't get me wrong there are lots of uses for setting strict timeouts and being able to do so is very important, but as a default no thanks.
There is no great OS solution for handling this. You kind of need to run async IO on the lowest layer, and at least still be able to receive the read readiness and associated close/reset notification that you can somehow forward to the application stack (maybe in the form a `CancellationToken`)
* Keepalive - Have the server ping back on a short timeout while it's working. Use a very long timeout for the server response.
* Asynchronous queues - Use queues for requests and discard traffic/error out when the queue becomes full.
* Idempotence - Send another request if the first one does not return in a reasonable amount of time.
* Broadcast - Don't fetch the information, have it sent to you through UDP. Great for cumulative metrics. If you miss one, no problem, the next packet has the same data.
* Cancellation - Cancel the request if you don't get an answer.
* Multiple requests - Send requests to multiple services and return the one that gets back first.
Forcing clients to pick timeouts amounts to punting a hard problem over to somebody who has even less idea how to solve it than you do.
Edit: clarity
That last suggestion (Multiple requests) can be tough to implement correctly, I think it's usually called Happy Eyeballs: https://youtu.be/oLkfnc_UMcE?t=290
You know we have a firewall that randomly disconnects connection and block traffic. A lots of apps just gets confused when that happened. And that happens a lot, once few minutes or shorter maybe.
When I work on my proxy, I had do define a new strategy to detect dead connections, such as to use separated timeout for Dial and Read. The Dial timeout will be a shorter value defaulted at 20 seconds, the Read timeout will be a rather normal one usually defaulted at 120 seconds.
I found many software just don't use any strategy. They just sends the connection and wait, assuming everything will be fine while it's actually hanging forever on user's end, until the OS kick them out. Many download system don't even have retry/resume mechanism: You downloaded 99% of a 600mb package (and it takes about 48 hours), then the connection EOF'ed, the software say "yeah you better download all of it again, hehe".
An example of good strategy can be found in `apt`. The software detects slow network, timed out connection, automatically retry downloads (not sure if it can resume download, could be great if it did). And all of that gives me a strong software that I can trust: I know when I run the command, the command will try it's best to get things done. And usually it did, causing far fewer issues than `npm`, `snap`, `git` and etc.
I suggest everybody give this mindset a try: When your software downloads data and puts it on user's computer, the copy of data is now owned by the user. You remove the data, you're looting the user from what they've got. It's like that you're making a dinner for your user: Everything been made (downloaded) is already on the table, one failed meal (packet) should not cause you to flipping the table. Instead, you retry and retry until it can't be done (For example, the source has changed or wait time is really too long).
When I have flaky connections, I get so many pauses (of songs I have downloaded) that I just deleted the app alltogether.
https://github.com/facebookarchive/augmented-traffic-control
Google, Apple and Microsoft also have ways to simulate bad connections but Facebook seems to be the best at it. I heard they even encourage their employees to switch to 2G from time to time.
This is especially noticeable in their web apps because they rarely reimplement the functionality which is built in to the browser. When I’m on the subway, mbasic.facebook.com has things like a working reload but the app and Facebook.com will both fail to do elementary error handling and will often discard whatever you entered.
So for one app I write the user (say a company executive who always complained his photo uploads failed but was too busy to give a detailed bug report or even a specific time it happened) would take pictures at home or office, which starts uploads on WiFi, then put the phone in their pocket and get into their car and drive away. So I ended up using relatively short timeouts (30 seconds) and automated retries.
The concept of a "soft mount" with a timeout was introduced to NFS but it's almost never recommended. This is because client programs have no idea how to handle a timeout from the filesystem. This article shows how every HTTP client has to be configured to handle failures. Imagine if every program that accesses a file, from /bin/cat all the way up, had to have error handling code to deal with timeouts and retries. A sane choice is to wait infinitely if there's nothing more intelligent that you can do.
Consider a startup script that hangs forever. Better it fail than hang. Or ls hanging forever when instead the filesystem could fail the operation after 30s.
There is a longstanding issue [1] to add a default timeout, but so far that hasn't happened yet.
Also important to consider what the client is advised to do in the case of a timeout. Retries, for instance should likely have backoff and jitter attached, or a retry budget.
Internal services have extremely low response time during normal operation (p99 around a second) but then the database will start a snapshot or a large analytics query hits on the week end (high IO) and the latency is through the roof for a short while. Too bad if services have short timeouts, they're all failing all requests now for no reason.
p99 is normal operation. Services shouldn't be configured to systematically fail for 1% of operations.
YES 1000% (I mean maybe not 'top' but it's up there)
languages must move connection pooling, timeouts, and retry semantics into the stdlib
API client libraries have to do a better job of documenting what happens when a request fails
systems need to do a better job of centralizing how timeouts are configured; this can't be left to chance
Also you can have millions of processes per core, with minimal performance regression, do you're likely to notice it in monitoring before it becomes a problem.
https ://hexdocs.pm/elixir/GenServer.html#call/3
Note that call is even more sophisticated, it incorporates a liveness check on its counterparty to quit out before timeout if there's been a catastrophe. Note this is effectively a simple form of backpressure management that degrades availability gracefully under stress in favor of ensuring the integrity of existing connections. Because fallibility is so baked into the runtime, every good library incorporates timeouts where it's sensible.
For connection pools, it's not explicitly part of the standard library but basically everyone (maybe not whatsapp) uses poolboy:
https://elixirschool.com/en/lessons/libraries/poolboy/
I have never used it myself, yet, but I trust more experienced devs have incorporated it successfully in ecto (rdbms) and Phoenix (web framework).
There are a dozen or more timeouts just for a TCP connection. There's the initialization timeout, the 3-way handshake timeout, the half-closed timeout, the time-wait timeout, the unverified reset timeout, the established connection timeout, the retransmission timeout, the timed wait delay, the delayed ack timer, the arp cache timeout, the arp cache minimum reference timeout, the keep-alive timeout, and more.
Every single person in the world depends upon default timeouts, so of course they matter. When they are picked intelligently, they improve the default behavior of the majority of system interactions. So we can trust default timeouts, when they are useful. But if we're building a system, it makes sense for us to determine what the appropriate timeout is for our system.
Uh what? Has he never heard of Fetch:
https://developer.mozilla.org/Web/API/Fetch_API
its been around for at least 5 years, and it returns a Promise.
Many of these things are used for one-off scripts, where it isn't worth thinking about. For many APIs, it isn't worth the trouble - if one of your dependent services is unresponsive, there isn't really any meaningful thing your application can do anyways. It doesn't become an issue until there are so many timeouts that it's impacting other resources. Best to leave it off until you know what you want to do with it.
net.netfilter.nf_conntrack_dccp_timeout_timewait = 240
net.netfilter.nf_conntrack_frag6_timeout = 60
net.netfilter.nf_conntrack_generic_timeout = 600
net.netfilter.nf_conntrack_gre_timeout = 30
net.netfilter.nf_conntrack_gre_timeout_stream = 180
net.netfilter.nf_conntrack_icmp_timeout = 30
net.netfilter.nf_conntrack_icmpv6_timeout = 30
and even for TCP, there is a timeout after the connection is closed. The fact that UDP has no state and therefore no 'connection' doesn't mean that just because TCP does, that conntrack only tracks it while the connection is open. Besides, you could sever a cable and TCP wouldn't know that anything happened. So you do need timeouts for anything in a NAT table. package http
var DefaultClient = &Client{}
It is a convenient global variable that uses whatever last settings were set upon it from any bit of code executed in any dependency.Never say never. What can be tuned in the system (obviously relevant only for server software) is better tuned there unless you really like (re-)negotiating tuning options with ops.