Defeat your 99th percentile with speculative task
bjankie1.github.io
bjankie1.github.io
(regarding "Speculative task") Also called "Hedged Request" here in an article called "The Tail at Scale"[1]
[1] http://www-inst.eecs.berkeley.edu/~cs252/sp17/papers/TheTail...
https://static.googleusercontent.com/media/research.google.c...
Does he elaborate on this?
* one of your servers is unhealthy, remove it from the pool.
* reject some requests to make the rest faster.
speculatively :: Int -> IO a -> IO (Either a a)
speculatively ms action = race action (threadDelay ms >> action)I mean if this is indeed helping it seems that it's not Redis that's the long pole, but maybe the request routing mesh or something.
I don't mean to be negative of the solution, I am just curious.
However, in my experience latencies are not static and depend on how far away the request is sent, the type of the resource requested, the size of the resource, current network load in that direction and other factors. Which gets tricky and complicated. At some point you need to store latest latency history for each request per each size group per each resource type per each node and dynamically calculate 90th percentile latency. But then things like size may not be predictable, so you may need to cap response sizes to a sufficiently small value. And so on.
If your responses are small, it's easier to just always send two requests in parallel to different servers and choose the fastest one.
For all you know you're in the middle of a horizontal scaling event and hit an overloaded box, and there is no problem.
There is no point in having your service try to speculate what happened when that isn't its job.
As long as your service doesn't need to ensure things happen only once or you build in that recovery mechanism (identifying unique events and throwing them out if your system has already seen them), being able to toss work to another instance is great in my mind.
Like philsnow says in a different comment (https://news.ycombinator.com/item?id=16832566)
> If your backend is Redis, how is starting more speculative hits to Redis going to help, since it's single threaded?
I agree; really understanding this problem is a lot better than just blindly retrying because it seems to work. Is there something wrong with the network that'll affect every service?
Retrying could be the solution, but it should be so with the acknowledgement that it's incurring technical debt.
You can do none, one of them, or both. You may have team skills to tackle none, one, or both. You may or may not control / have access to one of the sides. You may have a timeline for solving this problem where you decide which task is faster to finish. Finally, you may have a recurring problem which makes you lose a specific amount of money every time and a speculative request stops that now, while putting people on a performance debugging mission may only potentially have a positive return in the future.
Sure, it's a technical debt. It's not illegal - you just need a really good reason to commit to it.
Maybe. But again, this condition can occur with no flaws. It may just be you're on the edge of an autoscale event or a box with a bad xen peer. There may not a be a problem. It doesn't matter.
So sure, folks should keep to their SLA, but it doesn't have a lot of bearing on this. It's just best practice.
Oh also, because so many people use languages where even a basic production-quality binary tree is not easy to write. in.
Usually, it doesn't matter why the consumer finds it slow; it could be anything from an unusual I/O issue (AWS is notorious for brief but painful networking issues) to a VM peer suddenly becoming ill behaved before being throttled.
They're also notorious becuass ECS is SO BAD.